ISO 27005 / EBIOS RISK MANAGER / DEBIAN 13
I assess the risk first.
Then I deploy SysWarden.
ISO 27005 and EBIOS RM compared on the same case. Risk analysis, residual risk acceptance, then a complete rollout procedure for 500 Debian 13 servers.

01 / THE DECISION
I start with the service I need to keep running.
Five hundred servers is an interesting deployment target. It is also five hundred opportunities to repeat the same mistake. Before I put SysWarden on that many machines, I want a clear answer to a less glamorous question: what business loss am I trying to prevent, and what new failure could my own deployment introduce?
As SysWarden's creator, I care about its technical capabilities. I also have a responsibility to describe their limits. Installing a security product changes a system. It introduces privileged code, configuration, dependencies, operational work and a new part of the software supply chain. I include those changes in the risk assessment with the threats the product is intended to address.
In this article, I work through a fictional deployment of 500 Debian 13 AMD64 servers. I first perform an ISO/IEC 27005-guided risk assessment, then analyse the same case with EBIOS Risk Manager and compare the results. I take the treatment and residual risk decisions to the authorised owners. Only after that acceptance do I proceed to the complete staged deployment procedure.
This is my proposed engineering approach. The fleet, ratings, budgets, thresholds and schedules are illustrative. They are not a customer case study, a claim that SysWarden has been benchmarked on a 500-server estate, or a report of an accredited assessment. A production team would replace every assumption with its own evidence and obtain the necessary decisions from its risk owners.
The result I want
I want to be able to explain why the deployment is justified, which controls it relies on, what I have actually tested, what remains exposed and who has accepted that exposure. A successful package installation is one input to that decision.
I have organised the article so that a security leader can follow the decision trail and a Linux operator can follow the execution trail. Both need to meet at the same point: a service that remains usable, observable and recoverable.
02 / TWO COMPLEMENTARY TOOLS
The two approaches I will apply and compare.
ISO/IEC 27005:2022 provides guidance for managing information security risks in support of an ISMS. Its scope includes assessment, treatment, communication, monitoring and review. I use that lifecycle to keep the assessment connected to decisions and ongoing operations. The ISO publication page identifies the fourth edition, published in October 2022. [1] ISO/IEC 27005
ANSSI's EBIOS RM method combines a security baseline with scenario analysis. Its five workshops move through scope and baseline, risk sources, strategic scenarios, operational scenarios, and risk treatment. The scenario approach focuses on intentional, targeted threats. I use those workshops for the second analysis, connecting business consequences to plausible attack paths. [2] ANSSI method overview
I keep accidental outages, operator mistakes, component failure and capacity exhaustion in the wider risk register as well. I do not force a broken disk or a mistaken configuration into the role of an adversary. Those risks still influence deployment design, recovery and acceptance.
| Question | How I address it | Record I retain |
|---|---|---|
| What matters, and how will I judge risk? | Define the scope, criteria, responsibilities and decision process. | Assessment charter and acceptance criteria. |
| How could an adversary cause serious harm? | Use EBIOS RM workshops to develop and challenge scenarios. | Scenario sheets and explicit assumptions. |
| What will I change? | Choose treatments, owners, resources and tests. | Treatment plan linked to each risk. |
| What remains after implementation? | Reassess against observed control effectiveness. | Residual risk decision with an expiry and review triggers. |
I do not describe this mapping as an official equivalence table. It is a practical way to organise this example. I refer to the actual standard and ANSSI publications when conducting a formal engagement; this article does not reproduce their full requirements, scales or workshop instructions.
I also separate the organisation's assurance from the product's assurance. Using SysWarden does not confer ISO certification. Using an EBIOS-inspired analysis does not create ANSSI approval. SysWarden release IVV and an organisation's deployment acceptance answer different questions. An intermediate release's IVV evidence can inform my assessment; it cannot decide whether my particular business service should accept the remaining risk. The project's release assurance documentation explains the distinction between IVV and IVVQ.
03 / PART I: ISO 27005
I establish the context for my ISO 27005 assessment.
I start with the guidance in ISO/IEC 27005, before applying EBIOS RM to the same case. I use one fictional organisation, one set of business objectives and one evidence baseline for both analyses. Otherwise, a difference in the results could simply reflect a different starting point. The following records are my own implementation of the risk process, not templates prescribed by ISO.
My decision is whether, and under which conditions, to introduce SysWarden across 500 Debian 13 AMD64 hosts supporting a transaction service. I include confidentiality of records, integrity of transactions, service availability and trusted administration. The scope includes delivery, identity, monitoring, backup and supplier dependencies, as well as the servers themselves.
I appoint a business risk owner for each service, a coordinator for the assessment and owners for technical treatments. I record the organisation's obligations and existing commitments, the decision horizon and the authority required to accept each risk level. I agree which evidence can be collected and how it will be protected. My own technical recommendation does not replace the business owner's decision.
I then define the baseline: the controls in operation today, their known gaps and the evidence supporting them. I assess both the risk of leaving the current estate unchanged and the risk introduced by the proposed security deployment. That comparison stops the project from receiving credit merely because its purpose is security.
04 / MAKE THE RATINGS MEAN SOMETHING
I agree the risk criteria before seeing the results.
I use four impact levels and four likelihood levels for this example. I keep their meanings stable across scenarios. The business owner supplies thresholds suited to the service, such as tolerable interruption, recovery objectives, material data loss and contractual consequences. I avoid assigning arbitrary currency values without that input.
| Level | Impact | Likelihood within the study horizon |
|---|---|---|
| 1 | Limited disruption, contained within routine operations. | Unlikely: substantial prerequisites, with strong supporting evidence for the barriers. |
| 2 | Noticeable service degradation requiring coordinated recovery. | Plausible: a credible path exists, but meaningful barriers remain. |
| 3 | Major interruption or significant loss affecting business commitments. | Likely: attainable prerequisites and material control weaknesses. |
| 4 | Critical harm threatening a core mission or creating intolerable loss. | Very likely: a readily accessible path with weak or ineffective barriers. |
These labels are ordinal judgments. Level four is not twice the probability of level two. I do not multiply them and then present the result as a measured annual loss. Where I have defensible quantitative data, I can use a quantitative model, but I would show its assumptions, uncertainty and sensitivity separately.
| Impact / Likelihood | 1 | 2 | 3 | 4 |
|---|---|---|---|---|
| 4: Critical | High | Critical | Critical | Critical |
| 3: Major | Moderate | High | High | Critical |
| 2: Noticeable | Low | Moderate | Moderate | High |
| 1: Limited | Low | Low | Moderate | Moderate |
In this example, a critical risk prevents production rollout. A high risk requires treatment and a documented decision by the designated senior risk owner before any limited exception. A moderate risk may be accepted for a defined period with named actions and monitoring. A low risk still needs an owner and a review date. Legal, contractual and safety obligations remain independent constraints; a coloured cell cannot waive them.
I record confidence beside the rating. A plausible scenario supported by a recent restoration test is different from the same rating based on an old design diagram. If uncertainty could change the deployment decision, I investigate before accepting the risk. Missing evidence is a reason to qualify the judgment, not a reason to assume the best outcome.
I maintain a short rationale for disagreements. If the application owner expects rapid recovery but the platform team has never restored that application, the discrepancy becomes a test requirement. This is where a workshop starts producing useful engineering work.
05 / ISO 27005: ASSESSMENT
I identify, analyse and evaluate the risks.
I write each risk as a cause, an event and a business consequence. For example: an overprivileged delivery account could distribute an unauthorised package, compromising the trustworthy operation of several services. I link the statement to affected assets, weaknesses, existing safeguards and a named owner. A list of vulnerabilities alone would leave out that chain of consequences.
| Record | Cause, event and consequence | Initial judgment |
|---|---|---|
| R1: Delivery | Compromised privileged distribution permits unauthorised code and widespread loss of trusted operation. | Impact 4, likelihood 3: Critical. |
| R2: Data access | An application foothold crosses excessive network and identity permissions, exposing protected records. | Impact 4, likelihood 3: Critical. |
| R3: Hostile policy change | A malicious privileged actor changes policy and interrupts service or administration. | Impact 3, likelihood 3: High. |
| R4: Concealed failure | An attacker suppresses monitoring and delays detection of a compromise. | Impact 3, likelihood 3: High. |
| O1: Accidental change | A configuration mistake propagates through automation and interrupts legitimate traffic. | Impact 3, likelihood 3: High. |
| O2: Dependency failure | A delivery or intelligence dependency becomes unavailable, degrading installation or protection freshness. | Impact 2, likelihood 3: Moderate. |
I analyse likelihood using the prerequisites and safeguards for each event, then assess the business consequence. These illustrative judgments expose priorities; they are not measured probabilities. I would support a real rating with incident history, access reviews, dependency tests, recovery results and informed specialist judgment. I record uncertainty instead of disguising it with precise-looking arithmetic.
I evaluate the results against the criteria agreed above. R1 and R2 require further treatment before production deployment. R3, R4 and O1 need an explicit treatment decision. O2 still needs an owner and a dependency strategy. I also examine concentration: one controller reaching all 500 hosts is a common cause, not 500 independent small risks that disappear when averaged.
My output is a prioritised risk register with rationale, confidence, current controls and evidence gaps. I can already make useful decisions with this ISO-guided assessment. I have not yet claimed that a proposed control works or that any residual risk is accepted.
06 / ISO 27005: TREATMENT AND REVIEW
I turn the assessment into decisions I can fund and verify.
For R1, I propose independent release authentication, separated approval and execution, and deployment credentials limited by service boundary. For R2, I propose reviewed connectivity and application authorisation. For O1, I propose small batches, independent health probes and demonstrated recovery. For O2, I propose authenticated package staging and a tested policy for unavailable or stale intelligence. These are engineering choices for this example, not a product feature checklist.
I compare those choices with avoiding an unsuitable deployment profile, redesigning a dependency, sharing some consequences through a contract, or retaining an exposure through an authorised decision. I cost implementation, ongoing operation and recovery. A cheap control that nobody can maintain can create a different risk.
I separate the expected residual risk from the risk supported by evidence after treatment. I require existing organisational controls to be effective and the proposed technical design to pass representative tests before requesting production authorisation. Where a control will only exist on a host after rollout, the decision records that dependency and the deployment verifies it before admitting the host to service.
I communicate the result to the owners, record unresolved disagreements and define review triggers. The register stays alive when assumptions change. I now have an ISO-guided assessment and treatment proposal. Next, I use the five EBIOS RM workshops on the same case to examine the intentional scenarios more closely, reconcile any differences and strengthen the same decision record.
07 / PART II: EBIOS RM, WORKSHOP 1
I define the boundary before I inventory the servers.
I now apply EBIOS RM to the same case. I reuse the approved scope and criteria rather than silently changing the comparison. For my example, the business mission is to deliver a customer-facing transaction service, protect the records it processes and recover within agreed service objectives. The study covers the Linux hosts and the management, monitoring, backup and delivery systems that can affect them. A supplier is not outside the analysis simply because its infrastructure is outside my control.
I set a twelve-month decision horizon and review the assessment after material changes. I use a shorter operational horizon for each deployment wave. These are choices for the example, not mandatory ISO or ANSSI periods. The business owner, security lead, platform lead, application owners, incident response lead and supplier manager take part in the relevant decisions.
| Service class | Hosts | What I need to preserve |
|---|---|---|
| Public web and reverse proxy | 160 | Customer access, TLS termination and controlled backend connectivity. |
| Application and worker services | 180 | Authenticated processing, queues and restricted service dependencies. |
| Data services | 80 | Integrity, confidentiality, replication and tested recovery. |
| Management and observability | 40 | Administration, alert delivery and evidence collection. |
| Shared infrastructure | 40 | DNS, time, package distribution and other explicitly inventoried dependencies. |
| Total | 500 | These are planning counts, not a capacity claim. |
A server count does not tell me how much simultaneous change is tolerable. I map hosts to applications, clusters, failure domains, sites and management paths. Two servers in different inventory groups may still be the only two members of the same critical service. My orchestration must understand that relationship before either machine enters a deployment batch.
I distinguish business values from supporting assets. Customer transactions, trustworthy records and the ability to administer the service are business concerns. Debian hosts, identities, configuration repositories, package stores, network paths and backup systems support those concerns. Losing an individual host matters because of the consequences it can create.
I write feared events in business language: transactions unavailable beyond the agreed recovery window; unauthorised disclosure of protected records; undetected alteration of transaction data; and loss of trusted administration during an incident. Each event has an accountable owner. I ask that owner to explain the operational and financial consequences, rather than letting a vulnerability score stand in for business impact.
I make the baseline explicit.
My baseline includes supported operating systems, timely patching, strong administrative authentication, controlled privileges, segmentation, tested backups, central monitoring and change control. For each measure, I record whether it is implemented, partly implemented, absent or unverified. I attach an evidence reference and a date. An undocumented assumption stays unverified.
I also inventory what already changes host policy. That includes nftables, compatibility rules, container networking, VPNs, configuration management, local security agents, logging and scheduled maintenance. I assign an owner to each shared resource. If two systems both believe they own the entire firewall, I resolve that design conflict before adding a third.
The output of this workshop is a bounded study, a dependency map, a list of feared events and a baseline gap register. I can then explain which problems SysWarden is being evaluated to help address and which require other work.
08 / WORKSHOP 2
I pair credible risk sources with their objectives.
For the intentional threat analysis, I select risk sources and objectives that make sense for the fictional service. I consider motivation, capability, access and relevance. I do not keep an exotic adversary in scope merely because it makes a presentation look serious.
| Risk source | Objective | Why I retain it |
|---|---|---|
| Financially motivated criminal group | Extort the organisation through disruption or data theft. | The service has valuable records and availability commitments. |
| Attacker controlling a supplier account | Use a trusted relationship to reach privileged systems. | External dependencies connect to management or delivery workflows. |
| Malicious person with legitimate access | Alter policy, steal data or suppress evidence. | Privileged paths exist and require accountable controls. |
| Attacker targeting software delivery | Distribute a malicious privileged component. | A common deployment channel can affect many hosts. |
I give each pair a short selection rationale and retain the reasons for exclusions. An opportunistic scanner may be handled primarily through the baseline; a targeted compromise of the automation controller deserves a deeper scenario because of the concentration of privilege. I revisit that selection when the organisation's exposure changes.
I ask a practical question for every retained objective: what would success look like to the attacker? For extortion, it might be a sufficiently disruptive loss of service plus leverage over recovery. For data theft, it might be sustained access to the application data path. That prevents me from treating every malicious packet as if it had the same business significance.
I keep accidental change risk in a parallel record. An administrator can deploy an incorrect rule without malicious intent and still cause a serious outage. The prevention and response may overlap with the malicious-policy scenario, but the initiating event and likelihood evidence differ.
09 / WORKSHOP 3
I follow trust through the ecosystem.
I draw the relationships that let work happen: developers publish changes, automation distributes packages, operators access hosts, hosts obtain intelligence data, applications reach databases, and monitoring collects events. Every relationship carries a purpose and a degree of trust. I look for the relationships whose compromise could cross several boundaries at once.
My first strategic scenario is a delivery compromise. An attacker obtains control of a privileged distribution path and uses it to place a harmful package or configuration on a large part of the estate. The business consequence is a fleet-wide loss of trustworthy operation. A clean firewall on one host does not address the authority of the system delivering code to every host.
My second scenario starts with an internet-facing application. An attacker gains an application foothold, reaches a more sensitive service through excessive connectivity, then steals or alters valuable data. Host filtering can restrict paths, but application authorisation, credential boundaries and data protection remain central to the scenario.
My third scenario involves a service provider or support relationship. A compromised identity reaches the management plane and changes policy or disables evidence collection. I examine access scope, session accountability, expiry, supplier incident notification and the organisation's ability to revoke the relationship without losing recovery access.
- DeliverySource, build, release identity and internal distribution.
- AdministrationPeople, automation credentials and independent recovery.
- Host enforcementLocal policy, application dependencies and privileged runtime.
- EvidenceMonitoring, independent probes and protected records.
I ask whether the same identity can publish an artifact, approve its deployment and erase its evidence. I ask whether one management account reaches all 500 hosts. I ask whether recovery depends on the same provider account being attacked. These answers shape the treatment plan before I choose a batch size.
At this stage, I can already reduce exposure through organisational and architectural measures: separate approval from execution, restrict distribution privileges, define supplier escalation, divide administration into appropriate scopes and maintain an independent recovery path. Those are decisions around SysWarden, not features I attribute to it.
10 / WORKSHOP 4
I turn scenarios into observable tests.
I now describe the operational paths in enough detail to identify prerequisites, controls and useful observations. I keep the description focused on defensive validation. A scenario sheet should tell an operator what must be tested and why, without becoming a catalogue of speculative attack techniques.
| ID and path | Control hypothesis | Evidence I would require |
|---|---|---|
| R1: Compromised delivery identity distributes unauthorised privileged code. | Independent artifact verification and deployment approval interrupt distribution. | Wrong identity, altered package and unapproved version are rejected before installation. |
| R2: Compromised public application reaches an unauthorised data service. | Segmentation limits the permitted path; application credentials limit permitted actions. | Negative connectivity tests plus application authorisation tests, including IPv6 where enabled. |
| R3: A malicious privileged policy change blocks administration and service traffic. | Small batches, independent probes and a tested recovery path contain the change. | Observed stop condition, successful console access and timed restoration rehearsal. |
| R4: An attacker suppresses collection to conceal a compromise. | Independent health monitoring and alert delivery identify loss of visibility. | A deliberately interrupted test collector triggers a received, actionable alert. |
For R2, I separate the application's network path from the application's authority. An allowed connection may still carry an unauthorised request. A valid service credential may allow a harmful action over a perfectly legitimate connection. I test those controls with the application owner; I do not give a host firewall credit for enforcing business authorisation.
Alongside the intentional R3 scenario, I retain the accidental O1 scenario from my first assessment. A rule that blocks a malicious source can also block a partner, a monitoring probe or a shared egress address if the policy is poorly chosen. I establish the legitimate traffic baseline and the exception process before enabling broader blocking policies.
For R4, I measure the complete alert path. A local log line is only the beginning. The event must reach the intended collector, match the intended rule, arrive in a queue someone watches and carry enough context to act. I also test non-malicious collector failure separately and test the absence of events: a silent agent and a quiet environment must not be indistinguishable.
I estimate likelihood against the controls that demonstrably exist today. I then record a separate target rating for proposed treatments. I keep the target visibly marked as unverified until implementation and testing support it. NIST's risk assessment guidance is a useful additional reference for communicating assumptions and uncertainty, although I am not replacing the ISO and EBIOS structure with a third methodology here. [3] NIST SP 800-30 Rev. 1
11 / WORKSHOP 5
I give every treatment an owner and a test.
My treatment plan can reduce a risk, avoid the activity creating it, share part of its consequences through an appropriate arrangement, or retain it through an authorised decision. I document the reasoning. A support contract can improve access to assistance; it does not remove my organisation's responsibility for service continuity.
For each action, I record the linked scenario, accountable owner, expected benefit, cost, dependencies, due date, implementation evidence and effectiveness test. ANSSI's treatment method sheet recommends practical, measurable actions with ownership and scheduling, maintained through the system's lifecycle. [4] ANSSI treatment guidance
| Action | Accountable role | Evidence required before rollout |
|---|---|---|
| T1: Authenticate and freeze the release bundle. | Release engineering lead | Verified signatures, provenance policy and immutable distribution record. |
| T2: Define and test each host policy profile. | Platform lead | Approved configuration, positive and negative traffic tests, dependency review. |
| T3: Rehearse recovery without normal SSH. | Infrastructure lead | Console access and timed restoration of a representative service. |
| T4: Make failures visible outside the affected host. | Detection and response lead | Received alerts for enforcement, feed, collector and application failures. |
| T5: Limit deployment authority and concurrency. | Automation lead | Scoped credentials, explicit inventory, approval gates and stop-condition rehearsal. |
| T6: Validate the full product lifecycle. | Platform lead | Installation, upgrade, reboot, recovery, removal and reinstallation results for the approved paths. |
I budget for those tests as part of the deployment. Five hundred installations may be inexpensive to automate, while five hundred poorly understood recoveries are not. My estimate includes engineering time, test capacity, duplicate storage, maintenance windows, monitoring work and the people needed to investigate a failed wave.
I also reserve resources for independent review. A code review can examine trust boundaries and privileged behaviour. An authorised penetration test can challenge the deployed design. Both need a scope, suitable test systems, a remediation budget and a retest. I never present an audit that has not taken place as a control already implemented.
Where I put SysWarden in that plan.
I evaluate SysWarden as a Linux host defence component. Its configuration and enforcement must fit the service architecture, the existing firewall ownership model and the organisation's operational procedures. I use the project's architecture guide and the exact release's source and documentation to determine the applicable behaviour.
I do not use it as a substitute for application security, identity governance, patching, upstream network protection or backups. A host rule cannot restore corrupted business data. A local agent cannot reliably attest a host already controlled by a privileged attacker. A server cannot filter its way out of a saturated upstream link without help elsewhere in the network.
I decide separately whether optional VPN, HA, geographic policy, ASN policy or external integrations are justified. An option earns its place through a service requirement and a tested operating model. I would keep an unnecessary option out of the initial deployment rather than create another dependency to support.
12 / THE COMPARISON
What each approach changes in my SysWarden decision.
I compare the approaches using the same scope, criteria and starting evidence. ISO/IEC 27005 is guidance for managing information security risk; EBIOS RM provides a workshop-based method. I can implement the former using the latter, but I do not pretend their formats or levels of prescription are identical.
| Dimension | My ISO 27005-guided assessment | My EBIOS RM analysis |
|---|---|---|
| Starting question | What could prevent the service meeting its objectives, and how should I manage that risk? | Which feared events and targeted scenarios challenge the security baseline? |
| Structure | I design the assessment records and decision workflow around the organisation's risk process. | I organise participation and outputs around five workshops. |
| Coverage in this example | I include intentional, accidental and dependency risks in one register. | I deepen intentional scenarios; I keep accidental risks in the wider register. |
| Participants | I involve risk owners and specialists according to the decision and evidence needed. | I use business and ecosystem discussions first, then focused technical scenario work. |
| Most useful outputs | Criteria, prioritised register, treatment decisions, acceptance and review records. | Risk-source objectives, strategic and operational scenarios, treatment measures and follow-up. |
| Effort I need | A repeatable process, credible evidence and accountable decisions. | The same discipline plus prepared workshops and people who understand the ecosystem and attack paths. |
| Failure mode I watch | A spreadsheet detached from real technical and business conditions. | Convincing scenarios with weak evidence, omitted dependencies or no funded treatment. |
| Value for this deployment | I keep rollout risk, availability and long-term ownership in the decision. | I challenge concentration of trust in delivery, management and monitoring. |
R1 makes the difference concrete. My first assessment identified privileged distribution as a critical risk. The EBIOS scenario work asks which actor could compromise which relationship, how that actor would reach the controller, and whether an independent barrier would actually interrupt the path. If the same identity controls verification and execution, I revise my treatment instead of counting two nominal controls.
I do not expect a method to produce a conveniently lower score. If both analyses use the same facts, an unchanged rating can be the correct result. When EBIOS reveals an additional path, I update the existing risk rather than double-count it as an unrelated item. When the wider assessment reveals an accidental failure, I retain it even though it is outside the intentional scenario analysis.
For this fleet, I would combine the lifecycle discipline with the workshop detail. I would tailor effort to complexity and uncertainty, not prescribe a universal number of consulting days. My comparison ends with one reconciled treatment plan and one accountable acceptance process.
13 / ACCEPTANCE BEFORE PRODUCTION
I obtain acceptance before authorising deployment.
Before the production rollout, I reassess using implemented organisational safeguards and representative preproduction evidence. A planned control has an estimated effect; it is not evidence of an already protected production host. I distinguish the original baseline risk, the target I hoped to reach and the residual risk I can now justify. I retain the reasoning when a rating changes. If the treatment mainly reduces likelihood, I do not also reduce impact without evidence that the consequences have been contained.
| Risk | Baseline | Residual assessment | Decision still required |
|---|---|---|---|
| R1: Delivery compromise | Impact 4, likelihood 3: Critical | Impact 4, likelihood 1: High, if independent verification, approval and scoped deployment authority are effective. | Senior owner decision; high risk remains and may require further treatment. |
| R2: Application foothold reaches data | Impact 4, likelihood 3: Critical | Impact 4, likelihood 2: Critical until application and identity controls close the remaining path. | Rollout for that affected service is blocked by this example's criteria. |
| R3 / O1: Hostile or accidental policy change | Impact 3, likelihood 3: High | Impact 2, likelihood 2: Moderate only if bounded batches and demonstrated recovery limit consequences. | Time-limited acceptance with recovery readiness and wave monitoring. |
| R4: Loss of useful visibility | Impact 3, likelihood 3: High | Impact 3, likelihood 1: Moderate if independent health alerts and response are demonstrated. | Acceptance with alert-path exercises and coverage review. |
This first review rejects rollout for the affected service because R2 remains critical. I return to treatment. In the worked example, I assume the application owner implements narrower service identities and authorisation, and representative tests demonstrate that the remaining path is substantially harder. R2 is then reassessed at impact 4, likelihood 1: High. Its impact has not magically disappeared.
For the continuation of this fictional example, I explicitly assume that the designated senior risk owner accepts the remaining High R1 and R2 exposures for a 30-day controlled rollout, with scoped authority, monitoring and recovery conditions. The service owners accept Moderate R3, R4 and O1. O2 remains Moderate and is accepted with tested dependency-failure behaviour and freshness monitoring. These are illustrative decisions, not real approvals or a recommended appetite for every organisation. If the authorised owner declines any of them, I stop here.
My acceptance record identifies the risk, affected service, business owner, current rating, confidence, evidence references, remaining weaknesses, compensating measures, review date and conditions that invalidate the decision. It states who is authorised to accept that level of risk. The engineer implementing a control does not automatically have that authority.
An acceptance record I can revisit
For O1, my example record accepts a 30-day Moderate exposure before the first production wave, based on a successful preproduction rehearsal. The record would name the remaining configuration-error exposure, the tested recovery time, the approved concurrency, outstanding actions and the incident or drift conditions requiring immediate review. Thirty days is an example review period, not a methodological requirement.
I keep acceptance separate from the treatment backlog. An accepted risk can still have improvement actions, and a completed action can still leave an unacceptable risk. I also distinguish acceptance of a limited pilot from acceptance of the full estate. The scope of the decision matters as much as the signature.
The gate into the deployment procedure
I require the signed risk decisions, funded treatment plan, authenticated release record, approved configuration profiles, successful representative tests, available recovery route and explicit wave limits. I accept the remaining exposure before production execution. I verify the conditions on every host and reopen the decision if observations contradict them.
I distinguish this permission to perform a controlled rollout from final operational acceptance of all 500 hosts. The former comes now; the latter follows observed results. The next sections describe the complete execution and evidence procedure under that authorisation, including rechecking the lab evidence on which the decision depends.
I do not promise zero residual risk. Privileged compromise, application flaws, upstream outages, supply-chain concentration and human error remain possible. My responsibility is to make the remaining exposure visible and support an informed decision about it.
14 / PART III: AUTHORISED DEPLOYMENT
I make the fleet describable and reproducible.
With the risk decision accepted and its conditions recorded, I begin the deployment procedure. I create an inventory row for every host. At minimum, it contains the service owner, role, environment, architecture, Debian release, kernel, failure domain, current SysWarden state, administration route, recovery method, policy profile and maintenance constraints. I add references to the current backup and monitoring records. I keep credentials and private keys in the organisation's secret-management system, outside this inventory.
I distinguish fresh hosts from upgrades and from machines with unresolved installation state. I group upgrades by the exact source version and relevant configuration history. I do not assume that two Debian 13 hosts are equivalent merely because their operating-system labels match. Their service units, firewall loaders, VPNs, package state and operator customisation may differ.
For each profile, I define allowed business flows by source, destination, protocol and purpose. I cover IPv4 and IPv6 deliberately. I include DNS, time, package repositories, intelligence sources, monitoring, backups, replication and management. The final policy is reviewed by the service owner and tested from the appropriate network locations.
I identify shared services before introducing automation. An administrator-managed firewall, existing Fail2ban deployment or application security control keeps its own ownership boundary. SysWarden is not a reason to remove an unrelated protection. I need evidence that the intended controls can coexist, reload and survive a reboot without overwriting each other.
Debian 13 has release-specific details worth checking. Its release notes describe changes involving temporary storage, sysctl configuration and network interface naming. I use those notes to review the actual host configuration; I do not assume that every upgraded image has the same defaults as a fresh installation. [5] Debian 13 operational changes
The AMD64 requirement is deliberate. Debian supports several architectures, but distribution support does not imply availability of a SysWarden package for every architecture. I verify the product's published package matrix independently. [6] Debian architectures and SysWarden release inventory.
I keep the configuration review ahead of installation.
A native package can run privileged installation scripts and change policy during configuration. I therefore review the exact package's scripts, dependencies and configuration-loading behaviour before running APT. I do not assume that I can safely install first and decide on SSH, firewall or hardening settings afterwards.
I prepare the release-supported configuration through a reviewed, version-specific procedure in the lab. I record which files the package supplies, which files the product generates and which files the organisation owns. I test the actual precedence rules. I avoid cloning host identity, VPN keys or unrelated secret material as part of a supposedly reusable template.
For administration, I test the consequences of SSH forwarding restrictions, privilege changes, group membership and systemd confinement against the organisation's operating model. If a host profile depends on a behaviour that the selected release changes, I resolve the conflict before approval. I do not silently weaken the host's hardening to make an installation pass.
15 / ONE APPROVED RELEASE
I authenticate once, then distribute the same bytes.
I select an exact published release whose documented scope matches the deployment. I freeze the package version, asset names, hashes, signing identities, source references, configuration revision and verification-tool versions in a release record. An internal release ID ties them together. I do not resolve a moving latest target independently on 500 hosts.
The current installation guide provides the release inventory and version-specific authentication links. I follow the procedure for the selected release, including native package authentication and applicable signed metadata. I check the complete inventory, not just a package filename that looks right.
A digest answers whether the bytes match an expected value. I still need a trustworthy origin for that expected value. GitHub artifact attestations provide signed provenance claims; their value comes from verifying them against an expected identity and policy. They do not prove that the code is free of defects. [7] GitHub artifact attestations
My provenance policy constrains the repository, signing workflow and approved source reference or digest as appropriate to the release record. I account for the distinction between the product build commit and a later publication commit. I retain the verification result and reject an unexpected identity even when its signature is mathematically valid. The GitHub CLI documents these verification controls. [8] Attestation verification options
I import the authenticated bundle into an access-controlled internal artifact store. I keep the original bytes and detached evidence. Copying into an internal location does not replace upstream authentication; it adds a controlled distribution step. Each host checks its staged package against the approved digest before installation.
For disconnected environments, I plan the verification material and trust updates explicitly. GitHub documents offline verification using the artifact, attestation bundle and trusted root material. I record when those roots were obtained and how revocations or rotations will reach the disconnected environment. [9] Offline provenance verification
I treat intelligence feeds as a separate dependency from package delivery. An internal package mirror does not make runtime feeds available. I test the selected release's documented behaviour for source disagreement, unavailability, invalid input and stale local data, then decide what operational state my organisation will accept.
16 / THE HOST ADMISSION GATE
I stop unsuitable hosts before they enter a wave.
I begin with read-only inspection. The following commands are a starting point for an operator, not a complete fleet audit. Their output can reveal infrastructure details, so I keep it in the private deployment evidence store with restricted access.
cat /etc/os-release
dpkg --print-architecture
uname -r
sudo dpkg --audit
systemctl --failed --no-pager
sudo sshd -t
sudo nft -j list ruleset
systemctl is-active ssh.service
findmnt -no TARGET,FSTYPE,OPTIONS /
df -h / /var /var/tmp
I expect Debian 13, AMD64 and a package database with no unresolved transactions. I inspect existing failures instead of clearing them to get a green screen. I verify sufficient disk space and inodes for package staging, logs and recovery material. The actual resource envelope comes from measurement under representative load, not from the host count alone.
I check that APT sources are approved, dependency resolution is understood, and another package operation will not overlap the deployment. I preserve the installed package list and relevant service state. For a previously installed SysWarden host, I capture its exact version and inspect the corresponding migration guidance before deciding whether it belongs in this wave.
I test a fresh administrative connection from the approved management path. An already established SSH session can survive a policy problem that prevents all new connections. I also demonstrate console access through an independent route. I keep recovery credentials protected and available to the people responsible for the window.
I capture the effective firewall and its persistent loaders, relevant service definitions and overrides, configuration files, scheduler entries, and application probes. I label these records by owner. A backup of configuration is useful only if I know how to restore it without replacing an administrator's later changes.
I create an application-consistent recovery point appropriate to the workload. For a database, I involve the database owner and account for replication, write ordering and external dependencies. A VM snapshot alone is not proof that the service can recover correctly. I rehearse restoration on representative systems and measure how long it takes to restore useful service.
The admission decision has three outcomes: ready, excluded pending remediation, or outside the supported profile. I record the reason. A host that fails admission is not quietly counted as deployed, and it does not inherit the acceptance of a neighbouring machine.
17 / PROVE THE PROCEDURE FIRST
I enforce the preproduction evidence required by acceptance.
The acceptance above depends on this preproduction campaign already passing for the approved release and profiles. I describe its exact scope here so an operator can reproduce it. If the package, configuration or profile changes, I rerun the affected checks and revisit acceptance before production. I select representatives for each service profile and materially different starting state. That includes a fresh host, each approved upgrade origin, strict hardening, shared firewall ownership, relevant container or VPN networking, and the actual monitoring and backup paths. I also test the least powerful supported host profile to understand resource pressure.
I run the same package and configuration that I intend to deploy. The tests include installation, reconfiguration, service restart, policy reload, feed refresh and a real reboot. I repeat application probes and administration checks at the points where state can change. I include both expected allowed traffic and expected denied traffic.
I test failures deliberately in the lab: an altered package, an unexpected signing identity, unavailable dependencies, an interrupted maintenance step and loss of a required external service. I verify the documented failure result and recovery procedure. I preserve the failed run alongside the repaired run so the acceptance record shows what changed.
I distinguish product release evidence from deployment evidence. Release IVV may establish a specific implementation property on a specific matrix. My lab establishes whether the chosen package, configuration, automation and service dependencies work together in this organisation. I do not turn either result into an unsupported claim about every possible host.
I also validate the exit path before expansion. I test native removal and purge according to their separate contracts, and standalone uninstall only for an applicable standalone installation. I verify unrelated controls after removal, reload and reboot. This is lifecycle evidence, not a promise that removal reverses every earlier operating-system change.
I then rehearse an aborted wave. The automation stops; an operator receives the failure; the affected host remains excluded; recovery uses the approved route; and the inventory records the actual final state. If that rehearsal depends on someone remembering an undocumented command, I improve the procedure before production.
18 / FROM FIVE HOSTS TO FIVE HUNDRED
I expand only when the previous wave earns it.
I use seven production waves in this example: 5, 20, 50, 100, 100, 100 and 125 hosts. They total 500. A wave is an approval boundary, not a concurrency setting. A 100-host wave can still progress in batches of five, with a stricter limit for a particular service or failure domain.
| Wave | Added | Cumulative | Evidence before expansion |
|---|---|---|---|
| 1: Canary | 5 | 5 | One suitable representative per profile, independent access and application checks. |
| 2: Diversity | 20 | 25 | Additional dependencies and failure domains; one full operational cycle observed. |
| 3: Pilot | 50 | 75 | Resource behaviour, alert handling, reload and reboot results. |
| 4 | 100 | 175 | Review of previous results, exceptions and remaining capacity. |
| 5 | 100 | 275 | Same evidence gate, with no unresolved common-cause failure. |
| 6 | 100 | 375 | Service-owner confirmation and support readiness. |
| 7 | 125 | 500 | Final wave approval, followed by fleet reconciliation. |
I select the canaries deliberately. Five identical low-risk web servers would tell me little about a fleet with five different service classes. I choose replaceable members where possible, while preserving quorum and redundancy. A singleton or a host outside the tested profiles gets an individual maintenance decision.
I observe at least one relevant operational cycle before broadening the pilot. In this example that includes a scheduled feed refresh, a backup operation, a business traffic peak and the applicable reboot test. I might need more than a day to see those events. The observation period follows the service, rather than a countdown that expires while the system is idle.
I separate fleet waves from application draining. A stateless node can often leave a load balancer before maintenance. A database member requires the database owner's failover and replication procedure. I do not apply a generic restart loop to both and call it standardisation.
My automation stops at the wave boundary.
If I use Ansible, I invoke a reviewed play against a frozen inventory containing only the approved wave. I use serial to limit batches and explicit error handling to stop expansion. Ansible documents that a list of serial sizes continues with its last size; it does not create human approval between successive groups. [10] Ansible execution strategy
The following is an orchestration outline, not a supplied SysWarden role or a ready-to-run installer. The named task files are organisation-owned procedures that must exist, be reviewed and pass the lab campaign. The inventory is a single approved wave. Those constraints are part of the design.
# Design outline only. Implement and review every included task file.
- name: Deploy one explicitly approved wave
hosts: approved_wave
become: true
gather_facts: true
serial: 5
any_errors_fatal: true
tasks:
- name: Verify admission evidence and release identity
ansible.builtin.import_tasks: admission.yml
- name: Apply the tested release-specific installation procedure
ansible.builtin.import_tasks: install-approved-release.yml
- name: Verify host state and independent service probes
ansible.builtin.import_tasks: verify-service.yml
- name: Record the observed result
ansible.builtin.import_tasks: record-result.yml
I do not set ignore_errors around the installation or verification. I make unreachable hosts visible and stop the controller's next approval stage when results are incomplete. I handle rescue logic carefully: a successful rescue block can change Ansible's failure handling, so recovery must still produce an explicit failed-or-recovered deployment verdict that the external wave gate evaluates. [11] Ansible error handling
I set the controller's credentials and concurrency to the minimum justified scope. I keep the approval record outside the host being changed. I also prevent a second controller, unattended package job or scheduled product updater from racing this controlled operation, using a reviewed scheduling policy that preserves unrelated security updates.
19 / THE PER-HOST PROCEDURE
I use the same observable sequence on every admitted host.
- Recheck admission. I confirm the host still matches its approved inventory and that its recovery point, management path and service health remain valid.
- Prepare the service. I drain or fail over according to its own runbook, verify remaining capacity and prevent conflicting maintenance.
- Stage the authenticated bundle. I copy the approved package into a controlled local directory, verify its exact digest and retain the release-record reference.
- Prepare the reviewed configuration. I apply the version-specific configuration procedure already tested in the lab, preserving separate administrator ownership.
- Review the package transaction. I inspect the proposed dependency changes, then run the approved native installation operation and retain its result.
- Verify locally and independently. I check package state, configuration, services, effective policy, new administration sessions, application behaviour and alert delivery.
- Complete the lifecycle gate. I perform the profile's required reload and reboot checks, then restore the host to service only when all required evidence passes.
- Reconcile the record. I mark the host accepted, failed, recovered or excluded. The next batch follows that observed state.
For a staged, authenticated Debian package, APT can install a local file. The example below intentionally requires a path supplied by the approved release record. It does not download a release, authenticate a signature or prepare configuration. Those steps must already have succeeded, and the command changes the host.
# Run only after admission, authentication and configuration approval.
set -eu
# APPROVED_DEB must be an absolute path to the exact verified local package.
: "${APPROVED_DEB:?Set the approved absolute package path}"
case "$APPROVED_DEB" in
/*.deb) ;;
*) printf '%s\n' 'Expected an absolute .deb path' >&2; exit 1 ;;
esac
sudo apt-get --simulate install -- "$APPROVED_DEB"
I review the simulation and resolve every unexpected operation before authorising installation. In the same prepared shell, I then run:
sudo apt-get install -- "$APPROVED_DEB"
The APT simulation helps me review package operations. It does not execute the package's installation scripts or prove that firewall and application behaviour will be correct. I use the authenticated package and the actual lab execution for that evidence. APT's own documentation explains its simulation and installation behaviour. [12] Debian apt-get manual
After installation, I require the exact expected package version and the install ok installed state. I use dpkg --audit to detect incomplete package state, validate the installed configuration using the release-supported command and inspect relevant service failures. A green package status is necessary, but it is only one acceptance condition. [13] Debian dpkg manual
I run the business probes from outside the changed host. I test successful customer transactions, necessary backend flows and intentionally denied paths. Where VPN connectivity is part of the profile, I test useful traffic over the tunnel rather than relying on an interface being present. Where IPv6 is enabled, it has its own positive and negative checks.
I record CPU, memory, disk growth, service latency and alert volume against the baseline. I set thresholds from the application's objectives and the lab measurements. An example stopping condition is a material service-level breach, any unexpected administrative lockout, an authenticity failure, or a repeated unexplained error across hosts. I do not invent a universal acceptable latency penalty for SysWarden.
If a check fails, I stop the batch and preserve the evidence. I do not rerun an installer repeatedly until its exit code changes, erase ownership records or flush the firewall to make a dashboard look healthy. The runbook determines whether I can make a bounded repair, restore the recovery point or rebuild and rejoin the service.
20 / A FAILURE MUST HAVE AN ENDING
I distinguish recovery, rollback and removal.
When I say rollback, I mean a tested return to a compatible service state. That can require configuration, runtime state, package metadata and application data together. Installing an older binary over newer state is not automatically a supported rollback. I use the selected release's documented recovery boundary.
If a deployment causes an availability incident, I first stop expansion and preserve access through the approved independent route. I identify the affected service and containment boundary, retain the relevant evidence, and have the incident lead choose the recovery action. A suspected compromise may require isolation and forensic preservation before restoration; a simple configuration error may permit a smaller correction.
I recover one representative affected host first and verify useful service before applying that action to others. I test new connections, application integrity, monitoring and persistent policy after a reboot. I update the release or configuration record if the recovery changed the approved procedure. The repaired cohort does not silently inherit the previous acceptance result.
Removal is a separate lifecycle operation. Debian distinguishes removing a package from purging its package configuration. Neither is a general command to return the machine to an earlier operating-system snapshot. I follow the product's exact lifecycle contract and preserve unrelated administrator state. I do not use a package purge as a substitute for an application recovery plan.
Before calling removal complete, I verify that product-owned active components and persistent rules are gone, deliberately retained recovery material is private and inactive, and unrelated protection still works. I test the shared loader and services after reload and reboot. An unresolved ownership conflict is an unfinished operation with a next step, not a successful cleanup.
I also retain the ability to reinstall from an authenticated release after removal. That requires a consistent package database, an explicit configuration decision and fresh verification. I keep this lifecycle evidence with the deployment record because changing direction later is part of operating responsibly.
21 / DAY TWO AND EVERY DAY AFTER
I keep the evidence useful after the rollout.
At the end of each wave, I reconcile the expected inventory with observed results. My report shows accepted hosts, failed hosts, recovered hosts and exclusions, with reasons. Five hundred reachable machines are not necessarily five hundred accepted deployments. I require an accountable final state for each inventory entry.
I monitor configuration drift, package versions, enforcement health, feed freshness, service failures, resource use, false positives and recovery readiness. I track whether alerts reach the response team and whether that team can resolve them. I also maintain an exception register with owners and expiry dates so a temporary allowance does not quietly become permanent.
I separate product health from business health. A running process can coexist with a broken customer journey. A quiet event stream can coexist with a failed collector. Independent service probes and control-plane checks help me recognise those differences.
I review residual risk after a major service change, a new exposure, an important vulnerability, a failed recovery exercise, a supplier change or a material incident. I also schedule a periodic review even when nothing appears wrong. The twelve-month horizon is not a reason to wait twelve months before reacting.
I retain evidence proportionately. Private configuration exports, network inventories and operational logs belong in a controlled store with access rules, retention limits and a deletion process. Public engineering material can describe independently verified product behaviour without exposing the organisation's infrastructure or confidential feedback.
My final handover includes the approved release and configuration identities, service-owner decisions, residual risk register, support contacts, monitoring routes, recovery runbooks, maintenance schedule and the evidence for each accepted host. I ask the operations team to demonstrate the procedures it will own. That conversation is more valuable than a ceremonial handover of an unread document.
This is the standard I want to hold my own project to: a deployment that an operator can explain, a business owner can judge and a future maintainer can recover. The useful outcome is a stronger service with a decision trail that still makes sense when conditions change.
22 / PRIMARY SOURCES AND SCOPE
The references behind the method.
I checked these primary references on 7 October 2026. The ISO reference identifies the published standard; it is not a substitute for obtaining the full text for a formal assessment. ANSSI's English method guide is also available through its official English guide page; I consult the current French guidance and method sheets for updates. The fictional scenarios and deployment choices in this article are my own worked example.
- ISO/IEC 27005:2022: scope and edition of the information security risk management guidance.
- ANSSI: EBIOS Risk Manager: method, five workshops and scenario scope.
- NIST SP 800-30 Rev. 1: complementary risk assessment reference.
- ANSSI: structuring risk treatment measures: action planning and lifecycle follow-up.
- Debian 13 release notes: operational issues: distribution-specific changes to check.
- Debian 13 supported architectures: distribution architecture scope.
- GitHub artifact attestations: provenance purpose and limits.
- GitHub CLI: attestation verification: identity and provenance-policy controls.
- GitHub: offline attestation verification: bundles and trusted root material.
- Ansible: execution strategies: batches, ordering and concurrency.
- Ansible: error handling: stopping and rescue semantics.
- Debian: apt-get(8): package operations and simulation.
- Debian: dpkg(1): package state and audit behaviour.
For the product itself, I keep the current release and installation guide, architecture guide and release assurance record as the canonical starting points. I select the exact version-specific procedure before execution. This article defines a deployment approach; it does not announce or approve an unpublished release.