@@ -322,59 +322,8 @@ Matching tests for these requirements are listed in [6.3.6 Metrics tests].
### 5.3.7 Data minimisation
Data minimisation is achieved with understanding what data is collected [REQ-METRICS-2], and how long it is reasonably stored [REQ-METRICS-3].
### 5.3.8 High Availability
High availability starts from the running process.
In a modern cluster runtime environment used in large system deployments, the process rarely can control the loss of underlying resources.
Administrative actions can shutdown unexpectedly the node without a preseeding announcement.
It is up to the software design whether or not such interruptions can be tolerated.
Modern design is often distributed, but depending on the implementation and runtime context, a singular process can also provide the targeted service availability, if the process was implemented correctly, and a self-healing system can launch a replacement within a given time window.
Following general risks apply if the NMS can be reached by a DDoS attack: For an NMS in large deployment scenarios, e.g. telecom or enterprise use cases, a DDoS is not practical due to the physical separation of payload network and management network.
An attacker would need to physically access a high number of the managed devices, manipulate them, build a bridge to the management network, and be able to orchestrate then DDoS on the NMS.
As this is considered not practical in real scenarios, such an NMS itself will not be affected by DDoS.
Of course, such NMS are able to detect DDoS scenarios on its managed devices.
NMS in these large scenarios are usually also part of the DDoS defense means: The NMS gets input from defense systems and executes then mitigating configurations on the managed devices affected by the DDoS attack.
In these scenarios and for the CRA conformity, the NMS manufacturer can argue: a DDoS is not applicable.
For other scenarios, where the NMS itself can be affected by DDoS, the NMS may cease services for a configurable timeframe, apply predefined mitigation means, such as a reset, and return then to normal operation.
Most DDoS scenarios last less than 10 minutes.
However, such event needs to be notified, reported and logged.
Of course, for forensic analysis of this event, i.e. to identify the source of DDoS orchestration, sources of the flooding etc., a logging of all available data about this event is required.
Generally, DDoS attack vary greatly in duration, but following some online statistics, 60-80% last less than one hour, the most thereof less than 10 minutes.
Nevertheless, single publicly known campaigns can last 12 hours, the longest known 2025/12 lasted more than 7 days, but those are usually rare.
These are example data for which public statistics are available [i.17].
For low risk:
***[REQ-HA-0]** The expected availability shall be defined for each relevant product component.
***[REQ-HA-1]** System updates and changes shall not be considered as exceptions in the product general availability definition.
***[REQ-HA-6]** The product shall emit security events about detected issues that affects the high availability.
For medium risk:
***[REQ-HA-2]** The product shall tolerate loss of resources within the limits of the defined availability.
***[REQ-HA-6]** The product shall minimise the impact to other systems when anomalies occur.
***[REQ-HA-3]** Recovery capabilities shall be sufficiently implemented to match the expected availability targets.
<mark>REQ-HA-3 needs documentation trick for: "made available in the technical documentation"</mark>
For high risk:
DDoS mitigations:
* port switching
* traffic redirections
* service termination for a recommended and configurable time
* disabling of the affected ports and interfaces for a configurable time
***[REQ-HA-4]** The prodct shall implement coordinated brute‑force and overload protection mechanisms that not only detect excessive authentication attempts or inbound traffic surges but also enforce active mitigation actions including, but not limited to connection throttling, temporary IP blocking, message buffering, QoS parameterisation.
***[REQ-HA-5]** The product shall implement recovery or failover to mitigate overload attempts.
***[REQ-HA-8]** If the product can be exposed to DDoS, the product shall implement DDoS mitigations like listed above.
# 6 Conformity assessments and tests
## 6.1 General requirements assessments
@@ -816,134 +765,6 @@ DDoS mitigations:
### 6.3.8 High availability tests
#### 6.3.8.0 REQ-HA-0
**Objective:**
**Preparation:**
1. Have the product initialised and available with the default configuration and required credentials.
**Activities:**
1. Study the metrics for availability information provided in the technical documentation
2. Check in the system whether the features are available as described in the technical documentation
**Verdict:**
1. Pass, if all relevant features are available as defined and tracked.
1. Fail otherwise.
**Supporting Evidence:**
1. Relevant metrics described in the technical documentation.
1. Screen shots of metrics being visualised in the dashboards.
#### 6.3.8.1 REQ-HA-1
**Objective:** Events by changes of the product itself that impact the product availability do not render the product behaviour into the unpredicatable. The product keeps the availability time definitions.<br/>
**Preparation:**
1. Have the product initialised and available with the default configuration and required credentials.
**Activities:**
1. Intensionally terminate randomly an NMS-internal process, a processing node or simulate a loss of a datacenter.
2. Repeat the previous step 1 enough often with varying scopes to demonstrate conformance.
**Verdict:**
1. Pass, if the effect of the loss of a chosen resource or process termination matches the availability and service description, and that the NMS meets the availability time period definitions.
1. Fail otherwise.
**Supporting Evidence:**
1. Structured log output or other documentation that shows made actions and perceived operative response.
**Objective:** System updates and changes are included in the availability definition.<br/>
#### 6.3.8.2 REQ-HA-2
**Objective:** The user understands how the sytem behaves under different conditions and can make a disaster recovery plan for the operation.<br/>
**Preparation:**
1. Have the product initialised and available with the default configuration and required credentials.
**Activities:**
1. Study the availability definition from the documentation.
2. Disconnect or make unavailable a service, a compute node, a rack, or a datacenter based on the tolerable definitions provided in the technical documentation.
3. Wait for stability.
4. Reconnect or make available the previously removed resource.
5. Wait for stability.
**Verdict:**
1. Pass if recovery expectations are clearly defined,
2. and the system returns to stability after removal of the the maximum capasity mentioned in the toleration description,
3. and the deleted or removed resources returns operation after they are reintroduced to the system,
4. and the defined high availability is hold.
5. Fail otherwise.
> NOTE: The selection of the resource to be removed should meet the principal of the most needed or having the highest priority resource for the normal NMS operations.
**Supporting Evidence:**
1. Pointers to the documentation.
2. Description of the test procedure and relevant bits of log showing the phases.
#### 6.3.8.3 REQ-HA-3
**Objective:** The technical documentation provides explanations on how the sytem behaves under different conditions, enabling the user to develop a disaster recovery plan for the operation.<br/>
**Preparation:** None <br/>
**Activities:**
1. Cross-reference the recovery capabilities with the distribution description.
2. Cross-reference the recovery capabilities with the deployment plan and system architecture.
**Verdict:**
1. Pass if recovery expectations are clearly defined
1. and the distribution of the system operations is within reasonable scope in regards to the production computational resources.
1. Fail otherwise.
**Supporting Evidence:**
1. Pointers to the documentation.
#### 6.3.8.4 REQ-HA-4
**Preparation:**
1. Ensure brute force protection is enabled and configured as per manufacturers instructions <br/>
**Activities:**
1. From a test host, generate repeated failed SSH / HTTP (as applicable) logins to exceed the configured threshold.
2. Observe the scheduler‑based detection and automatic IP block of the attacking host
**Verdict:**
1. Pass if the offending IP is blocked for the configured duration and legitimate connections are still served.
1. Fail otherwise.
#### 6.3.8.5 REQ-HA-5
**Preparation:**
1. Identify queues/handlers that can buffer incoming work (Eg: task queues, batch operations)
**Activities:**
1. Feed they system with large volume of batch jobs or messages
1. Execute concurrent operations in accordance
1. Confirm the operations are batched and system is still serving requests and is not unresponsive
**Verdict:**
1. Pass if buffering and task queuing is performed within the queue depth or metrics as per the manufactured defined limits. The system load reduces as the tasks complete.
1. Fail otherwise.
# Annex A (informative): Mapping with essential requirements of the CRA
> Table mapping technical cybersecurity requirements from Section 5 of the present document to essential cybersecurity requirements in Annex I of the CRA. The purpose of this is to help identify missing technical cybersecurity requirements.
@@ -1288,10 +1288,56 @@ This clause addresses the requirements in the CRA [\[i.1\]](#_ref_i.1) Annex 1 P
## 5.10 Availability protection
<mark>_Proposed ESR code: AP_</mark>
This clause addresses the requirements in the CRA [\[i.1\]](#_ref_i.1) Annex 1 Part 1 (2) (h).
High availability starts from the running process.
In a modern cluster runtime environment used in large system deployments, the process rarely can control the loss of underlying resources.
Administrative actions can shutdown unexpectedly the node without a preseeding announcement.
It is up to the software design whether or not such interruptions can be tolerated.
Modern design is often distributed, but depending on the implementation and runtime context, a singular process can also provide the targeted service availability, if the process was implemented correctly, and a self-healing system can launch a replacement within a given time window.
Following general risks apply if the NMS can be reached by a DDoS attack: For an NMS in large deployment scenarios, e.g. telecom or enterprise use cases, a DDoS is not practical due to the physical separation of payload network and management network.
An attacker would need to physically access a high number of the managed devices, manipulate them, build a bridge to the management network, and be able to orchestrate then DDoS on the NMS.
As this is considered not practical in real scenarios, such an NMS itself will not be affected by DDoS.
Of course, such NMS are able to detect DDoS scenarios on its managed devices.
NMS in these large scenarios are usually also part of the DDoS defense means: The NMS gets input from defense systems and executes then mitigating configurations on the managed devices affected by the DDoS attack.
In these scenarios and for the CRA conformity, the NMS manufacturer can argue: a DDoS is not applicable.
For other scenarios, where the NMS itself can be affected by DDoS, the NMS may cease services for a configurable timeframe, apply predefined mitigation means, such as a reset, and return then to normal operation.
Most DDoS scenarios last less than 10 minutes.
However, such event needs to be notified, reported and logged.
Of course, for forensic analysis of this event, i.e. to identify the source of DDoS orchestration, sources of the flooding etc., a logging of all available data about this event is required.
Generally, DDoS attack vary greatly in duration, but following some online statistics, 60-80% last less than one hour, the most thereof less than 10 minutes.
Nevertheless, single publicly known campaigns can last 12 hours, the longest known 2025/12 lasted more than 7 days, but those are usually rare.
These are example data for which public statistics are available [i.17].
For **low** risk:
***AP_HA-0** The expected availability shall be defined for each relevant product component.
***AP_HA-1** System updates and changes shall not be considered as exceptions in the product general availability definition.
***AP_HA-6** The product shall emit security events about detected issues that affects the high availability.
For **medium** risk:
***AP_HA-2** The product shall tolerate loss of resources within the limits of the defined availability.
***AP_HA-6** The product shall minimise the impact to other systems when anomalies occur.
***AP_HA-3** Recovery capabilities shall be sufficiently implemented to match the expected availability targets.
<mark>REQ-HA-3 needs documentation trick for: "made available in the technical documentation"</mark>
For **high** risk:
DDoS mitigations:
* port switching
* traffic redirections
* service termination for a recommended and configurable time
* disabling of the affected ports and interfaces for a configurable time
***AP_HA-4** The prodct shall implement coordinated brute‑force and overload protection mechanisms that not only detect excessive authentication attempts or inbound traffic surges but also enforce active mitigation actions including, but not limited to connection throttling, temporary IP blocking, message buffering, QoS parameterisation.
***AP_HA-5** The product shall implement recovery or failover to mitigate overload attempts.
***AP_HA-8** If the product can be exposed to DDoS, the product shall implement DDoS mitigations like listed above.
## 5.11 Non-interference
<mark>_Proposed ESR code: IM_</mark>
@@ -1591,6 +1637,134 @@ Verify that:
## 6.10 Availability protection
### 6.10.0 REQ-HA-0
**Objective:**
**Preparation:**
1. Have the product initialised and available with the default configuration and required credentials.
**Activities:**
1. Study the metrics for availability information provided in the technical documentation
2. Check in the system whether the features are available as described in the technical documentation
**Verdict:**
1. Pass, if all relevant features are available as defined and tracked.
1. Fail otherwise.
**Supporting Evidence:**
1. Relevant metrics described in the technical documentation.
1. Screen shots of metrics being visualised in the dashboards.
### 6.10.1 REQ-HA-1
**Objective:** Events by changes of the product itself that impact the product availability do not render the product behaviour into the unpredicatable. The product keeps the availability time definitions.<br/>
**Preparation:**
1. Have the product initialised and available with the default configuration and required credentials.
**Activities:**
1. Intensionally terminate randomly an NMS-internal process, a processing node or simulate a loss of a datacenter.
2. Repeat the previous step 1 enough often with varying scopes to demonstrate conformance.
**Verdict:**
1. Pass, if the effect of the loss of a chosen resource or process termination matches the availability and service description, and that the NMS meets the availability time period definitions.
1. Fail otherwise.
**Supporting Evidence:**
1. Structured log output or other documentation that shows made actions and perceived operative response.
**Objective:** System updates and changes are included in the availability definition.<br/>
### 6.10.2 REQ-HA-2
**Objective:** The user understands how the sytem behaves under different conditions and can make a disaster recovery plan for the operation.<br/>
**Preparation:**
1. Have the product initialised and available with the default configuration and required credentials.
**Activities:**
1. Study the availability definition from the documentation.
2. Disconnect or make unavailable a service, a compute node, a rack, or a datacenter based on the tolerable definitions provided in the technical documentation.
3. Wait for stability.
4. Reconnect or make available the previously removed resource.
5. Wait for stability.
**Verdict:**
1. Pass if recovery expectations are clearly defined,
2. and the system returns to stability after removal of the the maximum capasity mentioned in the toleration description,
3. and the deleted or removed resources returns operation after they are reintroduced to the system,
4. and the defined high availability is hold.
5. Fail otherwise.
> NOTE: The selection of the resource to be removed should meet the principal of the most needed or having the highest priority resource for the normal NMS operations.
**Supporting Evidence:**
1. Pointers to the documentation.
2. Description of the test procedure and relevant bits of log showing the phases.
### 6.10.3 REQ-HA-3
**Objective:** The technical documentation provides explanations on how the sytem behaves under different conditions, enabling the user to develop a disaster recovery plan for the operation.<br/>
**Preparation:** None <br/>
**Activities:**
1. Cross-reference the recovery capabilities with the distribution description.
2. Cross-reference the recovery capabilities with the deployment plan and system architecture.
**Verdict:**
1. Pass if recovery expectations are clearly defined
1. and the distribution of the system operations is within reasonable scope in regards to the production computational resources.
1. Fail otherwise.
**Supporting Evidence:**
1. Pointers to the documentation.
#### 6.10.4 REQ-HA-4
**Preparation:**
1. Ensure brute force protection is enabled and configured as per manufacturers instructions <br/>
**Activities:**
1. From a test host, generate repeated failed SSH / HTTP (as applicable) logins to exceed the configured threshold.
2. Observe the scheduler‑based detection and automatic IP block of the attacking host
**Verdict:**
1. Pass if the offending IP is blocked for the configured duration and legitimate connections are still served.
1. Fail otherwise.
#### 6.10.5 REQ-HA-5
**Preparation:**
1. Identify queues/handlers that can buffer incoming work (Eg: task queues, batch operations)
**Activities:**
1. Feed they system with large volume of batch jobs or messages
1. Execute concurrent operations in accordance
1. Confirm the operations are batched and system is still serving requests and is not unresponsive
**Verdict:**
1. Pass if buffering and task queuing is performed within the queue depth or metrics as per the manufactured defined limits. The system load reduces as the tasks complete.