Digital twin driven multi-agent collaborative operation method for optical transport network

By constructing a real-time mapped digital twin model and a collaborative multi-agent operation and maintenance method, the problems of untimely fault analysis and mishandling in optical transmission networks have been solved, and safe and reliable fault handling and service recovery have been achieved.

CN122339946APending Publication Date: 2026-07-03CHENGDU XIONGBO TECH DEV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU XIONGBO TECH DEV
Filing Date
2026-06-03
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing fault handling methods for optical transmission networks rely on discrete alarms and human experience, making it difficult to integrate topology, physical link performance, and service carrying relationships in a timely manner. This leads to untimely fault analysis, potential erroneous automatic handling actions, and the possibility of mishandling or secondary impacts when the twin state is unreliable.

Method used

Construct a real-time mapped digital twin model, identify root cause candidates and handling candidates through fault propagation graphs and business carrying constraint matrices, and use twin state difference gates and handling control tokens to coordinate multiple agents for operation and maintenance, ensuring safe execution of actions and verification feedback.

Benefits of technology

It enables timely analysis and safe handling of optical transmission network faults, reduces the risk of misoperation and secondary impact, and ensures the verifiability and reliability of service recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122339946A_ABST
    Figure CN122339946A_ABST
Patent Text Reader

Abstract

This invention discloses a digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks, belonging to the field of optical transmission network operation and maintenance control technology. The method collects optical network topology, device status, link performance, alarm information, and service bearer relationships; constructs a real-time mapped digital twin model and fault propagation graph; filters executable actions based on the service bearer constraint matrix; releases the action control token through a twin state difference gate; and has the action executed by the action agent and the verification agent perform multi-source verification and versioned write-back. This scheme can reduce the risk of mishandling and concurrent conflicts, and improve the reliability of fault location, automatic handling, and closed-loop verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical transmission network operation and maintenance control technology, specifically to a digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks. Background Technology

[0002] Optical transmission networks typically include optical modules, line cards, tributary cards, optical amplifiers, ROADM equipment, OTN cross-connect equipment, protection switching units, fiber optic links, channels, and service carrying paths. The operational status of these components is distributed across network management systems, telemetry acquisition systems, alarm systems, performance monitoring systems, and service resource systems. When a fault or link degradation occurs, maintenance personnel usually need to manually compare topology, port status, alarm timing, optical power, bit error rate, protection switching status, and service paths across multiple systems to determine the root cause of the fault and the scope of affected services.

[0003] Existing fault handling methods for optical transmission networks suffer from the following main problems: First, fault analysis relies on discrete alarms and human experience, making it difficult to integrate topology, physical link performance, equipment status, and service carrying relationships in a timely manner. Second, potential link degradation, before triggering a hard alarm, typically manifests as gradual changes in state variables such as received optical power, optical signal-to-noise ratio (OSNR), bit error rate (BER), forward error correction (FEC) margin, and bit error count, which traditional alarm-driven methods struggle to identify in a timely manner. Third, automated handling actions often focus on command issuance or policy generation, but fail to fully consider service carrying paths, protection switching status, alternative link margins, and control conflicts on the same port, link, board, protection group, or cross-connection. Fourth, after handling actions are executed, judging successful recovery solely based on equipment command acknowledgments or temporary alarm clearing may not confirm whether link performance and service carrying have truly recovered.

[0004] Digital twin technology can map the physical network state to the virtual side. However, in commercial optical transmission networks, differences may arise between the physical network state and the twin model due to telemetry delays, asynchronous sampling, equipment interface refresh cycles, instantaneous changes during protection switching, or cross-connection adjustments. If actions are directly issued when the twin state is unreliable or there is a state version conflict, it may lead to misoperation, duplicate processing, or secondary service impacts. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks, comprising the following steps: Step 1: Collect the optical network topology, device status, link performance, alarm information and service carrying relationships of the physical optical transmission network, and construct a real-time mapped digital twin model that includes the twin object status version, sampling timestamp, source tag and fault object closed-loop status fields; Step 2: Based on the real-time mapped digital twin model, the link performance, the device status, the alarm timing, and the service carrying relationship, construct a fault propagation graph with propagation direction, node confidence, and service impact weight; Step 3: Based on the fault propagation graph, determine the root cause candidate objects, affected service paths, and candidate handling actions. Based on the service bearer constraint matrix that represents the service path, alternative link margin, required link margin, and protection group status constraining the candidate handling actions, eliminate candidate handling actions that do not meet the service bearer constraints to obtain executable handling actions. Step 4: Before issuing the executable action, determine the control object and exclusive scope of the executable action. The exclusive scope is the set of physical objects that have control conflicts with the control object. Based on the latest physical network state snapshot formed by the device state and the real-time mapped digital twin model, perform a twin state difference gate judgment on the state difference items of the control object and the exclusive scope. When no state difference item is obtained, or the physical objects corresponding to all the state difference items do not belong to the control object or the exclusive scope and do not change the judgment result of whether the control object and the exclusive scope meet the control security conditions, release the disposal control token. When at least one physical object corresponding to the state difference item belongs to the control object or the exclusive scope, or at least one state difference item changes the judgment result of whether the control object or the exclusive scope meets the control security conditions, do not release the disposal control token, prohibit the execution of the executable action through the device control interface, and update the latest physical network state snapshot by rereading the device state, link performance, protection switching status, cross-connection status, or service bearing relationship. Step 5: Schedule the disposal agent and enable the disposal agent to execute the executable disposal action through the device control interface after obtaining the disposal control token. Within the configured closed-loop verification window, the disposal agent that has not obtained the disposal control token is prohibited from executing conflict disposal actions on physical objects within the exclusive scope. Step Six: Schedule the verification agent and have it read the device receipt, configuration activation status, link performance recovery status, alarm convergence status, and service recovery verification results within the closed-loop verification window. The service recovery verification results are formed from the verification results configured as the basis for service recovery verification in the service OAM detection results and bit error test results, thus forming a verification feedback status. When the verification feedback status is "successful handling," update the closed-loop status field of the fault object to "successful handling," and write the updated twin object status version confirmed by successful handling back to the real-time mapped digital twin model. When the verification feedback status is "partially successful," "failed handling," or "state conflict," retain the conflict record and trigger secondary positioning, action rollback, or manual takeover.

[0006] Furthermore, the construction of a real-time mapped digital twin model including the twin object's state version, sampling timestamp, source tag, and fault object closed-loop state fields includes: The nodes, ports, links, channels, fiber segments, protection groups, cross-connections, and service paths in the physical optical transmission network are mapped to twin objects, and each twin object is configured with twin object status version, sampling timestamp, source tag, acknowledgment status, controllable attributes, and fault object closed-loop status fields.

[0007] Furthermore, the link performance includes received optical power, transmitted optical power, optical signal-to-noise ratio (OSNR), bit error rate (BER), forward error correction (FEC) margin, bit error count, and historical performance trend; the device status includes temperature, current, voltage, port status, board status, and optical module status associated with the link; the link performance and the device status are used to form an optical link degradation status quantity, and the link degradation confidence is calculated based on the optical link degradation status quantity.

[0008] Furthermore, the fault propagation graph includes topological adjacency edges, protection switching associated edges, shared resource edges, service bearer edges, alarm timing edges, and propagation direction edges; wherein, the propagation direction edges are determined based on the optical signal propagation direction, cross-connection relationships, and protection switching status, and are updated when the protection switching status or cross-connection status changes.

[0009] Furthermore, the step of determining the root cause candidate, affected service path, and candidate handling action based on the fault propagation graph includes: Based on the node confidence in the fault propagation graph, the link degradation confidence calculated based on the link performance and the device status, the alarm timing, and the service impact weight, the root cause device, root cause link, propagation object, and the affected service path are distinguished, and the candidate actions for handling are formed for the root cause device or the root cause link.

[0010] Furthermore, the service bearing constraint matrix is ​​generated from the affected service path, the service bearing relationship, the allowed service impact range, the alternative link margin, the required link margin, the protection group status, and the action risk level; and excludes candidate actions that would cause the affected service path to exceed the allowed service impact range, cause the alternative link margin to be less than the required link margin, or conflict with the protection group status.

[0011] Furthermore, the step of performing a twin state difference gate determination based on the latest physical network state snapshot formed by the device state and the real-time mapped digital twin model for the state difference items of the controlled object and the exclusive range includes: Based on the latest physical network state snapshot formed by the device state, the twin object state version, the sampling timestamp, the source tag, and the obtained device interface receipt, the state of the object to be controlled, the protection switching state, the cross-connection state, and the port state in the latest physical network state snapshot are compared with the corresponding states in the real-time mapped digital twin model to obtain the state difference items; when no state difference items are obtained, or when the physical objects corresponding to all the state difference items do not belong to the controlled object or the exclusive scope and do not change the judgment result of whether the controlled object and the exclusive scope meet the control security conditions, the judgment is passed.

[0012] Furthermore, the disposal control token includes the locked object, action type, exclusive scope, timeout condition, rollback condition, and twin object state version before rollback; the locked object is a port, link, board, protection group, or cross-connection; the control conflict includes physical objects within the exclusive scope belonging to the same port, link, board, protection group, or cross-connection as the controlled object, or there being a control dependency between them; within the closed-loop verification window, the disposal agent that has not obtained the disposal control token shall not perform conflict disposal actions on physical objects within the exclusive scope.

[0013] Furthermore, the closed-loop verification window is configured based on device type, service level, performance counter refresh cycle, alarm clearing delay, protection switching convergence time, and service detection cycle. When the device acknowledgment is valid, the configuration effective status is consistent with the executable action, the link performance recovery status and alarm convergence status meet the system configuration conditions, and all verification results configured as the basis for service recovery verification meet the system configuration service recovery conditions, the verification feedback status is "processing successful".

[0014] The beneficial effects of this invention are: This invention enables nodes, ports, links, channels, fiber segments, protection groups, cross-connections, and service paths in a physical optical transmission network to correspond to the object states in the twin model by configuring state versions, sampling timestamps, and source tags for twin objects. This provides a data foundation for subsequent pre-control difference determination and post-processing version write-back.

[0015] This invention constructs a fault propagation graph that includes topological adjacency edges, protection switching associated edges, shared resource edges, service carrying edges, alarm timing edges, and propagation direction edges. This integrates optical network topology, protection switching status, cross-connection relationships, channel or fiber sharing relationships, alarm timing, optical link degradation status quantities, and service carrying relationships into a single analysis structure. This facilitates the differentiation of root cause devices, root cause links, propagation objects, and affected services, reducing the risk of misjudging downstream alarms or affected services as root causes.

[0016] This invention incorporates the affected service path, service bearing relationship, allowable service impact range, alternative link margin, and protection group status into the candidate action screening process through a service bearing constraint matrix. This ensures that automatic handling actions are subject to the joint constraints of service bearing relationship and link physical margin, reducing the risk of secondary impact caused by blind switching, resetting, rerouting, or cross-connection adjustments.

[0017] This invention uses a twin state difference gate to compare the latest state snapshot formed by the device state in the physical optical transmission network with the real-time mapped digital twin model before issuing device control actions. It can keep the disposal control token from being released when the twin state is outdated, the protection switching state is inconsistent, the cross-connection state is conflicting, or the port state change affects the controlled object. It also prohibits the issuance of disposal actions that will change the device configuration, protection switching state, cross-connection state, port state, or service bearer state, and allows retesting operations that do not change the above states to be executed in read-only retesting mode, thereby reducing the risk of mis-disposal based on untrusted twin states.

[0018] This invention limits the locked object, action type, exclusive scope, timeout conditions, rollback conditions, and twin object state version before rollback by using control tokens. It also limits the locked object to ports, links, boards, protection groups, or cross-connections, so that multiple operation and maintenance agents cannot concurrently execute conflicting actions on physical objects with control conflicts within the same closed-loop verification window.

[0019] This invention forms a verification feedback state by considering at least one of the following: device acknowledgment, configuration activation status, link performance recovery status, alarm convergence status, service OAM detection results, and bit error rate test results, which are configured as the basis for service recovery verification. When the verification feedback state is "successful handling," the closed-loop status field of the fault object is updated to "successful handling," and the updated twin object status version confirmed by successful handling is written back. When the verification feedback state is "partially successful," "failed handling," or "state conflict," conflict records are retained, and subsequent processing is triggered. This invention can distinguish between successful command issuance, successful configuration activation, successful link performance recovery, and successful service bearer recovery, forming a verifiable closed-loop handling process. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks. Detailed Implementation

[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.

[0022] The features and performance of the present invention will be further described in detail below with reference to embodiments.

[0023] In this embodiment, the physical optical transmission network includes nodes, ports, optical modules, line cards, tributary cards, optical amplifiers, ROADM devices, OTN cross-connect devices, protection switching units, channels, fiber segments, links, protection groups, cross-connects, and service paths that carry optical signal transmission and service switching. The optical network topology includes the physical connections between the above-mentioned objects and the logical connections formed by cross-connects, protection switching, channel multiplexing, shared fiber segments, and service carrying paths.

[0024] A real-time mapped digital twin model is maintained in the network controller, network management platform, or operation and maintenance control system. This real-time mapped digital twin model maps nodes, ports, links, channels, fiber segments, protection groups, cross-connections, and service paths in the physical optical transmission network to twin objects. For each twin object, it maintains a twin object status version, sampling timestamp, source tag, confirmation status, controllable attributes, link health field, degradation confidence field, action execution status field, verification feedback status field, and fault object closed-loop status field. The real-time mapping does not require absolute zero-latency consistency between the physical state and the twin state at any given time. Instead, it enables the twin object state to be continuously updated with telemetry sampling, alarm updates, device acknowledgments, and verification feedback. Furthermore, it allows for the determination of whether the current twin state is available for control actions based on the twin object status version and the sampling timestamp.

[0025] The twin object status version is used to distinguish the status of the same twin object at different times. The twin object status version is updated after the twin object status is confirmed by effective telemetry sampling, device feedback, or verification feedback. The sampling timestamp records the collection time, reporting time, or confirmation time corresponding to when telemetry data, alarm information, device status, service bearing relationship, or device feedback enters the real-time mapped digital twin model. The source tag records the data source, which includes telemetry acquisition interface, alarm system, network management system, service resource database, device control interface feedback, service OAM detection interface, or bit error rate test interface.

[0026] The device interface feedback includes device status query feedback and device control action feedback. When the device interface feedback is used for twin state difference gate determination, it is a device status query feedback; when the device feedback is used for closed-loop verification, it is a device control action feedback. The confirmation status records whether the twin object status has been confirmed by telemetry sampling, device feedback, or verification feedback. The controllable attribute records whether the twin object is allowed to be locked as a disposal control token object. The confirmation status and the controllable attribute are used to enable the twin state difference gate to distinguish between object states used only for display or analysis and object states that can participate in device control actions.

[0027] The device status includes board status, port status, optical module status, protection group status, cross-connection status, device power status, interface connectivity status, temperature, current, and voltage. The link performance includes received optical power, transmitted optical power, optical signal-to-noise ratio (OSNR), bit error rate (BER), forward error correction (FEC) margin, bit error count, performance counters, and historical performance trends. Operation, management, and maintenance (OAM) probing is used to detect service connectivity, service path status, or service layer maintenance information; bit error rate testing is used to verify the bit error recovery capabilities of the link or service channel.

[0028] The optical link degradation status parameters are formed by received optical power, transmitted optical power, OSNR, BER, FEC margin, bit error count, temperature, current, voltage, and historical performance trends. These parameters describe the trend of the optical link changing from a normal state to a potential fault state, rather than simply determining whether a triggered hard alarm exists. A degradation confidence level is formed based on these parameters. This confidence level can be generated by comparing the current performance status with historical performance trends, similar link health baselines, adjacent port status, or shared resource status, and is used to update the link health field and degradation confidence level field in the real-time mapped digital twin model.

[0029] The fault propagation graph uses physical objects, service objects, and alarm objects as nodes, and is constructed as a directed graph or a directional association graph using topology adjacency edges, protection switching association edges, shared resource edges, service bearer edges, alarm timing edges, and propagation direction edges. The topology adjacency edges are formed based on the connection relationships between nodes, ports, links, channels, and fiber segments; the protection switching association edges are formed based on protection group status, working paths, protection paths, and switching status; the shared resource edges are formed based on channels, fiber segments, shared links, shared amplifiers, or other shared optical layer resources; the service bearer edges are formed based on the bearer relationships between service paths and nodes, ports, links, channels, fiber segments, protection groups, and cross-connections; the alarm timing edges are formed based on alarm occurrence time, alarm clearing time, alarm object, and alarm source; and the propagation direction edges are formed based on the optical signal propagation direction, cross-connection relationships, and protection switching status.

[0030] The service bearer relationship records the bearer relationships between service paths and nodes, ports, links, channels, fiber segments, protection groups, and cross-connections. The service bearer constraint matrix constitutes an action screening structure with at least the affected service path, service bearer relationship, allowable service impact range, alternative link margin, required link margin, protection group status, and candidate actions as constraint elements. It may further include service level weights, protection attribute coefficients, action risk levels, and action impact flags. The service bearer constraint matrix is ​​used to determine whether a candidate action will cause the affected service path to exceed the allowable service impact range, whether it will cause the alternative link margin to be less than the link margin required to execute the corresponding candidate action, or whether it conflicts with the protection group status.

[0031] Before the executable action is issued, the twin state difference gate is determined. The twin state difference gate, based on the latest state snapshot, the twin object state version, the sampling timestamp, the source tag, and the device interface receipt, compares the state of the object to be controlled, the protection switching state, the cross-connection state, and the port state in the latest state snapshot with the corresponding states in the real-time mapped digital twin model to obtain state difference items. When no state difference items are obtained, or when the physical objects corresponding to all state difference items do not belong to the controlled object or exclusive scope and do not change the judgment result of whether the controlled object and exclusive scope meet the control security conditions, the disposal control token is released. When the physical object corresponding to at least one state difference item belongs to the controlled object or exclusive scope, or when at least one state difference item changes the judgment result of whether the controlled object or exclusive scope meets the control security conditions, the disposal control token is kept unreleased, and disposal actions that would change the device configuration, protection switching state, cross-connection state, port state, or service bearer state are prohibited from being executed through the device control interface. When a state difference item causes the disposal control token to remain unreleased, the device state, link performance, protection switching state, cross-connection state, or service bearer relationship are reread to update the latest state snapshot.

[0032] The disposal control token is a control permission object granted to the disposal agent to execute a specified disposal action. It includes the locked object, action type, exclusive scope, timeout condition, rollback condition, and the twin object state version before rollback. The locked object can be a port, link, board, protection group, or cross-connection. The exclusive scope is the set of physical objects that have control conflicts with the locked object. Control conflicts include physical objects belonging to the same port, link, board, protection group, or cross-connection as the locked object, or physical objects having control dependencies on the locked object. The closed-loop verification window is the time range or event range used to collect and judge verification feedback after the disposal action is executed. The closed-loop verification window is configured based on device type, service level, performance counter refresh cycle, alarm clearing delay, protection switchover convergence time, and service detection cycle.

[0033] The verification feedback status is formed based on at least one result configured as the basis for service recovery verification from the device receipts, configuration activation status, link performance recovery status, alarm convergence status, service OAM detection results, and bit error rate test results collected within the closed-loop verification window. The verification feedback status includes successful handling, partial success, handling failure, and state conflict. The operation and maintenance intelligent agent is a software functional unit in the operation and maintenance control system that executes specific input, judgment, action, and output tasks. It may include a predictive intelligent agent, a localization intelligent agent, a handling intelligent agent, and a verification intelligent agent. The operation and maintenance intelligent agent can be implemented based on rules, models, or workflows, but the boundaries of the operation and maintenance intelligent agent's participation in this implementation method are defined by input data, output status, control permissions, and interface timing.

[0034] The fault closed-loop handling system of this invention can be deployed in the network management system, operation and maintenance controller, digital twin control platform, or server connected to the above systems in the optical transmission network. The system acquires the optical network topology, device status, link performance, alarm information, and service carrying relationships in the physical optical transmission network through telemetry acquisition interfaces. Telemetry acquisition interfaces may include SNMP interfaces, NETCONF / YANG interfaces, Telemetry interfaces, TL1 interfaces, gNMI / gRPC interfaces, vendor network management northbound interfaces, or other interfaces capable of providing network status data. The system issues handling actions through device control interfaces, which may include device southbound configuration interfaces, protection switching control interfaces, cross-connection configuration interfaces, optical power adjustment interfaces, port reset interfaces, service rerouting interfaces, or test loopback interfaces.

[0035] A fault closed-loop handling system includes at least a telemetry acquisition module, a twin mapping module, a fault propagation graph construction module, a candidate action generation module, a state difference gate module, an agent scheduling module, a handling control module, and a verification write-back module. The telemetry acquisition module acquires physical network data; the twin mapping module maintains the real-time mapped digital twin model; the fault propagation graph construction module integrates topology, protection switching, shared resources, service carrying capacity, alarm timing, optical signal propagation direction, and link performance status; the candidate action generation module determines root cause candidates, affected service paths, and executable handling actions; the state difference gate module determines state differences before issuing handling actions; the agent scheduling module allocates prediction, location, handling, and verification tasks; the handling control module executes constrained handling actions through the device control interface; and the verification write-back module generates verification feedback states and writes them back to the real-time mapped digital twin model.

[0036] During system operation, state changes in the physical optical transmission network are first entered into the real-time mapped digital twin model via the telemetry acquisition module. The fault propagation graph construction module constructs a fault propagation graph based on the real-time mapped digital twin model. The candidate action generation module determines the root cause candidate objects and affected service paths based on the fault propagation graph, obtains the corresponding candidate actions for the root cause candidate objects, as well as the allowable service impact range, alternative link margin, required link margin, and protection group status. Based on the service bearer constraint matrix generated from the affected service paths, service bearer relationships, allowable service impact range, alternative link margin, required link margin, and protection group status, it excludes candidate actions that would cause the affected service path to exceed the allowable service impact range, make the alternative link margin less than the link margin required to execute the corresponding candidate action, or conflict with the protection group status, thus obtaining executable action. Before issuing executable action, the state difference gate module determines whether the corresponding twin state can be used for equipment control. The agent scheduling module schedules the disposal agent and the verification agent; the disposal control module issues control actions after the disposal agent obtains the disposal control token, and prohibits the disposal agent that has not obtained the disposal control token from performing conflict disposal actions within the configured closed-loop verification window; the verification write-back module enables the verification agent to read multi-source verification feedback within the closed-loop verification window, and updates the twin object status version and the closed-loop status field of the faulty object or retains the conflict record based on the verification feedback status.

[0037] Example 1

[0038] like Figure 1 As shown, the digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks includes the following steps: Step 1: Collect the optical network topology, device status, link performance, alarm information and service carrying relationships of the physical optical transmission network, and construct a real-time mapped digital twin model that includes the twin object status version, sampling timestamp, source tag and fault object closed-loop status fields; Step 2: Based on the real-time mapped digital twin model, the link performance, the device status, the alarm timing, and the service carrying relationship, construct a fault propagation graph with propagation direction, node confidence, and service impact weight; Step 3: Based on the fault propagation graph, determine the root cause candidate objects, affected service paths, and candidate handling actions. Based on the service bearer constraint matrix that represents the service path, alternative link margin, required link margin, and protection group status constraining the candidate handling actions, eliminate candidate handling actions that do not meet the service bearer constraints to obtain executable handling actions. Step 4: Before issuing the executable action, determine the control object and exclusive scope of the executable action. The exclusive scope is the set of physical objects that have control conflicts with the control object. Based on the latest physical network state snapshot formed by the device state and the real-time mapped digital twin model, perform a twin state difference gate judgment on the state difference items of the control object and the exclusive scope. When no state difference item is obtained, or the physical objects corresponding to all the state difference items do not belong to the control object or the exclusive scope and do not change the judgment result of whether the control object and the exclusive scope meet the control security conditions, release the disposal control token. When at least one physical object corresponding to the state difference item belongs to the control object or the exclusive scope, or at least one state difference item changes the judgment result of whether the control object or the exclusive scope meets the control security conditions, do not release the disposal control token, prohibit the execution of the executable action through the device control interface, and update the latest physical network state snapshot by rereading the device state, link performance, protection switching status, cross-connection status, or service bearing relationship. Step 5: Schedule the disposal agent and enable the disposal agent to execute the executable disposal action through the device control interface after obtaining the disposal control token. Within the configured closed-loop verification window, the disposal agent that has not obtained the disposal control token is prohibited from executing conflict disposal actions on physical objects within the exclusive scope. Step Six: Schedule the verification agent and have it read the device receipt, configuration activation status, link performance recovery status, alarm convergence status, and service recovery verification results within the closed-loop verification window. The service recovery verification results are formed from the verification results configured as the basis for service recovery verification in the service OAM detection results and bit error test results, thus forming a verification feedback status. When the verification feedback status is "successful handling," update the closed-loop status field of the fault object to "successful handling," and write the updated twin object status version confirmed by successful handling back to the real-time mapped digital twin model. When the verification feedback status is "partially successful," "failed handling," or "state conflict," retain the conflict record and trigger secondary positioning, action rollback, or manual takeover.

[0039] Specifically, the system collects information from the physical optical transmission network, including optical network topology, device status, link performance, alarm information, and service carrying relationships. The optical network topology can include relationships between nodes and ports, ports and links, links and fiber segments, fiber segments and channels, protection groups, cross-connections, and service paths. Device status can include board status, port status, protection switching status, cross-connection status, optical module status, device temperature, current, voltage, and interface connectivity. Link performance can include received optical power, transmitted optical power, OSNR, BER, FEC margin, bit error count, and performance counters. Alarm information can include alarm type, alarm level, alarm occurrence time, alarm clearing time, alarm object, and alarm source. Service carrying relationships can include service paths, service levels, service protection attributes, and the nodes, ports, links, channels, fiber segments, protection groups, and cross-connections traversed by the service.

[0040] The twin mapping module maps nodes, ports, links, channels, fiber segments, protection groups, cross-connects, and service paths in the physical optical transmission network to twin objects. Each twin object includes at least an object identifier, object type, corresponding physical object identifier, twin object status version, sampling timestamp, source tag, acknowledgment status, and controllable attributes. The object type identifies whether the twin object belongs to a node, port, link, channel, fiber segment, protection group, cross-connect, or service path. The twin object status version records the number of twin object status updates or the sequence of status updates. The sampling timestamp records the acquisition, reporting, or acknowledgment time of the corresponding status data. The source tag identifies whether the status data comes from a telemetry acquisition interface, alarm system, network management system, service resource database, device receipt interface, or verification interface. The acknowledgment status distinguishes whether the status has been confirmed by device receipts, verification feedback, or other reliable sources. The controllable attributes identify whether the twin object is allowed to be locked as a control token. The real-time mapped digital twin model also includes a fault object closed-loop status field, which is used to record the closed-loop processing status of the fault object, such as not handled, being handled, handled successfully, partially successful, handled unsuccessfully, or in conflict with the status.

[0041] When the same twin object receives status data from different sources, the twin mapping module determines whether the status data can be used to update the twin object based on the sampling timestamp, source tag, and confirmation status. If the status data is consistent with the current twin object status, or if the source of the status data has the confirmation capability for the corresponding controlled object, the twin object status version is updated. If there are conflicts between the status data, and the conflict affects the controlled object or exclusive scope of subsequent actions, the corresponding twin object is marked as having a status conflict. The twin object status version is not updated directly, and the system waits for rereading of status data, device feedback, or manual confirmation results.

[0042] For twin objects corresponding to links, ports, channels, or fiber segments, the system generates optical link degradation status quantities based on link performance and the status of devices associated with the link. The link performance may include received optical power, transmitted optical power, OSNR, BER, FEC margin, bit error count, and historical performance trends; the device status associated with the link may include temperature, current, and voltage. This data can come from real-time telemetry of the devices, periodic sampling from performance counters, or performance files or alarm-related data reported by the network management system.

[0043] The system performs state-based processing on optical link degradation status variables. This state-based processing may include: comparing the current received and transmitted optical power with the link's historical stable state; comparing the current OSNR, BER, FEC margin, and bit error count with the link's historical performance trends; comparing temperature, current, and voltage with the normal operating conditions of similar devices; and associating the current link state with the states of adjacent ports, shared fiber segments, or links within the same protection group. Based on these comparison results, the system generates a link health field and a degradation confidence field. The link health field indicates whether the link is currently in a normal, observed, degraded, or faulty state. The degradation confidence field indicates the degree of confidence that the link is trending towards degradation.

[0044] In one implementation, the system first normalizes the performance parameters in the optical link degradation state variables. For the link... The There are several performance parameters, and the current sampled value is recorded as follows: Historical health baseline is The parameter degradation direction coefficient is When the parameter value increases, it indicates degradation. When a parameter value decreases, it indicates degradation. The first... The normalized degradation deviation of each performance parameter is:

[0045] in, For the first A performance parameter at time... For the link Normalized degradation deviation; This is the current sampled value; Historical health baseline; This is the degradation direction coefficient; Non-zero protection parameters configured for the system are used to prevent division by zero. This indicates that the calculation results are limited to between 0 and 1. For received optical power, transmitted optical power, OSNR, and FEC margin, a decrease in parameter value indicates degradation; for abnormal deviations in BER, bit error count, temperature, current, and voltage, the system normalizes them according to the degradation direction coefficient of the corresponding parameter.

[0046] For performance parameters such as temperature, current, and voltage, which may exhibit both high-side and low-side anomalies, the system can set high-side degradation direction coefficients and low-side degradation direction coefficients separately, and calculate the high-side normalized degradation deviation and low-side normalized degradation deviation separately. Alternatively, the system can use the absolute deviation of the current sampled value relative to the historical healthy baseline to form the normalized degradation deviation, according to the system configuration rules. The results of these processing steps are used as the corresponding performance parameters. Input degradation confidence level calculation.

[0047] The system further calculates the link The trend of deterioration:

[0048] in, For the first The trend of deterioration of each performance parameter; The sampling interval configured for the system; This represents the parameter value at the previous sampling time. The trend degradation is used to characterize the change in the degradation direction of the same link between adjacent sampling times.

[0049] For links A set of neighboring links sharing fiber segments, channels, amplifiers, protection groups, or adjacent ports. ,when At that time, the system calculates the neighborhood association degradation amount according to the following formula:

[0050] in, For link The amount of neighborhood association degradation; For link A set of neighboring links that share resources or have propagation associations; The number of links in the neighborhood link set; For any link in the neighborhood link set; Neighborhood Links The degradation confidence level at the previous sampling time or the previous update cycle. When At that time, the system will Take 0.

[0051] link The degradation confidence level is calculated using the following formula:

[0052] in, For link At any moment The confidence level of degradation; The number of performance parameters involved in the calculation; For the first System configuration weights for deviations of performance parameters; For the first The system configuration weights of each performance parameter trend quantity; The system assigns weights to the neighborhood association degradation. The system is based on... Update the link health field and degradation confidence field in the real-time mapped digital twin model, and... As a link in the fault propagation diagram The confidence level of the corresponding node is input.

[0053] When initially calculating the link degradation confidence, if the link Historical health baseline Since this is not yet established, the system can use the device's factory configuration values, the average health sample value confirmed by the operation and maintenance system, or the system configuration baseline of similar links as the initial health baseline. During subsequent operation, the system can adjust the baseline based on historical sample values ​​identified as being in a normal state. Perform a sliding update. If the neighborhood association degradation was not present when initially calculating the neighborhood association degradation... If so, the previous cycle degradation confidence of the corresponding neighboring link will be set to 0 or the initial confidence of the system configuration.

[0054] The formation of degradation confidence does not depend on a single alarm. If states such as decreased optical power margin, decreased OSNR, increased BER, decreased FEC margin, or increased bit error count show a consistent trend on the same link or related links, the degradation confidence of the corresponding link increases. If, after the action is taken, optical power, OSNR, BER, FEC margin, or bit error count recovers to the healthy state range configured by the system, and the corresponding alarm converges or service probe passes, the degradation confidence of the corresponding link decreases. The specific healthy state range and trend determination granularity can be configured according to network standard, equipment type, and operation and maintenance strategy; this invention does not require fixed numerical thresholds as a necessary condition.

[0055] The optical link degradation state quantity also serves as input to the node confidence in the fault propagation graph. For link nodes, port nodes, or channel nodes with degraded link performance, the fault propagation graph construction module can increase the node confidence as a root cause candidate based on the degradation confidence. For other nodes that have optical signal propagation direction, shared resources, or service carrying relationships with the degraded link, the fault propagation graph construction module can update the node confidence and propagation direction edges based on the corresponding propagation relationships.

[0056] The fault propagation graph construction module constructs a fault propagation graph based on the port, link, channel, fiber segment, cross-connect, and protection switching status in the real-time mapped digital twin model. This is combined with alarm timing, optical link degradation state quantities based on link performance and the status of devices associated with the links, and service carrying relationships. Nodes in the fault propagation graph can include device nodes, port nodes, link nodes, channel nodes, fiber segment nodes, protection group nodes, cross-connect nodes, service path nodes, and alarm nodes. Edges in the fault propagation graph include topology adjacency edges, protection switching association edges, shared resource edges, service carrying edges, alarm timing edges, and propagation direction edges.

[0057] Topology adjacency edges are formed based on the connection relationships of nodes, ports, links, channels, and fiber segments in the optical network topology. Protection switching association edges are formed based on protection group status, working path, protection path, and switching status. Shared resource edges are formed based on resource sharing relationships such as channels, fiber segments, shared links, or shared amplifiers. Service carrying edges are formed based on the carrying relationships between service paths and nodes, ports, links, channels, fiber segments, protection groups, and cross-connections. Alarm timing edges are formed based on the timing relationships between alarm occurrence time, alarm clearing time, and alarm objects. Propagation direction edges are formed based on the optical signal propagation direction, cross-connection relationships, and protection switching status.

[0058] Each node in the fault propagation graph can carry a node confidence level and a service impact weight. The node confidence level is determined based on the alarm status, link degradation confidence level, alarm timing location, topology location, protection switching status, and correlation with other abnormal nodes of the object corresponding to the node. The service impact weight is determined based on the number of services carried by the object corresponding to the node, the service level, the service protection attributes, and the number of affected service paths.

[0059] In one implementation, for nodes in the fault propagation graph The system calculates node confidence based on evidence of physical degradation, alarm timing, propagation consistency, and business impact weights. The node confidence level is:

[0060] in, For nodes At any moment The node confidence level; For nodes Evidence of physical degradation of the corresponding object; For nodes Timing evidence corresponding to the alarm; For nodes Evidence of consistent propagation between upstream and downstream nodes; For nodes The weight of business influence; , , and Configure weights for the system.

[0061] In one implementation, According to nodes The maximum, average, or system configuration weighted value of the confidence score for link degradation is used to form the link degradation confidence score; when a node When not associated with link performance status, You can take a node Abnormal status flags for associated ports, boards, or protection groups. Based on nodes The occurrence order, duration, and alarm level of the corresponding alarms in the current alarm sequence form normalized temporal evidence. Based on nodes The abnormal state is determined by the percentage of edges whose directions are consistent with those of topological adjacent edges, protection switching associated edges, shared resource edges, service carrying edges, and propagation direction edges; when the abnormal propagation direction is consistent with the propagation direction edge. Increase when the abnormal propagation direction is inconsistent with the propagation direction edge. reduce.

[0062] In one implementation, physical degradation evidence It can be formed according to the following formula:

[0063] in, For nodes A set of associated links; For nodes An abnormal status flag for the corresponding device, port, board, or protection group.

[0064] In one implementation, alarm timing evidence It can be formed according to the following formula:

[0065] in, This is evidence at the alarm level. Evidence of the order in which alarms occurred. Evidence of alarm duration , and Configure weights for the system.

[0066] In one implementation, consistent evidence is propagated. It can be formed according to the following formula:

[0067] in, For nodes The number of edges propagating in the same direction as the abnormal state. For nodes The number of associated edges that can be used to propagate consistency judgments. When When dividing by zero, the denominator is set to 1. This indicates an abnormal status flag for the device, port, board, or protection group. This can be determined based on the state of the corresponding object; when the object is in an abnormal, invalid, offline, failover, or interface receipt abnormal state, When the object is in a normal state, When an object is in a state of pending confirmation or conflict, Take the intermediate value of the system configuration. Alarm level evidence. Alarm levels can be mapped to normalized values; evidence of alarm occurrence sequence. Alarms can be ranked according to their order in the current alarm sequence, with earlier alarms located upstream in the propagation direction having higher values; alarm duration evidence. This can be calculated based on the ratio of the alarm duration to the system configuration observation window. Consistent propagation direction means that the direction from the abnormal object to the affected object is the same as the direction of the propagation direction edge.

[0068] node The business impact weight is calculated according to the following formula:

[0069] in, For the nodes Or affected by nodes The affected business set; For business Business level weighting; For business The protection attribute coefficient; For nodes For business The indicator or degree of influence. When a node Located in business When on the working path, protection path, or fault propagation impact path Take 1 or the impact value of the system configuration; when the node Not part of business When the working path and protection path are not on the fault propagation impact path. .when At that time, node The system will not carry out tasks that do not affect business processes. Set to 0. For unprotected or high-priority services, the system can be configured with a higher value. or In order to increase its weight in the business impact.

[0070] In the fault propagation diagram, if a link node exhibits an abnormal optical link degradation status, and its upstream port, cross-connection, or protection group status simultaneously displays abnormal alarms, then that link node, upstream port node, or associated protection group node can be elevated to a root cause candidate. If a downstream service path node only exhibits service interruption or alarm propagation, but the corresponding link physical status does not show a root cause anomaly, then that service path node can be identified as an affected service, but not directly as a root cause object.

[0071] When the protection switching status changes, the fault propagation graph construction module updates the protection switching associated edges and propagation direction edges according to the new working path and protection path. When the cross-connection status changes, the fault propagation graph construction module updates the topology adjacency edges, service carrying edges, and propagation direction edges according to the new ingress port, egress port, and service path. Therefore, the fault propagation graph can be updated according to changes in protection switching and cross-connection status, avoiding the use of historical propagation directions for current actions.

[0072] The candidate action generation module determines root cause candidates and affected service paths based on the fault propagation graph. It then obtains candidate actions for each root cause candidate, along with the allowable service impact range, alternative link margin, required link margin, and protection group status corresponding to those actions. Based on a service bearer constraint matrix generated from the affected service paths, service bearer relationships, allowable service impact range, alternative link margin, required link margin, and protection group status, it excludes candidate actions that would cause the affected service path to exceed the allowable service impact range, make the alternative link margin less than the required link margin for executing the corresponding candidate action, or conflict with the protection group status. This results in executable action choices. Root cause candidates may include root cause devices, root cause links, root cause ports, root cause boards, root cause protection groups, or root cause cross-connections. Candidate action choices may include power adjustment, protection switching, service rerouting, port reset, cross-connection adjustment, and test loopback. Whether these actions are included in the executable action set is determined by device control permissions, device status, the service bearer constraint matrix, and the current network policy.

[0073] The system distinguishes between root cause devices, root cause links, propagation objects, and affected services based on node confidence in the fault propagation graph, link degradation confidence calculated from optical link degradation state variables, alarm timing, and service impact weights. Root cause devices or root cause links are used to locate the target objects for handling actions. Propagation objects describe intermediate objects in which the fault impact spreads along topology, protection switching, shared resources, or service bearer relationships. Affected services describe the service paths where service quality degradation, interruption, or protection switching occurs due to anomalies in the root cause object or propagation object. The system writes the correspondence between root cause devices, root cause links, propagation objects, and affected services into the service bearer constraint matrix.

[0074] The service bearer constraint matrix is ​​composed of affected service paths, service bearer relationships, permissible service impact range, alternative link margin, required link margin, protection group status, and candidate actions. The service bearer constraint matrix can be implemented using tables, matrices, graph edge attributes, or rule sets. For each candidate action, the service bearer constraint matrix records the service paths that the action may affect, the corresponding service bearer relationships, alternative link status, alternative link margin, whether it may trigger a secondary switchover, whether it may exceed the permissible service impact range, and whether it conflicts with the protection group status.

[0075] In one implementation, the service bearer constraint matrix is ​​used to handle candidate actions. and affected businesses For indexing, record business level weight, protection attribute coefficient, and action impact flag. Alternative link margin Required link margin Action risk level Conflict signs of protection group and the scope of permitted business impact .in, Indicates the handling of candidate actions It will affect business , This indicates that it will not affect business. .

[0076] Action Influence Marker Can be handled according to candidate actions The set of control objects and business The question determines whether the sets of objects they carry intersect; if they do intersect, or if the candidate actions will change the business logic. When specifying working paths, protection paths, or cross-connect relationships... ,otherwise .

[0077] Protection Group Conflict Markings It can be determined according to the following formula:

[0078] Permissible scope of business impact Can be handled by candidate actions The maintenance strategy for the service level, protection attributes, and system configuration within the affected service set is determined. For high-level services or unprotected services, a lower permissible service impact range is configured; for low-level services or services with available protection paths, a higher permissible service impact range can be configured. Service impact score. and the scope of permitted business impact Using the same normalized scoring scale, Configured to an upper limit of permissible impact between 0 and 1. Action risk level. Risk values ​​are mapped to normalized values ​​based on action type; for example, test loopback, power regulation, service rerouting, protection switching, cross-connect adjustment, and port reset can be configured with different risk values ​​based on their degree of change to service paths and equipment states. System Configuration At that time, and They are all on the same normalized scoring scale. The above action risk levels are only used for action screening and are not used as fixed experimental parameters.

[0079] Handling candidate actions The normalized business impact score is calculated according to the following formula:

[0080] in, To handle candidate actions The normalized business impact score; To handle candidate actions The range of businesses that may be affected; Risk weights configured for the system.

[0081] link Available link margin The margin can be formed by weighting optical power margin, OSNR margin, BER margin, FEC margin, and bit error count margin according to system configuration weights. The optical power margin characterizes the remaining margin of the current received or transmitted optical power relative to the system configuration health range; the OSNR margin characterizes the remaining margin of the current OSNR relative to the OSNR required by the service bearer; the BER margin characterizes the remaining margin of the current BER relative to the system configuration BER condition; the FEC margin characterizes the remaining margin of the current FEC margin relative to the FEC margin required by the service bearer; and the bit error count margin characterizes the remaining margin of the current bit error count relative to the system configuration bit error condition. In one implementation, each margin component is formed according to a normalized remaining margin of the current link state relative to the system configuration health condition, and is limited to between 0 and 1. For example, optical power margin is formed by the remaining amount of the current optical power relative to the system's configured healthy optical power range; OSNR margin is formed by the remaining amount of the current OSNR relative to the OSNR required by the service bearer; BER margin is formed by the remaining amount of the current BER relative to the system's configured BER condition; FEC margin margin is formed by the remaining amount of the current FEC margin relative to the FEC margin required by the service bearer; and bit error count margin is formed by the remaining amount of the current bit error count relative to the system's configured bit error condition. For link performance parameters whose values ​​decrease, indicating degradation, the system forms the corresponding margin component based on the remaining amount of the current value relative to the system's configured minimum healthy condition; for link performance parameters whose values ​​increase, indicating degradation, the system forms the corresponding margin component based on the remaining amount of the current value relative to the system's configured maximum healthy condition.

[0082] link Available link margin It is formed according to the following formula:

[0083] in, This refers to the optical power margin. OSNR margin; For BER margin; This represents the FEC margin. This is the error counting margin; , , , and Configure weights for the system. Execute candidate actions for handling. Post-service in the link Required link margin It is formed according to the following formula:

[0084] in, To execute the candidate action Post-service in the link The required link margin; To execute the candidate action It is expected to be carried on the link The business set on; For business The system requires sufficient basic link margin. For high-priority or unprotected services, the system can be configured with higher margins. , or At the business set level, when the link... This is a candidate action for handling. When the influence path is... for A subset or by It is obtained by remapping the bearer relationship after rerouting, protection switching, or cross-connection adjustments. When At that time, the link In executing the candidate disposal action After that, the system will not carry the corresponding business, and the system will Take 0.

[0085] when At that time, the availability margin of alternative links is poor. Determined according to the following formula:

[0086] in, To handle candidate actions The set of links included in the proposed alternative paths. When a candidate action corresponds to multiple alternative paths, the system can construct a corresponding path for each alternative path. and calculate respectively Alternatively, each alternative path can be treated as an independent candidate action. For candidate actions that do not involve alternative paths, the system does not calculate... , or Set to system configuration default and do not As a basis for excluding actions.

[0087] The candidate action generation module excludes and processes candidate actions when any of the following conditions are met. :First, First, the business impact score of the candidate action exceeds the allowed business impact range; second, That is, the minimum availability margin difference among the candidate links involved in the proposed action is negative, indicating that at least one candidate link has an availability margin less than the required link margin; third, The following are possible scenarios: First, the candidate action for handling conflicts with the current protection group status; second, the target object of the candidate action for handling lacks controllable attributes, or the target object has been locked by the exclusive scope of the handling control token. Candidate actions for handling that are not excluded are added to the set of executable actions.

[0088] For protection switching actions, the candidate action generation module checks the protection path status, alternative link margin, and the level of affected services. If the protection path does not exist, is in a degraded state, or the high-level services carried by the protection path would be affected beyond the allowable service impact range, the protection switching action is excluded. For service rerouting actions, the candidate action generation module checks the link performance, service level, and service protection attributes traversed by the alternative path. If the performance of the alternative path is insufficient to carry the affected services, the service rerouting action is excluded. For port reset actions, the candidate action generation module checks the service path carried by the port and whether an available protection path exists. If a port reset would cause high-level services to be unprotected and interrupted, the port reset action is prohibited. For test loopback actions, the candidate action generation module checks whether the test object is isolated from online services, or whether the test action is only used to reread status data and does not change the status of online services; if the test action would affect online services, the test loopback action is excluded.

[0089] Before an executable action is issued, the status difference gate module performs a twin status difference gate determination based on the latest status snapshot formed by the device status in the physical optical transmission network, the twin object status version, the sampling timestamp, the source tag, and the obtained device interface receipt. The device status may include the port status, board status, protection switching status, cross-connection status, link performance status, control interface connectivity status, and related device receipt status of the object to be controlled.

[0090] The scope of the twin state difference gate is determined based on the controlled object and exclusive scope of the executable action. If the executable action is protection switching, the scope includes at least the corresponding protection group, working path, protection path, related ports, and affected service paths. If the executable action is cross-connect adjustment, the scope includes at least the corresponding cross-connect, input port, output port, carried service path, and protection attributes. If the executable action is port reset, the scope includes at least the corresponding port, board, adjacent links, carried service path, and whether a protection path exists. If the executable action is power adjustment, the scope includes at least the corresponding optical module, link, channel, adjacent channels, amplifier status, and affected service paths.

[0091] The state difference gate module determines the controlled object and exclusive scope based on executable actions. The exclusive scope is a set of physical objects that have control conflicts with the controlled object. These control conflicts include physical objects within the exclusive scope belonging to the same port, link, board, protection group, or cross-connection as the controlled object, or physical objects having control dependencies on the controlled object. The state difference gate module compares the state of the controlled object, protection switching state, cross-connection state, and port state in the latest state snapshot with the corresponding states in the real-time mapped digital twin model to obtain state difference items. When no state difference item is obtained, the state difference gate module determines that the executable action meets the control safety conditions. When state difference items are obtained and all physical objects corresponding to all state difference items do not belong to the controlled object or exclusive scope, and this does not change the judgment result of whether the controlled object and exclusive scope meet the control safety conditions, the state difference gate module also determines that the executable action meets the control safety conditions.

[0092] When at least one state difference item corresponds to a physical object that belongs to a controlled object or exclusive scope, or when at least one state difference item changes the judgment result of whether the controlled object or exclusive scope meets the control security conditions, the state difference gate module keeps the disposal control token from being released and prohibits the execution of disposal actions that would change the device configuration, protection switching status, cross-connection status, port status, or service bearer status through the device control interface. When a state difference item causes the disposal control token to remain unreleased, the state difference gate module may trigger a rereading of the device status, link performance, protection switching status, cross-connection status, or service bearer relationship to update the latest status snapshot. State difference items may include changes in the status of the port to be controlled, changes in the status of the protection group, changes in the cross-connection status, re-bearing of the service path, inconsistencies between the object status represented by the device interface receipt and the twin object status, expired twin object status version of the object to be controlled, and sampling timestamps later than the time when the executable disposal action was generated but not included in the analysis, etc.

[0093] In one implementation, for executable disposal actions The system first determines the set of controlled objects. and exclusive range set Collection of controlled objects This includes the ports, links, boards, protection groups, or cross-connections directly affected by the aforementioned action; the exclusive scope set. This includes physical objects that have control conflicts with any object in the set of controlled objects.

[0094] for Any object in The system compares the object states in the latest state snapshot formed by the device states in the physical optical transport network. Real-time mapping of object states in digital twin models The system generates a difference impact flag based on object state differences, twin object state versions, sampling timestamps, source tags, and acquired device interface receipts. :

[0095] in, For object Compared to executable actions The differences in the influence of the markers; For object At any moment The object state in the latest state snapshot formed by the device state in the physical optical transport network; For object At any moment The real-time mapping of object states in the digital twin model. The object states represented by the acquired device interface receipts can include port states, protection group states, cross-connection states, interface connectivity states, or control interface confirmation states returned in the device status query receipts. When the object... If the sampling timestamp corresponding to the state version of the twin object is earlier than the time when the executable action is generated and it has not been included in the analysis process of the executable action, or if the sampling timestamp exceeds the state validity window configured by the system, the system determines that the object... The twin object's state version is outdated.

[0096] For port objects, when the port's working state, reset state, or service-bearing state in the latest state snapshot formed by the device state in the physical optical transmission network is inconsistent with the corresponding state in the real-time mapped digital twin model, and the port belongs to the set of control objects or the exclusive scope set for which actionable actions can be performed, For protected group objects, when the working path, protection path, or switching state in the latest state snapshot formed by the device status in the physical optical transmission network is inconsistent with the corresponding state in the real-time mapped digital twin model, and such inconsistency would change the action path of the executable handling action, For cross-connected objects, when the ingress port, egress port, or cross-connection status in the latest state snapshot formed by the device status in the physical optical transport network is inconsistent with the corresponding status in the real-time mapped digital twin model, and the inconsistency affects the target connection for which a disposal action can be performed, For business path objects, when a business path is re-beared after an executable action is generated, and the re-bearing changes the exclusive scope set, .

[0097] The gate value of the twin state difference gate is calculated according to the following formula:

[0098] in, For actionable procedures At any moment The twin state difference gate value. Located in The state difference item, which does not change the judgment result of whether the controlled object and exclusive scope meet the control safety conditions, is not included. .when When there are no state difference items affecting the controlled object or exclusive scope, the state difference gate module releases the disposal control token; when When a status difference gate module keeps the disposal control token from being released, it prohibits the execution of disposal actions that would change the device configuration, protection switching status, cross-connection status, port status, or service bearer status through the device control interface. When a status difference item causes the disposal control token to remain unreleased, the device status, link performance, protection switching status, cross-connection status, or service bearer relationship are reread to update the latest status snapshot.

[0099] When the twin state difference gate passes the determination, the state difference gate module releases the disposal control token. The disposal control token includes the locked object, action type, exclusive scope, timeout condition, rollback condition, and the twin object state version before rollback. The locked object can be a port, link, board, protection group, or cross-connection. The exclusive scope is limited to physical objects that have control conflicts with the locked object, such as the same port, the same link, the same board, the same protection group, the same cross-connection, or adjacent objects that have control dependencies on it.

[0100] Within the configured closed-loop verification window, the system prohibits agents without a disposal control token from performing conflict resolution actions on physical objects within their exclusive scope. Conflict resolution actions may include simultaneously performing mutually exclusive protection switching on the same protection group, simultaneously performing reset and configuration adjustment on the same port, simultaneously performing different cross-connection configurations on the same cross-connection, and simultaneously performing control actions that change the link state on the same link. If an agent without a disposal control token requests to perform an action on a physical object within its exclusive scope, and the request would change the state of the physical object within that scope, the agent scheduling module will suspend or reject the request; if the request is only used to reread device status, link performance, protection switching status, cross-connection status, or service bearer relationship, the agent scheduling module will allow the execution of the request to update the latest state snapshot.

[0101] After the agent scheduling module schedules the action agent, the action agent, upon obtaining the action control token, executes an executable action on the root cause candidate object through the device control interface. Actions can include power adjustment, protection switching, service rerouting, port reset, cross-connection adjustment, or loopback testing. After the action control module issues the action, it records the action number, target object, action type, action issuance time, device acknowledgment, twin object state version before rollback, and associated service path.

[0102] In one implementation, the exclusive scope set Determined according to the following formula:

[0103] in, To control any object in the collection of objects; The type of action for which a disposal action can be performed; To control conflict functions. When the object With the controlled object Belonging to the same port, the same link, the same board, the same protection group, the same cross connection, or there are action types between the objects mentioned above. When changing control dependencies at the same time Otherwise, it is 0.

[0104] The agent scheduling module receives an action request from a disposal agent that has not obtained a disposal control token within the closed-loop verification window. At that time, calculate the set of control objects for the action request. .like And action request It will change the exclusive scope set The state of the internal physical object will then trigger the action request. Refuse or suspend; if the action request Read only the exclusive range set If the state of the internal physical object is not changed, the action request can be executed. The results will then be used to update the latest state snapshot.

[0105] If the device control interface returns that the action is unexecutable, the device is unresponsive, the action permissions are insufficient, or the target object's state changes, rendering the action inapplicable, the handling control module will not continue executing subsequent control actions and will write the handling action execution status into the real-time mapped digital twin model. If the handling action has been executed but fails to be verified within the closed-loop verification window, the system will trigger action rollback, secondary positioning, or manual takeover based on the timeout and rollback conditions in the handling control token. Action rollback will refer to the configuration and state corresponding to the twin object's state version before rollback, and will again determine whether the rollback action meets the control safety conditions through the twin state difference gate before rollback.

[0106] When the verification feedback status indicates successful handling, the agent scheduling module releases the handling control token. When the closed-loop verification window times out or the verification feedback status indicates handling failure, the agent scheduling module triggers action rollback or secondary positioning based on rollback conditions and the twin object state version before rollback. Therefore, the handling control token is used not only for control permission before action issuance but also for prohibiting conflict handling within the closed-loop verification window and for rollback control after handling failure.

[0107] After the handling action is executed, the agent scheduling module schedules the verification agent. Within the closed-loop verification window, the verification agent reads the verification results configured as the basis for service recovery verification from the device acknowledgments, configuration activation status, link performance recovery status, alarm convergence status, and service OAM detection results and bit error rate test results through the corresponding verification interface. The device acknowledgments confirm whether the device has received and executed the control command. The configuration activation status confirms whether protection switching, cross-connection adjustment, power regulation, port reset, service rerouting, or test loopback has taken effect on the device side. The link performance recovery status confirms whether received optical power, transmitted optical power, OSNR, BER, FEC margin, bit error rate count, etc., meet the system-configured link recovery conditions. The alarm convergence status confirms whether root cause alarms, propagation alarms, and affected service alarms meet the system-configured alarm convergence conditions. The service recovery verification results confirm whether the affected service path meets the system-configured service recovery conditions; these results are formed from the verification results configured as the basis for service recovery verification from the service OAM detection results and bit error rate test results.

[0108] Link recovery conditions, alarm convergence conditions, and service recovery conditions can be configured based on device type, service level, network standard, performance counter refresh cycle, alarm clearing delay, protection switching convergence time, and service detection cycle. This invention does not limit the above conditions to fixed numerical thresholds. Link recovery conditions can manifest as link performance recovering to the system's configured healthy state; alarm convergence conditions can manifest as root cause alarm clearing, propagation alarm stopping propagation, or alarm status degradation; and service recovery conditions can manifest as successful service OAM detection, successful bit error rate test, or service path recovery and availability.

[0109] In one implementation, the closed-loop verification window is based on executable actions. The action type, device type, and service level are determined. The duration of the closed-loop verification window is also determined. Configure as follows:

[0110] in, For actionable procedures The duration of the closed-loop verification window; Configure the time required for the action to take effect; For equipment The performance counter refresh cycle; For equipment Alarm clearing delay; To protect the time required for the convergence of switching or cross-connection; The business probing period for the affected business set is determined according to the following formula:

[0111] in, For business The business OAM detection or error test cycle; The minimum verification period configured for the system, when the system configuration does not require business probing. A value of 0 is acceptable. The above times are configured by the system based on device type, service level, and network standard. For verification items triggered by events, the system can enter the verification feedback state for judgment before the corresponding event is completed.

[0112] Within the closed-loop verification window, the verification agent generates the following verification flag: Device receipt flag. Configure activation flag Link performance recovery flag Alarm convergence indicator Business resumption indicator When the device receipt is valid, ,otherwise When the configuration status is consistent with the executable action, ,otherwise When the link performance recovery status meets the link recovery conditions configured in the system, ,otherwise When the alarm convergence status meets the alarm convergence conditions configured in the system, ,otherwise If the system configuration only uses service OAM detection as the basis for service recovery verification, then The determination is based on the business OAM detection results; if the system configuration only uses bit error rate testing as the basis for business recovery verification, then... The determination is based on the bit error rate test results; if the system configuration uses both the service OAM detection results and the bit error rate test results as the basis for service recovery verification, then if both meet the service recovery conditions configured in the system configuration... If any condition is not met When the business OAM probe results and the bit error rate test results contradict each other and the priority cannot be determined based on the system configuration, the verification feedback status is determined to be a status conflict. If both verification methods are unavailable, then... Instead of directly determining the verification feedback status as successful, the system determines the verification feedback status as a status conflict and triggers manual takeover.

[0113] The verification feedback status is determined according to the following formula:

[0114] in, For actionable procedures The verification feedback status. If multiple branches are satisfied simultaneously, the system prioritizes determining the final verification feedback status in the order of state conflict, successful handling, failed handling, and partial success; when a state conflict is established, the twin object's state version is not updated.

[0115] when At that time, the verification write-back module updates the twin object's state version according to the following formula:

[0116] in, For object The twin object state version before writeback; For object The state version of the twin object after the write-back; This represents the version increment function for system configuration. The version increment function is only used when... Called at time; when In cases of partial success, failure, or state conflict, the system maintains the corresponding twin object's state version unchanged. The verification write-back module simultaneously updates the closed-loop state field of the faulty object to "successful handling," and writes back the updated twin object state version (confirmed by successful handling), handling action number, verification feedback status, device receipt, configuration activation status, link performance recovery status, alarm convergence status, service recovery verification result, and handling control token release status to the real-time mapped digital twin model.

[0117] when In cases of partial success, handling failure, or state conflict, the verification write-back module does not update the twin object's state version but instead writes a conflict record. The conflict record includes the handling action number, handling control token number, twin object state version before rollback, failed verification item, conflicting object, triggered secondary location, and action rollback or manual takeover status. Secondary location or action rollback re-enters the twin state difference gate judgment process.

[0118] During action rollback, the system uses the pre-rollback twin object state version in the handling control token as a reference to determine the device configuration, protection switching state, cross-connection state, or service path state that needs to be restored. The rollback action still needs to pass the twin state difference gate judgment before being issued. After the rollback is completed, the verification agent re-enters the closed-loop verification window to confirm whether the rollback action has taken effect and whether the service state is stable.

[0119] Example 2

[0120] The system for collaborative operation and maintenance of optical transmission networks based on digital twin-driven multi-agent methods includes a telemetry acquisition module, a twin mapping module, a fault propagation graph construction module, a candidate action generation module, a state difference gate module, an agent scheduling module, a handling control module, and a verification and write-back module.

[0121] The telemetry acquisition module is used to collect optical network topology, device status, link performance, alarm information, and service carrying relationships. The telemetry acquisition module can receive data from devices, network management systems, performance acquisition systems, alarm systems, and service resource systems, and then send the collected data to the twin mapping module after attaching a sampling timestamp and source tag.

[0122] The twin mapping module is used to build and maintain real-time mapped digital twin models. It maps physical objects to twin objects and maintains the twin object's state version, sampling timestamp, source tag, confirmation status, controllable attributes, link health field, degradation confidence field, action execution status field, verification feedback status field, and faulty object closed-loop status field for each twin object. The twin mapping module also receives verification feedback status written back by the verification write-back module. When the verification feedback status indicates successful action, it updates the twin object's state version; when the verification feedback status indicates partial success, action failure, or status conflict, it maintains the twin object's state version and marks the corresponding twin object status.

[0123] The fault propagation graph construction module is used to construct a fault propagation graph based on a real-time mapped digital twin model, alarm timing, optical link degradation state variables formed by link performance and the status of devices associated with the link, and service carrying relationships. The module generates topological adjacency edges, protection switching association edges, shared resource edges, service carrying edges, alarm timing edges, and propagation direction edges, and configures node confidence and service impact weights for nodes in the fault propagation graph. The module also updates propagation direction edges when protection switching status or cross-connection status changes.

[0124] The candidate action generation module is used to determine root cause candidates and affected service paths based on the fault propagation graph, obtain candidate actions for handling root cause candidates, and the allowable service impact range, alternative link margin, required link margin, and protection group status corresponding to the candidate actions. Based on a service bearer constraint matrix generated from the affected service paths, service bearer relationships, allowable service impact range, alternative link margin, required link margin, and protection group status, representing the constraints of the service paths, alternative link margins, required link margins, and protection group status on the candidate actions, the module excludes candidate actions that would cause the affected service path to exceed the allowable service impact range, cause the alternative link margin to be less than the required link margin for executing the corresponding candidate action, or conflict with the protection group status, thus obtaining executable action actions. The candidate action generation module distinguishes root cause devices, root cause links, propagation objects, and affected services based on node confidence, link degradation confidence calculated based on optical link degradation state quantities, alarm timing, and service impact weights. It then filters candidate actions based on the service bearer constraint matrix, alternative link margins, protection group status, and action risk level.

[0125] The State Difference Gate module is used to determine the controlled object and exclusive scope based on the executable action before issuing the action. It then performs a twin state difference gate judgment based on the state difference items between the latest physical network state snapshot formed by the device state and the real-time mapped digital twin model for the controlled object and exclusive scope. When no state difference item is obtained, or when the physical objects corresponding to all state difference items do not belong to the controlled object or exclusive scope and do not change the judgment result of whether the controlled object and exclusive scope meet the control security conditions, the State Difference Gate module releases the action control token. When at least one physical object corresponding to a state difference item belongs to the controlled object or exclusive scope, or when at least one state difference item changes the judgment result of whether the controlled object or exclusive scope meets the control security conditions, the State Difference Gate module does not release the action control token, prohibits the execution of executable actions through the device control interface, and updates the latest physical network state snapshot by rereading the device state, link performance, protection switching status, cross-connection status, or service bearer relationship.

[0126] The agent scheduling module is used to schedule predictive agents, localizing agents, handling agents, and verifying agents. Predictive agents generate predicted events based on optical link degradation state variables. Localizing agents generate root cause candidates based on fault propagation graphs. Handling agents access the device control interface and execute executable handling actions only after the twin state difference gate has passed the judgment and the handling control token has been released. Verifying agents generate verification feedback states based on closed-loop verification windows. The agent scheduling module also maintains the handling control token state to prevent handling agents that have not obtained a handling control token from executing conflicting handling actions on physical objects within the same closed-loop verification window.

[0127] The handling control module enables the handling agent to execute executable handling actions. It distributes handling actions to the physical optical transmission equipment and records the action number, target object, action type, equipment acknowledgment, twin object state version before rollback, and associated service path. The handling control module stops subsequent actions when there is an abnormal equipment acknowledgment, insufficient action permissions, or a change in the target object's state, and returns the handling action execution status to the agent scheduling module and the verification write-back module.

[0128] The verification write-back module enables the verification agent to read device receipts, configuration activation status, link performance recovery status, alarm convergence status, and service recovery verification results within the closed-loop verification window. The service recovery verification results are formed from the verification results configured as the basis for service recovery verification in the service OAM detection results and bit error rate test results, thus generating a verification feedback status. When the verification feedback status indicates successful handling, the verification write-back module updates the closed-loop status field of the fault object to "successful handling" and writes the updated twin object status version, confirmed by successful handling, back to the real-time mapped digital twin model. When the verification feedback status indicates partial success, handling failure, or status conflict, the verification write-back module retains the conflict record and triggers secondary location, action rollback, or manual takeover.

[0129] The aforementioned system modules can be deployed within the same network controller or distributed across digital twin platforms, network management systems, automated operation and maintenance platforms, and device control platforms. Modules can interact via message queues, remote call interfaces, database interfaces, or control buses. Regardless of the deployment method, the aforementioned data flow and control flow relationships are maintained between telemetry acquisition, twin mapping, fault propagation graph construction, action generation, state difference gates, tokenization control, action execution, and verification write-back.

[0130] Example 3

[0131] In an improved embodiment, the fault propagation graph construction module maintains the object relationships in the optical transmission network hierarchically according to edge types. Topological adjacency edges record the connection relationships between nodes, ports, links, channels, and fiber segments. Protection switching association edges record the working paths, protection paths, and current switching states within protection groups. Shared resource edges record the sharing relationships of multiple service paths, links, or channels to the same fiber segment, amplifier, port, or link resource. Service bearer edges record the correspondence between service paths and underlying bearer objects. Alarm timing edges record the chronological and propagation relationships between different alarm events. Propagation direction edges record the propagation direction of optical signals in the current cross-connection state and protection switching state.

[0132] When the link degradation confidence level, protection switching status, cross-connection status, or service carrying relationship changes, the fault propagation graph construction module updates the confidence level of relevant nodes, service impact weights, and propagation direction edges. In this way, the fault propagation graph is not a static topology graph, nor is it a simple alarm cause-effect graph, but a cross-layer propagation structure that is updated as the physical state, service carrying status, and protection status of the optical transmission network change.

[0133] In an improved embodiment, the service bearer constraint matrix uses the affected service path and candidate actions as indexes to record the service level, protection attribute, alternative link margin, required link margin, protection group status, action risk level, allowed service impact range, and exclusion results. The exclusion results in the service bearer constraint matrix are used to indicate whether the corresponding candidate action can enter the set of executable actions.

[0134] For protection switching actions, the candidate action generation module checks the protection path status, alternative link margin, and the level of affected services. If the protection path does not exist, is in a degraded state, or the high-level services carried by the protection path would be affected beyond the allowable service impact range, the protection switching action is excluded. For service rerouting actions, the candidate action generation module checks the link performance, service level, and service protection attributes traversed by the alternative path. If the performance of the alternative path is insufficient to carry the affected services, the service rerouting action is excluded. For port reset actions, the candidate action generation module checks the service path carried by the port and whether an available protection path exists. If a port reset would cause high-level services to be unprotected and interrupted, the port reset action is prohibited. For test loopback actions, the candidate action generation module checks whether the test object is isolated from online services, or whether the test action is only used to reread status data and does not change the status of online services; if the test action would affect online services, the test loopback action is excluded.

[0135] Example 4

[0136] In an improved embodiment, the state difference gate module classifies the differences between the latest state snapshot formed by the device states in the physical optical transport network and the real-time mapped digital twin model into differences that affect control security and differences that do not affect control security.

[0137] Differences affecting control security include: inconsistencies between the port status of the object to be controlled on the physical device side and the twin model side; protection group status changes on the physical device side but the twin model is not updated; cross-connection status changes on the physical device side but the twin model retains the old connection; the performance status of the link to be controlled changes after the executable action is generated and affects the controlled object; the status of the object to be controlled represented by the obtained device interface receipt is inconsistent with the status of the corresponding twin object; and the object to be controlled is locked by other action control tokens. When the above differences occur, the state difference gate module keeps the action control token from being released; for executable actions that will change device configuration, protection switching status, cross-connection status, port status, or service bearer status, the state difference gate module prohibits the issuance of such actions; when a state difference item causes the action control token to remain unreleased, the state difference gate module rereads the device status, link performance, protection switching status, cross-connection status, or service bearer relationship to update the latest state snapshot.

[0138] Differences that do not affect control security include: object state changes that are not topologically related to the object to be controlled; explicit attribute changes that are not exclusive to the executable action; and auxiliary state changes that do not affect protection switching, cross-connection, or port control. When the above differences occur, the state difference gate module can continue to release the action control token, but the difference should be recorded in the twin object state record for subsequent verification write-back and state synchronization.

[0139] By classifying the differences as described above, the system can avoid reducing the automation coverage rate by prohibiting automatic processing for any state change, and it can also avoid continuing to execute automatic processing actions when there are critical differences that affect control security.

[0140] Example 5

[0141] In an improved embodiment, the disposal control token includes a token number, locked object, action type, exclusive scope, timeout condition, rollback condition, twin object state version before rollback, token release condition, and associated verification window. The locked object can be a port, link, board, protection group, or cross-connection. The exclusive scope is the set of physical objects that conflict with the locked object. The timeout condition limits the validity of the disposal control token within the closed-loop verification window. The rollback condition limits the rollback trigger conditions in case of disposal failure, state conflict, or closed-loop verification window timeout. The twin object state version before rollback is used to determine the state version referenced for the rollback operation.

[0142] When a disposal agent that has not obtained a disposal control token requests to perform a disposal action on a physical object within the exclusive scope within the closed-loop verification window, if the request is only used to reread the device status, link performance, protection switching status, cross-connection status, or service bearer relationship, and does not change the status of the physical object within the exclusive scope, the agent scheduling module allows the execution of the request to update the latest status snapshot; if the request would change the status of the physical object within the exclusive scope, the agent scheduling module rejects the request or suspends it until the current disposal control token is released before re-evaluating.

[0143] The predictive agent, localizing agent, handling agent, and verifying agent access the same real-time mapped digital twin model through the agent scheduling module. Each agent does not directly bypass the twin model, the twin state difference gate, and the handling control token to access the device control interface. The predictive agent reads the optical link degradation status quantity, link health field, and degradation confidence field to form a predicted event. The localizing agent reads the predicted event, alarm information, and fault propagation graph to form root cause candidates and the scope of affected services. The handling agent reads the root cause candidates, the service bearer constraint matrix, and executable handling actions, and accesses the device control interface after obtaining the handling control token. The verifying agent reads the handling action execution status and closed-loop verification window configuration, collects multi-source verification feedback, and forms a verification feedback status.

[0144] When multiple agents request disposal control tokens for the same port, link, board, protection group, or cross-connection, the agent scheduling module identifies conflicting requests based on the locked object and exclusivity scope. For conflicting requests, the agent scheduling module allows only one agent to obtain the disposal control token. Other agents' requests are suspended, rejected, or allowed to execute only to update the latest state snapshot if the request changes the state of the physical object within the exclusivity scope. The disposal control token is released after the closed-loop verification window ends, the verification feedback status is confirmed, the action rollback is completed, or manual takeover occurs.

[0145] Example 6

[0146] In an improved embodiment, the verification agent performs verification within a closed-loop verification window in the following order: device acknowledgment, configuration activation status, link performance recovery status, alarm convergence status, and service recovery verification result. This order ensures that whether the device executes an action, whether the action takes effect on the device side, whether link performance is restored, whether alarms are converged, and whether services are restored are all verified.

[0147] If the device feedback indicates that the action was not accepted by the device, the verification agent directly establishes a failure status and triggers secondary location or manual takeover. If the device feedback is valid but the configuration status is inconsistent with the executable action, the verification agent establishes a state conflict or failure status. If the configuration status is consistent with the executable action but the link performance recovery status does not meet the system-configured link recovery conditions, the verification agent marks the relevant link nodes in the fault propagation graph as unrecovered and triggers secondary location. If the link performance recovery status meets the system-configured link recovery conditions but the service recovery verification result does not meet the system-configured service recovery conditions, the verification agent marks the corresponding service path as unrecovered and triggers a service bearer relationship review. If the device feedback, configuration status, link performance recovery status, alarm convergence status, and service recovery verification all meet the corresponding conditions, the verification agent establishes a successful status.

[0148] After a successful resolution status is established, the verification write-back module updates the closed-loop status field of the fault object to "successful resolution," updates the corresponding twin object status version to the confirmed successful resolution twin object status version, and writes the resolution action number, verification feedback status, device receipt, configuration effective status, link performance recovery status, alarm convergence status, and service recovery verification result into the twin status write-back table. If a partially successful, failed resolution, or conflicting status is established, the verification write-back module does not update the twin object status version but retains the conflict record and triggers secondary location, action rollback, or manual takeover.

[0149] During action rollback, the system uses the pre-rollback twin object state version in the handling control token as a reference to determine the device configuration, protection switching state, cross-connection state, or service path state that needs to be restored. The rollback action still needs to pass the twin state difference gate judgment before being issued. After the rollback is completed, the verification agent re-enters the closed-loop verification window to confirm whether the rollback action has taken effect and whether the service state is stable.

[0150] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks, characterized in that: Includes the following steps: Step 1: Collect the optical network topology, device status, link performance, alarm information and service carrying relationships of the physical optical transmission network, and construct a real-time mapped digital twin model that includes the twin object status version, sampling timestamp, source tag and fault object closed-loop status fields; Step 2: Based on the real-time mapped digital twin model, the link performance, the device status, the alarm timing, and the service carrying relationship, construct a fault propagation graph with propagation direction, node confidence, and service impact weight; Step 3: Based on the fault propagation graph, determine the root cause candidate objects, affected service paths, and candidate handling actions. Based on the service bearer constraint matrix that represents the service path, alternative link margin, required link margin, and protection group status constraining the candidate handling actions, eliminate candidate handling actions that do not meet the service bearer constraints to obtain executable handling actions. Step 4: Before the executable action is issued, determine the control object and exclusive scope of the executable action. The exclusive scope is a set of physical objects that have control conflicts with the control object. Based on the latest state snapshot of the physical network formed by the device state and the real-time mapped digital twin model, perform a twin state difference gate judgment for the state difference items of the control object and the exclusive scope. When no state difference item is obtained, or when the physical objects corresponding to all the state difference items do not belong to the controlled object or the exclusive scope and the judgment result of whether the controlled object and the exclusive scope meet the control security conditions is not changed, the disposal control token is released. When at least one of the physical objects corresponding to the state difference items belongs to the controlled object or the exclusive scope, or when at least one of the state difference items changes the judgment result of whether the controlled object or the exclusive scope meets the control security conditions, the disposal control token is not released, the executable disposal action is prohibited from being executed through the device control interface, and the latest state snapshot of the physical network is updated by rereading the device status, link performance, protection switching status, cross-connection status or service bearing relationship. Step 5: Schedule the disposal agent and enable the disposal agent to execute the executable disposal action through the device control interface after obtaining the disposal control token. Within the configured closed-loop verification window, the disposal agent that has not obtained the disposal control token is prohibited from executing conflict disposal actions on physical objects within the exclusive scope. Step Six: Schedule the verification agent and have it read the device receipt, configuration activation status, link performance recovery status, alarm convergence status, and service recovery verification results within the closed-loop verification window. The service recovery verification results are formed from the verification results configured as the basis for service recovery verification in the service OAM detection results and bit error test results, thus forming a verification feedback status. When the verification feedback status is "successful handling," update the closed-loop status field of the fault object to "successful handling," and write the updated twin object status version confirmed by successful handling back to the real-time mapped digital twin model. When the verification feedback status is "partially successful," "failed handling," or "state conflict," retain the conflict record and trigger secondary positioning, action rollback, or manual takeover.

2. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The construction of a real-time mapped digital twin model, including fields for twin object state version, sampling timestamp, source tag, and closed-loop state of faulty object, includes: The nodes, ports, links, channels, fiber segments, protection groups, cross-connections, and service paths in the physical optical transmission network are mapped to twin objects, and each twin object is configured with twin object status version, sampling timestamp, source tag, acknowledgment status, controllable attributes, and fault object closed-loop status fields.

3. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The link performance includes received optical power, transmitted optical power, optical signal-to-noise ratio (OSNR), bit error rate (BER), forward error correction (FEC) margin, bit error count, and historical performance trends. The device status includes temperature, current, voltage, port status, board status, and optical module status associated with the link. The link performance and the device status are used to form an optical link degradation status quantity, and the link degradation confidence level is calculated based on the optical link degradation status quantity.

4. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The fault propagation graph includes topological adjacency edges, protection switching associated edges, shared resource edges, service bearer edges, alarm timing edges, and propagation direction edges; wherein, the propagation direction edges are determined based on the optical signal propagation direction, cross-connection relationships, and protection switching status, and are updated when the protection switching status or cross-connection status changes.

5. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The process of determining root cause candidates, affected service paths, and candidate actions based on the fault propagation graph includes: Based on the node confidence in the fault propagation graph, the link degradation confidence calculated based on the link performance and the device status, the alarm timing, and the service impact weight, the root cause device, root cause link, propagation object, and the affected service path are distinguished, and the candidate actions for handling are formed for the root cause device or the root cause link.

6. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The service carrying constraint matrix is ​​generated from the affected service path, the service carrying relationship, the allowed service impact range, the alternative link margin, the required link margin, the protection group status, and the action risk level; it excludes candidate actions that would cause the affected service path to exceed the allowed service impact range, cause the alternative link margin to be less than the required link margin, or conflict with the protection group status.

7. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The process of performing a twin state difference gate determination based on the latest physical network state snapshot formed by the device state and the real-time mapped digital twin model for the state difference items of the controlled object and the exclusive range includes: Based on the latest physical network state snapshot formed by the device state, the twin object state version, the sampling timestamp, the source tag, and the obtained device interface receipt, the state of the object to be controlled, the protection switching state, the cross-connection state, and the port state in the latest physical network state snapshot are compared with the corresponding states in the real-time mapped digital twin model to obtain the state difference items; when no state difference items are obtained, or when the physical objects corresponding to all the state difference items do not belong to the controlled object or the exclusive scope and do not change the judgment result of whether the controlled object and the exclusive scope meet the control security conditions, the judgment is passed.

8. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The disposal control token includes the locked object, action type, exclusive scope, timeout condition, rollback condition, and twin object state version before rollback; the locked object is a port, link, board, protection group, or cross-connection; the control conflict includes physical objects within the exclusive scope belonging to the same port, link, board, protection group, or cross-connection as the controlled object, or there being a control dependency between them; within the closed-loop verification window, the disposal agent that has not obtained the disposal control token shall not perform conflict disposal actions on physical objects within the exclusive scope.

9. The digital twin-driven multi-agent collaborative operation and maintenance method for optical transmission networks according to claim 1, characterized in that, The closed-loop verification window is configured based on device type, service level, performance counter refresh cycle, alarm clearing delay, protection switching convergence time, and service detection cycle. When the device acknowledgment is valid, the configuration effective status is consistent with the executable action, the link performance recovery status and alarm convergence status meet the system configuration conditions, and all verification results configured as the basis for service recovery verification meet the system configuration service recovery conditions, the verification feedback status is "processing successful".