Fault self-repairing processing method for power communication transmission network
Through intelligent fault analysis and self-repair processes, faults in the power communication transmission network can be quickly located, dynamic priority scheduling can be performed to ensure the reliability of service recovery, fault repair files can be generated, the self-healing capability of the power communication transmission network can be improved, and the problem of low efficiency in traditional manual processing can be solved.
Patent Information
- Application Number
- CN202511441969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-09
AI Technical Summary
The current fault handling of power communication transmission networks relies on manual experience, resulting in low processing efficiency and slow response speed, which cannot meet the requirements of high reliability and rapid self-healing.
By receiving abnormal event alarm signals, identifying alarm types, performing intelligent fault analysis, quickly locating affected services and optical path resources, dynamically distinguishing service priorities, verifying N-1 redundancy capabilities, enabling temporary circuit orchestration and protection channel switching, conducting performance data collection and service traffic testing, generating fault repair files, and updating the knowledge base.
It enables rapid and accurate fault location and self-repair in power communication transmission networks, shortens service recovery time, improves network reliability and self-healing capabilities, and forms an intelligent fault handling closed loop.
Smart Images

Figure CN121309313A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power transmission technology, and in particular to a method for self-repairing faults in power communication transmission networks. Background Technology
[0002] With the continuous improvement of the intelligence level of power systems, the scale of power communication transmission networks is expanding and the network structure is becoming increasingly complex. Their stable operation has become a crucial cornerstone for ensuring power system security. However, in actual operation, various equipment failures and service alarms inevitably occur in this network. Currently, the mainstream fault handling method relies on the manual experience of maintenance personnel for alarm analysis, fault location, and repair operations. This process typically involves manually sifting through massive amounts of alarms to identify key information, gradually diagnosing fault points, and manually executing repair commands. This model has inherent drawbacks such as low processing efficiency and slow response speed. Especially when a fault occurs, a large number of derivative alarm messages can create interference, making it difficult for maintenance personnel to quickly and accurately determine the root cause of the fault. Therefore, it cannot meet the stringent requirements of power services for high reliability and rapid self-healing in communication transmission networks.
[0003] In view of this, a self-healing method for fault handling in power communication transmission networks is proposed. Summary of the Invention
[0004] This invention provides a self-healing method for faults in power communication transmission networks, which addresses the problem of failing to meet the stringent requirements of power services for high reliability and rapid self-healing in communication transmission networks.
[0005] This invention provides a self-repair method for fault handling in power communication transmission networks, comprising: Receive abnormal event alarm signals and identify the alarm type; Based on the alarm type, perform fault analysis, query the affected circuit services and optical resource IDs in the associated resource view, and generate fault detection results within the set time window; Based on the fault detection results, a self-repair process is initiated to extract the attribute information of the affected services to distinguish service priorities and verify the N-1 redundancy capability of the services. If the verification result requires the use of a temporary channel, then the temporary circuit orchestration is started, resource verification is performed, and after the verification is passed, the service traffic is migrated to the temporary circuit or the protection channel is switched over. Once the original optical path fault is repaired, the service rollback process is initiated, and performance data collection and service traffic testing are performed on the repaired original optical path. If the test passes, the service traffic will be migrated back to the original optical path in a dual-channel parallel manner, the resources occupied by the temporary channel will be released, and the network topology and service routing table will be updated. After the temporary circuit is released, a fault repair file is generated, the root cause of the fault is analyzed, the processing procedure and resource usage details are recorded, performance data is archived to generate optimization suggestions and update the knowledge base.
[0006] Furthermore, the alarm types include device and link alarms, optical signal interruption alarms, service layer alarms, optical power abnormality alarms, and peer optical port alarms.
[0007] Furthermore, the step of performing fault analysis based on the alarm type, querying the associated resource view for affected circuit services and optical resource IDs, and generating fault detection results within a set time window includes: If the alarm type is a device and link alarm, then determine whether the optical path is abnormal by combining the network element performance data; If the alarm type is an optical signal interruption alarm, it is directly determined to be an optical path interruption; If the alarm type is a service layer alarm, then analyze the commonalities of service routes and performance data; If the alarm type is an optical power abnormality alarm, then the optical port fault is determined by combining the real-time port alarm. If the alarm type is a peer optical port alarm, then verify the optical port status in conjunction with performance data.
[0008] Furthermore, the extraction of attribute information of affected services to distinguish service priorities includes: Extract the attributes, topology level, and resource dependency information of the affected services; Based on the dynamic correlation map of business and resources, other potentially affected business chains that are strongly related to the business are automatically marked in order to provide early warning of chain impacts; Based on the extracted information, distinguish the priority and protection level of the business; By combining real-time peak traffic data, the priority of services can be dynamically fine-tuned.
[0009] Furthermore, the N-1 redundancy capability of the verification service includes: Initiate a health scan of redundant channels to assess the historical failure rate and recent performance fluctuations of the backup channels; If the service does not have a protection channel configured or the protection channel is in an abnormal state, a temporary circuit orchestration process will be initiated, and a fast query of the cross-domain resource pool will be triggered to prioritize the use of low-load resources. If the service has a protection channel configured and the channel is in normal condition, the automatic switching of the protection channel will be triggered directly, and a status snapshot containing the channel performance parameters at the moment of switching will be pushed to the operation and maintenance personnel.
[0010] Furthermore, the performance acquisition and service traffic testing of the repaired original optical path includes: Automatically trigger optical path performance acquisition, and collect basic indicators of the original optical path such as input optical power, output optical power and bit error rate; By associating the historical fault types and repair process records of the original optical path, a digital health profile of the optical path is generated to visually demonstrate the performance change trend before and after the repair. Based on the historical traffic fluctuation patterns of the affected services, dynamic simulated traffic is generated to simulate and test the original optical path.
[0011] Furthermore, the performance acquisition and service traffic testing of the repaired original optical path also includes: Extreme scenario simulations were incorporated into the simulation test to verify the stress resistance of the original optical path; Determine whether the simulated test passes the test criterion based on business experience similarity; If the test fails, the optical path is marked as potentially not fully repaired, and a suspected root cause analysis report is generated to notify maintenance personnel to intervene and investigate. If the test passes, an optical path repair verification report is generated and the switchback process is triggered.
[0012] Furthermore, the method of migrating service traffic back to the original optical path using a dual-channel parallel approach includes: In the back-switching process, the service traffic is transmitted simultaneously on the temporary channel and the original optical path; Real-time monitoring of the load rate and delay jitter of the temporary channel and the original optical path; Based on real-time monitoring data, the algorithm dynamically adjusts the distribution ratio of business traffic on the two channels.
[0013] Furthermore, the process of releasing the resources occupied by the temporary channel and updating the network topology and service routing table includes: After all service traffic is migrated back to the original optical path, the port and bandwidth resources occupied by the temporary channel are released. If similar business demand is predicted in the near future, the released resources will be marked as priority reserved resources; otherwise, they will be marked as ordinary idle resources. Update the network topology and service routing table to reflect the actual status after resource release and service rollback.
[0014] Furthermore, after the temporary circuit is released, a fault repair file is generated, the root cause of the fault is analyzed, the processing procedure and resource usage details are recorded, performance data is archived to generate optimization suggestions, and the knowledge base is updated, including: Generate a fault repair file that includes a root cause analysis of the fault, timestamps of the processing process, and details of resource usage; Archive performance data during the operation of the temporary channel and the service switchback test phase; Based on the aforementioned fault repair files and archived performance data, optimization suggestions are generated; Update the fault case, handling logic, and optimization suggestions to the knowledge base and associate them with similar fault modes.
[0015] As can be seen from the above technical solutions, the present invention has the following advantages: This invention receives abnormal event alarm signals and identifies alarm types, performs intelligent fault analysis based on these types, and quickly locates affected services and optical path resources. Fault detection is completed within a set time window, solving the problems of low efficiency and inaccurate fault location associated with traditional manual troubleshooting. By extracting service attributes to dynamically differentiate priorities and combining this with N-1 redundancy capability verification, precise resource scheduling and service assurance for fault handling are achieved, effectively avoiding service interruptions caused by protection channel anomalies. Temporary circuit orchestration and a triple resource verification mechanism ensure the rapid and reliable establishment of temporary channels, significantly shortening service recovery time. Performance data acquisition and real traffic testing after original optical path repair, along with a dual-channel parallel backswitching strategy, ensure the safety of the service backswitching process. Finally, by generating fault repair archives, archiving performance data, and updating the knowledge base, the system's predictive and autonomous capabilities for similar faults are significantly improved. This invention achieves full-process automation and intelligence from fault detection and intelligent handling to service backswitching and knowledge accumulation, effectively solving problems such as low fault handling efficiency, difficulty in determining root causes, and long service recovery cycles in power communication transmission networks, significantly improving network reliability and self-healing capabilities. Attached Figure Description
[0016] Figure 1 This is a schematic flowchart of an embodiment of a self-repair method for fault handling in a power communication transmission network according to the present invention; Figure 2 This is a schematic diagram of the alarm identification and type determination process in this invention; Figure 3 This is a schematic diagram of the alarm analysis and self-repair process in this invention; Figure 4 This is a schematic diagram of the original optical path repair and service rerouting process in this invention; Figure 5 This is a schematic diagram of the fault repair archiving and optimization process in this invention. Detailed Implementation
[0017] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “corresponding to,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0018] Example 1 The implementation method in this embodiment can be implemented in a system, on a server, or on a terminal; no specific limitation is made. The method in this application will be described below from the perspective of system implementation. Please refer to... Figure 1 The method provided in this application includes the following steps: S11. Receive abnormal event alarm signals and identify the alarm type; S12. Perform fault analysis based on alarm type, query the affected circuit services and optical resource IDs in the associated resource view, and generate fault detection results within the set time window; S13. Based on the fault detection results, initiate the self-repair handling process, extract the attribute information of the affected services to distinguish service priorities, and verify the N-1 redundancy capability of the services. S14. If the verification result requires the use of a temporary channel, then start the temporary circuit orchestration, perform resource verification, and after the verification is passed, migrate the service traffic to the temporary circuit or trigger the protection channel switchover. S15. After the original optical path fault is repaired, the service switchback process is initiated, and the performance of the repaired original optical path is collected and the service traffic is tested. S16. If the test passes, the service traffic will be migrated back to the original optical path in a dual-channel parallel manner, the resources occupied by the temporary channel will be released, and the network topology and service routing table will be updated. S17. After the temporary circuit is released, generate a fault repair file, analyze the root cause of the fault and record the processing process and resource usage details, archive performance data to generate optimization suggestions and update the knowledge base.
[0019] During implementation, the system receives abnormal event alarm signals and identifies the specific alarm type. Based on the alarm type, it performs intelligent fault analysis by associating the resource view, quickly locating the affected circuit services and optical resource IDs, and generating accurate fault detection results within a strict time window. Based on these results, the system automatically initiates a self-healing process, dynamically prioritizing services by extracting service attribute information and verifying the N-1 redundancy capability of the services. If a temporary channel needs to be used, a temporary circuit orchestration with a triple-check mechanism is initiated. After ensuring resource compatibility and reliability, service traffic is seamlessly migrated to the temporary channel or protection switching is triggered. When the original optical path is repaired, the system does not immediately switch back. Instead, it first performs rigorous performance data collection and simulates real traffic load testing. Only after the tests pass does it gradually and smoothly migrate services back to the original optical path using a dual-channel parallel approach, releasing temporary resources and updating the network status. Finally, the system generates detailed fault repair archives, archives full-process performance data, and updates the knowledge base, forming a continuous optimization closed loop. This enables the system to possess complete self-healing capabilities, from rapid fault response and intelligent handling to experience accumulation.
[0020] Example 2 Please see Figure 2 Alarm identification and type determination include the following: In this embodiment, alarm types include device and link alarms, optical signal interruption alarms, service layer alarms, optical power abnormality alarms, and peer optical port alarms.
[0021] The fault analysis logic for different alarm types is as follows: 1. If the alarm type is a device and link alarm: the system will retrieve the performance data (such as CPU load, memory utilization, port error count) of the network element (such as switch, router) that generated the alarm, and analyze the abnormal trends of these data to determine whether it is a device failure or a performance degradation of its associated optical path, so as to avoid misjudgment caused by a single alarm.
[0022] 2. If the alarm type is an optical signal interruption alarm: This alarm is a direct and clear indication of optical path interruption. Therefore, the system does not need to perform complex performance correlation analysis and can directly and quickly determine that it is an optical path interruption, which greatly shortens the location time of critical faults.
[0023] 3. If the alarm type is a service layer alarm: the system will analyze the routing information of multiple affected services, find common network elements or optical path segments on their paths, and combine the performance data of these common segments (such as latency and jitter) for comprehensive analysis, thereby locating the common fault points that cause batch service anomalies.
[0024] 4. If the alarm type is an optical power abnormality alarm: the system will immediately associate the real-time status information of the port that generated the alarm. If the port status is abnormal, it can be quickly determined that the problem is a hardware or configuration failure of the optical port, rather than a problem with the remote line.
[0025] 5. If the alarm type is a peer optical port alarm: the system will combine local performance data (such as received optical power and bit error rate) to verify the authenticity of the peer optical port alarm and the actual impact on local services, and distinguish whether it is a remote equipment failure or a link transmission quality problem.
[0026] After completing the preliminary fault analysis based on alarm type, the system will associate the resource view, accurately query all circuit services affected by the fault and their corresponding optical path resource IDs, and complete all diagnostic steps within a preset time window to generate fault detection results containing fault location and scope of impact, providing clear input for subsequent self-repair handling.
[0027] Example 3 Please see Figure 3 Extracting attribute information of affected services to differentiate service priorities includes the following steps: 1. Extract the attributes, topology level, and resource dependency information of the affected services; This step corresponds to extracting two nodes in the process: affected service attributes, topology level, and resource dependency information. After initiating the self-healing process, the system immediately extracts key information about the affected service from the resource management system; attributes include service ID, bandwidth, and the customer to which the service belongs; the topology level refers to the location of the service path in the network structure (such as the core layer or aggregation layer); and the resource dependency information specifies the specific optical paths, ports, and devices that the service depends on.
[0028] 2. Based on the dynamic relationship map of business and resources, automatically mark other potentially affected business chains that are strongly related to the business, in order to provide early warning of chain impact; This step involves intelligent in-depth processing of the extracted information, with the system utilizing a real-time updated service-resource relationship graph for analysis. For example, if the currently faulty service shares the same critical optical path with another important service, the system will automatically identify the important service as a potentially affected service and send a cascading impact warning to the operations and maintenance interface, thereby achieving a proactive assessment of the impact scope.
[0029] 3. Based on the extracted information, differentiate the priority and protection level of the business; This step corresponds to the node in the attached diagram that distinguishes between business priority and guarantee level. The system assigns initial priority (e.g., emergency, high, medium, low) and guarantee level to the business based on the preset strategy (e.g., prioritizing power production control business) and the extracted business attributes (e.g., customer level, business type).
[0030] 4. By combining real-time business traffic peak data, the priority of business operations can be dynamically fine-tuned.
[0031] This step introduces a dynamic adjustment mechanism based on the initial classification. The system will query the current traffic data of the service in real time. If it finds that a service, although initially classified as having a normal classification, experiences a sudden surge in traffic during a critical period, the system will temporarily increase its processing priority to ensure that resource scheduling strategies align with real-time network load.
[0032] In this embodiment, the N-1 redundancy capability node for the verified service and its subsequent startup temporary circuit orchestration process and triggering automatic failover of the protection channel are two branch paths. The specific implementation process is described below: 1. Initiate a health scan of redundant channels to assess the historical failure rate and recent performance fluctuations of the backup channels; After reaching the N-1 redundancy capability node for verification services, the system queries its historical fault records to assess reliability and checks recent performance index fluctuations (such as optical power and bit error rate) to confirm its current stability.
[0033] 2. If the service does not have a protection channel configured or the protection channel is in an abnormal state, the temporary circuit orchestration process will be started, and a fast query of the cross-domain resource pool will be triggered to prioritize the use of low-load resources; This branch corresponds to the "no" path in the attached diagram, leading to the initiation of the temporary circuit orchestration process. When the system determines that the service has no protected channel or the channel is unhealthy, it initiates the temporary circuit construction process. During this process, the system will query available resource pools across regions (such as different data centers or different network domains) and prioritize the resources with the lowest load rate (such as idle ports or available bandwidth) for use, in order to improve the orchestration success rate and the quality of the temporary circuit.
[0034] 3. If the service has been configured with a protection channel and the channel status is normal, the automatic switching of the protection channel will be triggered directly, and a status snapshot containing the channel performance parameters at the moment of switching will be pushed to the operation and maintenance personnel.
[0035] This branch corresponds to the verified path in the attached diagram, which triggers automatic switching of the protection channel. Upon successful verification, the system immediately executes the switching operation. Simultaneously, to facilitate traceability and analysis by maintenance personnel, the system captures and pushes key performance parameters of the protection channel (such as received optical power and transmitted optical power) at the moment the switching action occurs, creating a status snapshot.
[0036] Example 4 Please see Figure 4 The troubleshooting and optimization process includes the following: In this embodiment, performance data acquisition and service traffic testing are performed on the repaired original optical path, including: 1. Automatically trigger optical path performance acquisition, collecting basic indicators of the original optical path such as input optical power, output optical power, and bit error rate; 2. Link the historical fault types and repair process records of the original optical path to generate a digital health profile of the optical path that visually displays the performance change trend before and after repair; 3. Based on the historical traffic fluctuation patterns of the affected services, generate dynamic simulated traffic to simulate and test the original optical path.
[0037] This section corresponds to the two key nodes in the attached diagram: automatically triggering optical path performance acquisition and simulating real service traffic testing. In practice, after initiating the service rollback process, the system first automatically triggers performance acquisition (automatically triggering optical path performance acquisition), collecting basic performance indicators (input / output optical power, bit error rate) of the repaired optical path from network elements, providing a data foundation for assessing its physical layer status. Subsequently, the system associates the historical records of the optical path (such as past fault causes and the technology used in this repair) and generates a comprehensive digital profile of the optical path's health through data visualization technology. This profile allows for a direct comparison of performance curves before and after repair, assisting in judging the quality of the repair. Finally, to make the testing more closely resemble real service load, the system analyzes the historical traffic data of the services carried on the optical path (such as recent periodic fluctuations) and generates dynamically changing simulated traffic (simulating real service traffic testing) to test the optical path's carrying capacity, rather than simply performing a connectivity test.
[0038] In this embodiment, based on the above-mentioned parts, the performance acquisition and service traffic testing of the repaired original optical path are further limited, and it also includes: 1. Incorporate extreme scenario simulations into the testing to verify the resilience of the original optical path; 2. Determine whether the simulated test passes the test criterion based on the similarity of business experience; 3. If the test fails, the optical path is marked as potentially not fully repaired, and a suspected root cause analysis report is generated to notify the maintenance personnel to intervene and investigate. 4. If the test passes, an optical path repair verification report will be generated and the switchback process will be triggered.
[0039] This section corresponds to the test pass / fail decision node and its two branch paths in the attached diagram. In practice, firstly, based on regular simulation testing, the system injects peak traffic or simulates instantaneous surges and other extreme scenarios (extreme scenario rehearsal) to evaluate the stability of the optical path under pressure. When determining whether the test passes, in addition to checking basic communication parameters (connectivity, bit error rate), service experience similarity is introduced as a criterion. This involves comparing the similarity of key experience indicators (such as latency, jitter) between the test traffic and real service traffic to ensure that the optical path can provide qualified service carrying quality. This judgment corresponds to the test pass / fail decision point in the diagram: if it fails (no branch), the system automatically marks the optical path as potentially incompletely repaired and generates a root cause analysis report based on performance data and test anomalies (e.g., indicating that the attenuation of a certain fiber segment still exceeds the threshold), then triggers a notification to maintenance personnel for access and troubleshooting; if it passes (yes branch), the system generates a detailed repair verification report and executes the trigger switchback process shown in the diagram to enter the next stage.
[0040] In this embodiment, a dual-channel parallel approach is used to migrate service traffic back to the original optical path, including: 1. During the switchback process, service traffic is transmitted simultaneously on both the temporary channel and the original optical path; 2. Real-time monitoring of the load rate and delay jitter of the temporary channel and the original optical path; 3. Based on real-time monitoring data, the algorithm dynamically adjusts the distribution ratio of business traffic on the two channels.
[0041] This section corresponds to the dual-channel parallel migration of service nodes shown in the attached diagram. In practice, after triggering the switchback process, the system does not immediately disconnect the temporary channel. Instead, it simultaneously distributes service traffic to both the temporary channel and the original optical path (dual-channel parallel transmission). During this process, the system monitors the performance metrics of both paths in real time, particularly load balancing and latency jitter. Then, based on this real-time data, it dynamically adjusts the traffic allocation ratio using a preset scheduling algorithm (such as a weighted or QoS-based strategy). (For example, gradually increasing the traffic share of the original optical path from 10% to 100%), thereby achieving a smooth, seamless migration of services from the temporary channel to the original optical path, minimizing service interruption or quality degradation.
[0042] In this embodiment, the release of resources occupied by the temporary channel is limited, and the network topology and service routing table are updated, including: 1. After all business traffic has been migrated back to the original optical path, release the port and bandwidth resources occupied by the temporary channel; 2. If similar business demand is predicted in the near future, the released resources will be marked as priority reserved resources; otherwise, they will be marked as ordinary idle resources. 3. Update the network topology and service routing table to reflect the actual status after resource release and service rollback.
[0043] This section corresponds to the nodes in the attached diagram that release temporary channel resources, mark resources as idle and available, and update the network topology and service routing table after the service rollback is completed. In practice, once the system confirms that service traffic has been completely migrated back to the original optical path (service rollback completed), it immediately releases the specific resources (such as port numbers and bandwidth quotas) occupied by the temporary channel. Next, the system determines, based on historical data or predictive models: if similar faults or service demands are likely in the near future, these resources are marked as priority reservations for rapid response; otherwise, they are marked as ordinary idle resources and added to the global resource pool. Finally, the system synchronously updates the network topology view and service routing table to ensure that the resource status and service routing information presented by the network management system are consistent with the actual network situation, providing an accurate basis for subsequent network operations, thereby completing the entire rollback and resource cleanup process.
[0044] Example 5 Please see Figure 5 In this embodiment, after the temporary circuit is released, a fault repair file is generated, the root cause of the fault is analyzed, the processing procedure and resource usage details are recorded, performance data is archived to generate optimization suggestions, and the knowledge base is updated, including: 1. Generate a fault repair file that includes a root cause analysis of the fault, timestamps of the processing process, and details of resource usage; 2. Archive performance data during the temporary channel operation and service switchback testing phase; 3. Generate optimization suggestions based on fault repair files and archived performance data; 4. Update the knowledge base with this failure case, handling logic, and optimization suggestions, and associate it with similar failure modes. This section corresponds to the entire archiving optimization process shown in the attached diagram, from the completion of the temporary circuit release to the end of the process. Its specific implementation process is closely aligned with the nodes in the attached diagram, as explained below: Generate a fault repair file containing root cause analysis, processing timestamps, and resource usage details: This step corresponds to the nodes in the attached diagram, such as generating the fault repair file, analyzing the root cause of the fault, recording the processing process and timestamps, and recording resource usage details. In specific implementation, after the temporary circuit resources are released, the system first automatically generates a structured fault repair file. This file not only contains the final analysis conclusion of the fault root cause (such as fiber breakage or optical port failure), but also accurately records the time consumed at key nodes throughout the entire process, from alarm triggering, fault location, temporary circuit orchestration, service switchback to resource release, using timestamps. Simultaneously, the file records detailed resource details such as port identifiers, bandwidth, and routing paths used by the temporary circuit, forming a complete fault handling data packet.
[0045] Archive performance data during the operation of the temporary channel and the service switchback test phase: This step directly corresponds to the archived performance data node in the attached diagram. The system will classify, compress, and store the performance data of the temporary channel throughout the entire service cycle (such as average latency, peak bandwidth utilization, and bit error rate) and the test data of the original optical path during the service switchback test phase (such as performance under stress testing) for a long time, providing a data foundation for subsequent analysis.
[0046] Based on fault repair archives and archived performance data, optimization suggestions are generated. This step corresponds to the "Generate Optimization Suggestion" node and its subordinate logic in the attached diagram, such as suggestions to supplement N-1 protection paths. The system performs correlation analysis on archive data and performance data, automatically generating targeted optimization suggestions. For example, if the analysis finds that a certain type of service has a long recovery time due to a lack of protection paths, it is recommended to supplement it with N-1 protection paths; if the resource verification process frequently fails due to insufficient capacity, it proposes to adjust the resource allocation strategy or expand capacity; if a certain optical path repeatedly fails, it provides suggestions to strengthen the quality inspection of the optical path.
[0047] The system updates the current fault case, handling logic, and optimization suggestions to the knowledge base and associates them with similar fault modes: this step corresponds to the "updating the knowledge base and associating similar fault mode nodes" in the attached diagram. The system synchronizes the complete case information of this fault, the successful handling logic, and the generated optimization suggestions to the operations and maintenance knowledge base. Simultaneously, through a pattern matching algorithm, the system associates the characteristics of this fault (such as alarm type and performance degradation mode) with existing historical fault cases in the knowledge base based on similarity, forming a fault mode network. This enables the system to provide decision support for subsequent fault self-healing based on the knowledge base when a similar alarm is triggered again (corresponding to the judgment node in the attached diagram), and automatically recommends handling strategies (corresponding to the branch in the attached diagram), achieving the accumulation and self-optimization of handling experience, and improving the intelligence and efficiency of future fault handling.
[0048] It is understood that those skilled in the art can combine various implementation methods in the above embodiments under the guidance of the above examples to obtain technical solutions with multiple implementation methods.
[0049] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for power communication transmission network fault self-repair processing, characterized in that, The method comprises the following steps: receiving an abnormal event alarm signal and identifying an alarm type; performing fault analysis according to the alarm type, querying circuit service and optical path resource IDs affected by the alarm type in a resource view, and generating a fault detection result within a set time window; starting a self-repairing treatment process based on the fault detection result, extracting attribute information of the affected service to distinguish the priority of the service, and verifying the N-1 redundancy capability of the service; if the verification result requires temporary channel to be enabled, starting temporary circuit arrangement, performing resource verification, and after the verification is passed, migrating service traffic to the temporary circuit or triggering protection channel switching; when the original optical path is repaired, starting a service back-switching process, collecting performance of the repaired original optical path, and testing service traffic; if the test is passed, migrating service traffic back to the original optical path in a dual-channel parallel mode, releasing resources occupied by the temporary channel, and updating network topology and service routing table; after the temporary circuit is released, generating a fault repair archive, analyzing the root cause of the fault, recording the processing process and resource occupation details, archiving performance data to generate optimization suggestions, and updating the knowledge base.
2. The power communication transport network fault self-healing process method according to claim 1, characterized in that, The alarm type includes device and link alarm, optical signal interruption alarm, service layer alarm, optical power abnormal alarm, and opposite end optical port alarm.
3. The power communication transport network fault self-healing process method according to claim 2, characterized in that, The fault analysis according to the alarm type, the query of the affected circuit service and optical path resource ID in the resource view, and the generation of the fault detection result within the set time window comprise: if the alarm type is device and link alarm, combining network element performance data to determine whether the optical path is abnormal; if the alarm type is optical signal interruption alarm, directly determining that the optical path is interrupted; if the alarm type is service layer alarm, analyzing the common points of service routing and performance data; if the alarm type is optical power abnormal alarm, combining real-time port alarm to determine optical port failure; if the alarm type is opposite end optical port alarm, combining performance data to verify the optical port state.
4. The power communications transport network fault self-healing process method of claim 1, wherein, The extraction of attribute information of the affected service to distinguish the priority of the service comprises: extracting attribute, topology level and resource dependency information of the affected service; based on the dynamic association graph of the service and the resource, automatically labeling other potential affected service chains strongly associated with the service to perform chain effect early warning; based on the extracted information, distinguishing the priority and protection level of the service; combining real-time service traffic peak value data to dynamically fine-tune the priority of the service.
5. The power communications transport network fault self-healing process method of claim 4, wherein, The verification of the N-1 redundancy capability of the service comprises: starting redundancy channel health scanning, evaluating historical failure rate and recent performance index fluctuation of the standby channel; if the service is not configured with a protection channel or the protection channel state is abnormal, starting a temporary circuit arrangement process, and triggering fast query of cross-domain resource pool to preferentially call low-load resources; if the service is configured with a protection channel and the channel state is normal, directly triggering automatic switching of the protection channel, and pushing a state snapshot containing performance parameters of the channel at the moment of switching to the operation and maintenance personnel.
6. The power communications transport network fault self-healing process method of claim 1, wherein, The performance collection and service traffic test of the repaired original optical path comprise: automatically triggering optical path performance collection to collect input optical power, output optical power and basic error rate indicators of the original optical path; By associating the historical fault types and repair process records of the original optical path, a digital health profile of the optical path is generated to visually demonstrate the performance change trend before and after the repair. Based on the historical traffic fluctuation patterns of the affected services, dynamic simulated traffic is generated to simulate and test the original optical path.
7. The power communications transport network fault self-healing process method of claim 6, wherein, The process of collecting performance data and testing service traffic on the repaired original optical path also includes: Extreme scenario simulations were incorporated into the simulation test to verify the stress resistance of the original optical path; Determine whether the simulated test passes the test criterion based on business experience similarity; If the test fails, the optical path is marked as potentially not fully repaired, and a suspected root cause analysis report is generated to notify maintenance personnel to intervene and investigate. If the test passes, an optical path repair verification report is generated and the switchback process is triggered.
8. The power communications transport network fault self-healing process method of claim 7, wherein, The method of migrating service traffic back to the original optical path using a dual-channel parallel approach includes: In the back-switching process, the service traffic is transmitted simultaneously on the temporary channel and the original optical path; Real-time monitoring of the load rate and delay jitter of the temporary channel and the original optical path; Based on real-time monitoring data, the algorithm dynamically adjusts the distribution ratio of business traffic on the two channels.
9. The power communications transport network fault self-healing process method of claim 8, wherein, The process of releasing the resources occupied by the temporary channel and updating the network topology and service routing table includes: After all service traffic is migrated back to the original optical path, the port and bandwidth resources occupied by the temporary channel are released. If similar business demand is predicted in the near future, the released resources will be marked as priority reserved resources; otherwise, they will be marked as ordinary idle resources. Update the network topology and service routing table to reflect the actual status after resource release and service rollback.
10. The power communications transport network fault self-healing process method of claim 1, wherein, After the temporary circuit is released, a fault repair file is generated, the root cause of the fault is analyzed, the processing procedure and resource usage details are recorded, performance data is archived to generate optimization suggestions, and the knowledge base is updated, including: Generate a fault repair file that includes a root cause analysis of the fault, timestamps of the processing process, and details of resource usage; Archive performance data during the operation of the temporary channel and the service switchback test phase; Based on the aforementioned fault repair files and archived performance data, optimization suggestions are generated; Update the fault case, handling logic, and optimization suggestions to the knowledge base and associate them with similar fault modes.
Citation Information
Cited By
A database configuration driving-based heterogeneous platform alarm docking and forwarding method
CN122437884A