A method for fault tracing and root cause analysis of a dual-mode communication module
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHERN POWER GRID DIGITAL GRID RESEARCH INSTITUTE CO LTD
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-07
AI Technical Summary
(1)数据采集局限:未针对双模通信模块设计多维度数据采集机制,未覆盖工况、链路质量、拓扑关联、业务传输四类核心数据;
本申请公开的技术方案通过IPv6路由组网与电鸿化适配接口,构建边缘与云协同故障分析架构,实现边缘侧实时采集与云侧深度分析的高效协同;通过多维度数据协同采集机制,覆盖工况、链路质量、拓扑关联、业务传输四类数据,全面捕捉故障特征;通过多源数据时空融合提取核心特征,构建数据、拓扑与业务三维特征集,为根因分析提供精准支撑;通过规则推理与机器学习融合分析、边缘与云双向验证机制、云侧溯源与边缘侧实时验证结合,提升溯源准确性;通过可视化运维支撑体系生成溯源路径图与详细报告,支持远程故障处置指令推送。从而实现以下技术效果:
Smart Images

Figure CN122533918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power technology, and in particular to a method for fault tracing and root cause analysis of a dual-mode communication module based on edge and cloud collaboration. Background Technology
[0002] With the construction of new power systems and the advancement of "dual-carbon" goals, a large number of integrated sensing, computing, and control dual-mode communication modules have been deployed in low-voltage distribution areas. These modules include CCO (Central Coordinator, the core control node of the dual-mode communication network), PCO (Relay Coordinator, the relay node of the dual-mode communication network), and STA (Station, the terminal node of the dual-mode communication network, responsible for data acquisition and uploading), supporting high-speed communication for multiple services such as metering, power distribution, photovoltaics, and charging piles. These modules integrate complex functions such as PLC / RF (Power Line Communication / Radio Frequency Communication) dual-mode communication, IP-based networking, and e-commerce management. They exhibit diverse fault types (link interference, topology anomalies, e-commerce compatibility faults, etc.), with concealed propagation paths, making root cause localization difficult. The e-commerce operating system possesses unified IoT management and data interaction capabilities, providing underlying support for edge and cloud collaborative fault analysis. However, there is an urgent need to address the accuracy and efficiency issues of existing methods.
[0003] The patent document CN117933909A, entitled "A Distributed Cloud-Edge Collaborative Task Management System and Method," discloses the following technical solutions: providing distributed storage services through a cloud platform to reduce network transmission latency and support emergency tasks; the edge-side communication module includes a first line (wired transmission) and a second backup line (wireless transmission), and adopts load prediction technology to improve resource utilization efficiency; and achieving unified management of key data and task allocation through cloud-edge collaboration to ensure rapid response to emergency tasks.
[0004] The above technology has the following shortcomings: (1) Data acquisition limitations: A multi-dimensional data acquisition mechanism was not designed for the dual-mode communication module, and the four core data categories of working conditions, link quality, topology association, and service transmission were not covered; (2) Lack of collaborative analysis: It only focuses on task allocation and resource management, without building a dedicated edge-cloud collaborative architecture for fault analysis, and lacks data fusion and intelligent analysis capabilities; (3) Weak fault handling: There is no fault root cause identification and tracing mechanism, and it can only achieve task-level response, which cannot solve the problem of complex fault location of communication module; (4) Insufficient adaptability: It is not deeply adapted to the Dianhong operating system, does not support Dianhong-style management data interaction of dual-mode communication modules, and cannot be integrated into the unified operation and maintenance system of low-voltage distribution areas; (5) Lack of visualization and operation and maintenance support: The system does not provide visualization of fault tracing paths and remote handling command push functions, resulting in low operation and maintenance efficiency.
[0005] Therefore, there is an urgent need for a collaborative method that combines real-time data acquisition at the edge with intelligent analysis at the cloud to achieve rapid source tracing and accurate root cause analysis of dual-mode communication module failures, thereby improving the operation and maintenance efficiency and stability of low-voltage distribution area communication networks. Summary of the Invention
[0006] The technical problem to be solved by this application is to provide a method for fault tracing and root cause analysis of dual-mode communication modules that can achieve multi-dimensional collaborative data collection and preprocessing, comprehensive capture of fault characteristics, cloud-edge collaborative data fusion and intelligent analysis, accurate identification of fault root causes, automated fault propagation path tracing and initial node positioning, improve tracing efficiency, edge and cloud bidirectional verification and dynamic optimization, visualization report generation and support for remote operation and maintenance, and adapt to multiple business scenarios of source, network, load and storage.
[0007] This application provides a method for fault tracing and root cause analysis of a dual-mode communication module based on edge and cloud collaboration, including the following steps: S10, Edge-side data acquisition and preprocessing: The dual-mode communication module collects its own operating condition data, link quality data, topology association data and service transmission data in real time. It completes data cleaning, format standardization and time synchronization through the Dianhonghua adapter interface, and adds IPv6 address identifiers and high-precision timestamps. S20, Edge and Cloud Collaborative Data Transmission: The edge side adopts a low-latency distributed data transmission algorithm and uploads the pre-processed data to the cloud-side IoT management platform through IPv6 routing networking; the cloud side synchronously calls the historical fault database, the transformer area non-disruptive topology model and the electrical equipment archive data. S30, Multi-source data fusion and feature extraction: The cloud side performs spatiotemporal fusion of edge-uploaded data and cloud-stored data to extract link quality features, topology association features, operating condition anomaly features and service transmission features, and construct a three-dimensional feature set of data, topology and service. S40, Root Cause Analysis: Based on a pre-defined rule base, rule reasoning is performed to match known fault patterns; combined with a trained machine learning model, unknown fault features are classified, and the root cause type and confidence level are output. S50, edge and cloud collaborative fault tracing: The cloud side traces the fault propagation path and initial node based on the root cause analysis results and topology model; the edge side collects real-time data through non-disruptive topology identification technology, verifies the tracing results and feeds them back to the cloud side. S60, Visualization of source tracing results and report generation: The cloud-based network management system generates fault source tracing path diagrams, root cause analysis reports and handling suggestions, and supports fault node location marking and remote maintenance command push.
[0008] According to some embodiments, in step S10, the operating condition data includes module model, firmware version, online rate, operating status, Elec-Tech operating system adaptation status, and API call logs; the link quality data includes the signal-to-interference-plus-noise ratio of the power line communication link, the received signal strength indication of the radio frequency link, packet loss rate, communication delay, link switching frequency, and communication status; the topology association data includes node terminal device identifier, topology affiliation, ranging data, cluster node similarity, and box connection relationship; the service transmission data includes acquisition delay, message retransmission count, service type identifier, IPv6 addressing status, and datagram transport layer security authentication result.
[0009] According to some embodiments, in step S10, the preprocessing includes using a sliding window filter to remove outliers, unifying the format according to the electronic data specification, and achieving edge-side data time alignment based on high-precision clock synchronization technology, wherein the timestamp accuracy is ≤1ms.
[0010] According to some embodiments, in step S20, the collaborative transmission adopts an active reporting and timed query mechanism. Abnormal data of the edge module is reported in real time, and normal data is reported in batches every minute. The cloud side realizes dynamic management of edge node IP addresses through DHCPv6 service to ensure IP addressing and access for data transmission.
[0011] According to some embodiments, in step S30, the link quality characteristics include the fluctuation amplitude of the signal-to-interference-plus-noise ratio, the threshold exceedance duration of the received signal strength indication, and the abnormal coefficient of the link switching frequency; the topology association characteristics include node ranging deviation, abnormal frequency of intra-cluster similarity, and topology change rate; the abnormal operating condition characteristics include the magnitude of the sudden drop in online rate, the number of failed Dianhong API calls, and firmware version incompatibility identifier; the service transmission characteristics include the duration of the acquisition delay exceeding the threshold, the frequency of datagram transport layer security authentication failure, and the gradient change of packet loss rate.
[0012] According to some embodiments, the method also includes feature standardization processing, which normalizes the link quality features, the topology association features, the operating condition anomaly features, and the service transmission features to the [0,1] interval.
[0013] According to some embodiments, in step S40, the fault modes in the preset rule base include power line link interference fault mode, radio frequency signal attenuation fault mode, topology assignment error fault mode, electrical compatibility fault mode, hardware fault mode, protocol incompatibility fault mode, and data transmission timeout fault mode; the machine learning model adopts the random forest algorithm, is trained based on historical fault data, and supports automatic classification of root cause type and confidence output.
[0014] According to some embodiments, step S40 also includes fusion decision, which includes weighted fusion of rule reasoning and machine learning results. When the overall confidence level is ≥85%, the final root cause is output. When the overall confidence level is <85%, the feature weights and rule base are updated and reanalyzed.
[0015] According to some embodiments, step S50 further includes dynamic optimization, which includes pushing the result if the source tracing verification result is successful, and updating the feature weights and rule base on the cloud side if the source tracing verification result fails, and re-executing steps S30 to S50.
[0016] According to some embodiments, in step S60, the root cause analysis report includes the fault type, confidence level, propagation path, and impact on services.
[0017] The beneficial effects of this application are as follows: The technical solution disclosed in this application constructs an edge-cloud collaborative fault analysis architecture through IPv6 routing networking and Dianhonghua adaptation interfaces, achieving efficient collaboration between real-time edge data collection and in-depth cloud analysis. Through a multi-dimensional data collaborative collection mechanism, it covers four types of data: operating conditions, link quality, topology correlation, and service transmission, comprehensively capturing fault characteristics. By extracting core features through spatiotemporal fusion of multi-source data, it constructs a three-dimensional feature set of data, topology, and services, providing accurate support for root cause analysis. Through rule-based reasoning and machine learning fusion analysis, a two-way verification mechanism between the edge and cloud, and a combination of cloud-side tracing and edge-side real-time verification, it improves the accuracy of source tracing. Through a visualized operation and maintenance support system, it generates source tracing path diagrams and detailed reports, supporting remote fault handling command push. Thus, the following technical effects are achieved: (1) High source tracing efficiency: The fault diagnosis cycle is shortened from hours to within 15 minutes, significantly improving operation and maintenance efficiency; (2) Root cause identification accuracy: Overall confidence level ≥85%, covering complex fault modes, and unknown fault identification accuracy ≥80%; (3) Comprehensive data support: Multi-dimensional data is collected collaboratively to avoid the one-sidedness of analysis caused by single data; (4) Strong adaptability: It is deeply adapted to the Dianhong operating system, integrated into the "one network" operation and maintenance system of low-voltage distribution area, and compatible with multiple service scenarios of "source, network, load and storage"; (5) Dynamic optimization: Supports fault feature self-learning and rule base iteration to continuously improve analysis accuracy; (6) Convenient operation and maintenance: Visualized reports and remote handling instructions are pushed to reduce the workload of operation and maintenance personnel on site. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This diagram illustrates the system architecture of an edge-cloud collaborative fault tracing and analysis system for a dual-mode communication module for low-voltage distribution areas, based on an example embodiment. It shows the hierarchical relationship and data flow of the edge side (CCO / PCO / STA module), the collaborative transmission layer (IPv6 routing network), and the cloud-side platform (IoT management platform, network management system).
[0020] Figure 2 The flowchart illustrates the fault tracing and root cause analysis of a dual-mode communication module based on edge and cloud collaboration according to an example embodiment, demonstrating the entire process of data acquisition, preprocessing, transmission, fusion extraction, root cause analysis, tracing verification, and report generation.
[0021] Figure 3 This diagram illustrates the data acquisition dimension structure for fault diagnosis of PLC / RF dual-mode communication modules in low-voltage distribution areas, based on an example embodiment. It shows the specific indicators and acquisition requirements for four types of data: operating conditions, link quality, topology correlation, and service transmission.
[0022] Figure 4 A flowchart illustrating the fusion decision-making process for root cause analysis of a dual-mode communication module failure according to an example embodiment is provided, demonstrating the process of fusion decision-making between rule-based reasoning and machine learning, and indicating the confidence threshold.
[0023] Figure 5 The diagram illustrates the interactive flowchart of edge and cloud collaborative fault tracing and verification according to an example embodiment, demonstrating the closed-loop logic of cloud-side tracing path reasoning, edge-side real-time data verification, result feedback, and dynamic optimization. Detailed Implementation
[0024] The embodiments of this application will now be described in detail with reference to the accompanying drawings. It should be understood that the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0025] Those skilled in the art should understand that the following specific embodiments or implementation methods are a series of optimized configurations listed in this application to further explain the specific application content. These configuration methods can be combined or used in conjunction with each other, unless this application explicitly states that some or a specific embodiment or implementation method cannot be associated with or used in conjunction with other embodiments or implementation methods. Furthermore, the following specific embodiments or implementation methods are only considered as optimized configurations and are not intended to limit the scope of protection of this application.
[0026] Example 1 Figure 1 This diagram illustrates the system architecture of an edge-cloud collaborative fault tracing and analysis system for a dual-mode communication module for low-voltage distribution areas, based on an example embodiment. It shows the hierarchical relationship and data flow of the edge side (CCO / PCO / STA module), the collaborative transmission layer (IPv6 routing network), and the cloud-side platform (IoT management platform, network management system).
[0027] like Figure 1 As shown, the system is divided into three layers: the top layer is the edge perception layer on the edge side, which is responsible for collecting multi-dimensional data such as operating conditions, link quality, topology association, and service transmission through the data acquisition module, and performing filtering, formatting, and high-precision timestamp marking after local preprocessing; the middle layer is the collaborative transmission layer, which is based on IPv6 routing and uses a hybrid transmission mechanism of active reporting and timed query to transmit the preprocessed data to the cloud side with low latency (≤1 minute) and high reliability; the bottom layer is the platform service layer on the cloud side, which, on the one hand, uses the data intelligent analysis module, combined with the rule reasoning engine and random forest machine learning algorithm, to perform fault root cause analysis and bidirectional source tracing verification on the three-dimensional feature set after multi-source fusion; on the other hand, it uses the application service module to realize fault path visualization, analysis report generation, and remote operation and maintenance command push, thus forming a complete fault diagnosis link from edge data collection and collaborative transmission to cloud intelligent analysis and closed-loop optimization, effectively solving the pain points of low efficiency in traditional communication fault diagnosis, inaccurate root cause identification, and reliance on manual operation and maintenance.
[0028] Figure 2 The flowchart illustrates the fault tracing and root cause analysis of a dual-mode communication module based on edge and cloud collaboration according to an example embodiment, demonstrating the entire process of data acquisition, preprocessing, transmission, fusion extraction, root cause analysis, tracing verification, and report generation.
[0029] like Figure 2As shown, a method for fault tracing and root cause analysis of a dual-mode communication module based on edge and cloud collaboration is proposed. Starting with edge-side data acquisition and preprocessing, it simultaneously collects and processes operational, link, topology, and service data. Then, it achieves edge-cloud collaborative data transmission through IPv6 upload and cloud-side data retrieval. Next, it constructs a three-dimensional feature set of data, topology, and services for multi-source data fusion and feature extraction. Subsequently, it combines rule matching and machine learning fusion algorithms to conduct fault root cause analysis. Finally, it performs edge-cloud collaborative fault tracing and verification through cloud-side tracing and edge verification. If verification fails, it returns to the multi-source data fusion and feature extraction stage for iterative re-iteration. After successful verification, it visualizes the tracing results and generates a report, outputting the fault path, report, and handling suggestions. Ultimately, it achieves remote operation and maintenance, forming a closed-loop, optimizable fault diagnosis and operation and maintenance process. Through a six-step core process—edge-side data acquisition and preprocessing, edge-cloud collaborative data transmission, multi-source data fusion and feature extraction, fault root cause analysis, resource allocation effect verification and closed-loop optimization, and root cause visualization and report generation—fault tracing and root cause analysis of the dual-mode communication module are achieved, with deep collaboration with the Elec-Tech operating system throughout. Specific content includes: Step 1, edge-side data acquisition and preprocessing, including: (1) Multi-dimensional data acquisition: The dual-mode communication module (CCO / PCO / STA) acquires four types of data in real time: Operating data: module model, firmware version, online rate, running status, Elec-Tech OS adaptation status, API call logs; Link quality data: SNR (Signal-to-Interference-plus-Noise Ratio) for PLC (Power Line Communication, a communication method that transmits data over power lines), RSSI (Received Signal Strength Indication) for RF (Radio Frequency, a communication method that transmits data over radio waves), packet loss rate, communication delay, link switching frequency, and N-1 / N-2 communication status. Topology association data: Node TEI identifier (Terminal Device Identifier, a unique identifier for nodes in a dual-mode communication network), topology affiliation, ranging data, intra-cluster node similarity, intra-box connectivity; Service transmission data: collection latency, message retransmission count, service type identifier, IPv6 addressing status, DTLS (Datagram Transport Layer Security) authentication result; (2) Preprocessing process: Sliding window filtering (window size 5) is used to remove outliers, and the data is standardized into JSON format according to the electronic data specifications; time alignment is achieved based on high-precision clock synchronization, and the timestamp accuracy is ≤1ms; IPv6 address identifiers are added to ensure data traceability.
[0030] Figure 3This diagram illustrates the data acquisition dimension structure for fault diagnosis of PLC / RF dual-mode communication modules in low-voltage distribution areas, based on an example embodiment. It shows the specific indicators and acquisition requirements for four types of data: operating conditions, link quality, topology correlation, and service transmission.
[0031] like Figure 3 As shown, with "data acquisition dimension" as the core, four major categories of key data systems have been constructed: First, service transmission data, covering IPv6 addressing status, DTLS authentication results, service type identifiers, acquisition latency, and message retransmission counts, reflecting the reliability and security of communication services; second, operating condition data, including API call logs, Elec-Tech OS adaptation status, online rate / operation status, module model, and firmware version, reflecting the health of equipment operation and system compatibility; third, topology association data, involving intra-cluster node similarity, ranging data, node TEI identifiers, and topology affiliation, depicting the physical and logical connection structure of the network; and fourth, link quality data, collecting SNR values and packet loss rates for PLC communication, collecting RSSI values and communication latency for RF communication, while recording link switching frequency and N-1 / N-2 communication status, accurately reflecting the link performance of different communication methods. These four dimensions of multi-source data together constitute the complete data foundation for fault diagnosis, providing comprehensive support for subsequent construction of data, topology, and service three-dimensional feature sets, and realizing fault root cause analysis and accurate source tracing.
[0032] Step 2, edge and cloud collaborative data transmission, including: (1) Transmission mechanism: The mode of active reporting and timed query is adopted. Abnormal data is reported in real time, and normal data is reported in batches every minute; (2) Transmission optimization: The edge side adopts a low-latency distributed data transmission algorithm and uploads to the cloud-side IoT management platform through IPv6 routing network, with a transmission latency of ≤1 minute; (3) Cloud-side data access: Synchronously access the historical fault database (complete records from 2023 to 2025), the uninterrupted topology model of the transformer area (supports 200 nodes, and the mapping time is ≤3 hours), and the Dianhong equipment files (module model, firmware version, certification status, etc.).
[0033] Step 3, multi-source data fusion and feature extraction, includes: (1) Spatiotemporal fusion: Based on real-time data, it is aligned and fused with historical data of the same period in the past 7 days; (2) Feature extraction: Constructing a three-dimensional feature set of data, topology, and business: Link quality characteristics: SNR fluctuation amplitude, RSSI threshold breach duration, and link handover frequency anomaly coefficient; Topological association features: node ranging bias, frequency of intra-cluster similarity anomalies, and rate of change of topological structure; Abnormal operating conditions: sudden drop in online rate, number of failed API calls to Dexton, firmware version incompatibility flag; Service transmission characteristics: duration of data collection delay exceeding the threshold, frequency of DTLS authentication failures, and gradient changes in packet loss rate; (3) Feature standardization: normalize all features to the [0,1] interval to ensure the consistency of analysis.
[0034] Step 4, root cause analysis, including: (1) Rule reasoning engine: Based on the preset rule base, it matches known fault modes. The rule base covers PLC link interference, RF signal attenuation, topology belonging error, electrical adaptation fault, hardware fault, protocol incompatibility, data transmission timeout, etc. (2) Machine learning model: The trained random forest algorithm is used to classify the features of unknown faults and output the root cause type and confidence level; (3) Fusion decision: The results of rule reasoning and machine learning are weighted and fused. When the overall confidence level is ≥85%, the final root cause is output. If it is <85%, the feature weights and rule base are updated and reanalyzed.
[0035] Figure 4 A flowchart illustrating the fusion decision-making process for root cause analysis of a dual-mode communication module failure according to an example embodiment is provided, demonstrating the process of fusion decision-making between rule-based reasoning and machine learning, and indicating the confidence threshold.
[0036] like Figure 4 As shown, a self-learning, closed-loop optimization fault diagnosis mechanism is constructed by combining a preset rule base with a machine learning model. The process first matches the fault data with the rule base, which includes known fault modes such as PLC link interference, RF signal attenuation, topology attribution errors, electrical compatibility faults, and hardware faults / protocol incompatibility. If a match is successful, the root cause type and initial confidence level are directly output. If no known fault mode is matched, the data is fed into a random forest machine learning model trained on historical fault data and supporting self-learning of unknown faults for classification, outputting the root cause type and model confidence level. The results from both paths enter the fusion decision unit. When the overall confidence level is ≥85%, the final root cause result is output. If the confidence level is <85%, the feature weights and rule base are updated and re-input for analysis, forming a closed loop of continuous iterative optimization. This balances the rapid location of known faults with the ability to identify unknown faults, improving the accuracy and adaptability of fault diagnosis.
[0037] Step 5, resource allocation effect verification and closed-loop optimization, including: (1) Cloud-based source tracing: Based on the root cause analysis results and topology model, trace the fault propagation path and initial node; (2) Edge-side verification: Real-time topology data and link quality data are collected at the edge using non-disruptive topology identification technology to verify the tracing results; (3) Dynamic optimization: If the verification is successful, the result is pushed; if it fails, the cloud side updates the feature weights and rule base, and steps 3 to 5 are executed again.
[0038] Figure 5 The diagram illustrates the interactive flowchart of edge and cloud collaborative fault tracing and verification according to an example embodiment, demonstrating the closed-loop logic of cloud-side tracing path reasoning, edge-side real-time data verification, result feedback, and dynamic optimization.
[0039] like Figure 5 As shown, the collaborative logic among the four roles—cloud-side engine, topology model, edge side, and network management system—is presented: First, the cloud-side engine requests the fault propagation path from the topology model. After the topology model returns the fault loop and initial node, the cloud-side engine issues a verification command to the edge side. The edge side collects real-time data and feeds back the verification results. If the verification is successful, the cloud-side engine pushes the final source tracing result to the network management system, which then generates a fault report and handling suggestions. If the verification fails, the cloud-side engine optimizes the feature weights and rule base, and triggers the fault analysis process to be re-executed, forming a closed-loop optimization mechanism to ensure the accuracy and verifiability of the fault source tracing results.
[0040] Step 6, source result visualization and report generation, including: (1) Visualization: The cloud-side network management system generates a fault tracing path map and a fault heat map of the transformer area, marking the location of fault nodes and the scope of impact; (2) Report generation: Output a root cause analysis report, including fault type, confidence level, propagation path, and impact on business; (3) Operation and maintenance support: push processing suggestions (remote upgrade instructions, topology adjustment schemes, link optimization parameters), and support the issuance of remote maintenance instructions.
[0041] Example 2 In a low-voltage distribution area hardware environment consisting of 1 CCO central coordination module, 20 PCO agent coordination modules, and 480 STA terminal modules, all equipped with the Elec-Tech Operating System V2.0, the technical solution of this application is fully implemented: Implementation conditions: (1) Hardware configuration: a dual-mode communication module cluster integrating sensing, computing and control (1 CCO, 20 PCOs, 480 STAs), all compatible with the Dianhong operating system V2.0; (2) Data parameters: PLC link SNR≥100dB, RF link RSSI≥-70dBm, data transmission delay≤1 minute, timestamp accuracy≤1ms; (3) Software support: Dianhong system hardware abstraction layer, low latency distributed transmission algorithm, random forest classification model, transformer area non-disruptive topology model, cloud-side IoT management platform.
[0042] Implementation steps: (1) Edge-side data acquisition and preprocessing: The CCO module collects network operating conditions, link and topology data every 1 minute, and the STA module reports abnormal data in real time; SNR abnormal values are removed by sliding window filtering, and the unified format is JSON, with IPv6 identifier and high-precision timestamp added; (2) Edge and cloud collaborative transmission: The edge side uploads data through IPv6 routing network, and the cloud side calls the historical fault database, topology model and electro-hydraulic equipment file; (3) Multi-source data fusion and feature extraction: Real-time data is fused with historical data from the past 7 days to extract features such as SNR fluctuation amplitude and topological similarity anomaly frequency, and a three-dimensional feature set is constructed after normalization; (4) Root cause analysis: A certain STA module has a data acquisition delay of 150 seconds, SNR=85dB, and link switching 6 times / minute. The rule matching is "PLC link interference" (confidence level 82%). The machine learning verification output is the same (confidence level 89%). The fusion confidence level is 86%, and the root cause is determined. (5) Collaborative tracing and verification: The cloud side traces back to the STA module with TEI=105 (range deviation of 35 meters) based on the topology model, and the edge side verifies that the topology assignment is incorrect by non-disruptive topology identification, and the verification is successful; (6) Visualized Reporting and Handling: The cloud side generates a traceability path map and report, pushes the instruction "adjust cluster affiliation and upgrade firmware V2.0", and completes the handling within 15 minutes, restoring the module collection latency to 45 seconds.
[0043] Implementation results: (1) The troubleshooting cycle is ≤15 minutes, which is 80% shorter than the traditional method; (2) Root cause identification confidence level ≥86%, and unknown fault identification accuracy ≥80%; (3) Supports concurrent fault monitoring of 500 nodes, adapting to multiple service scenarios of "source, network, load and storage"; (4) After processing, the module acquisition delay is ≤60 seconds and the link quality is restored to normal (SNR≥100dB).
[0044] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for fault tracing and root cause analysis of a dual-mode communication module, characterized in that, Includes the following steps: S10, Edge-side data acquisition and preprocessing: The dual-mode communication module collects its own operating condition data, link quality data, topology association data and service transmission data in real time. It completes data cleaning, format standardization and time synchronization through the Dianhonghua adapter interface, and adds IPv6 address identifiers and high-precision timestamps. S20, Edge and Cloud Collaborative Data Transmission: The edge side adopts a low-latency distributed data transmission algorithm and uploads the pre-processed data to the cloud-side IoT management platform through IPv6 routing networking; the cloud side synchronously calls the historical fault database, the transformer area non-disruptive topology model and the electrical equipment archive data. S30, Multi-source data fusion and feature extraction: The cloud side performs spatiotemporal fusion of edge-uploaded data and cloud-stored data to extract link quality features, topology association features, operating condition anomaly features and service transmission features, and construct a three-dimensional feature set of data, topology and service. S40, Root Cause Analysis: Based on a pre-defined rule base, rule reasoning is performed to match known fault patterns; combined with a trained machine learning model, unknown fault features are classified, and the root cause type and confidence level are output. S50, edge and cloud collaborative fault tracing: The cloud side traces the fault propagation path and initial node based on the root cause analysis results and topology model; the edge side collects real-time data through non-disruptive topology identification technology, verifies the tracing results and feeds them back to the cloud side. S60, Visualization of source tracing results and report generation: The cloud-based network management system generates fault source tracing path diagrams, root cause analysis reports and handling suggestions, and supports fault node location marking and remote maintenance command push.
2. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 1, characterized in that, In step S10, the operating condition data includes module model, firmware version, online rate, operating status, Elec-Tech operating system adaptation status, and API call logs. The link quality data includes the signal-to-interference-plus-noise ratio of the power line communication link, the received signal strength indication of the radio frequency link, packet loss rate, communication delay, link switching frequency, and communication status. The topology association data includes node terminal device identifiers, topology affiliation relationships, ranging data, intra-cluster node similarity, and intra-box connectivity relationships. The service transmission data includes collection delay, message retransmission count, service type identifier, IPv6 addressing status, and datagram transport layer security authentication result.
3. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 2, characterized in that, In step S10, the preprocessing includes using a sliding window filter to remove outliers, unifying the format according to the electronic data standard, and achieving edge-side data time alignment based on high-precision clock synchronization technology, wherein the timestamp accuracy is ≤1ms.
4. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 1, characterized in that, In step S20, the collaborative transmission adopts an active reporting and timed query mechanism. Abnormal data from the edge module is reported in real time, and normal data is reported in batches every minute. The cloud side uses DHCPv6 service to realize dynamic management of edge node IP addresses, ensuring IP-based addressing and access for data transmission.
5. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 1, characterized in that, In step S30, the link quality characteristics include the fluctuation amplitude of the signal-to-interference-plus-noise ratio, the threshold breach duration of the received signal strength indication, and the abnormal coefficient of the link switching frequency. The topological association features include node ranging bias, frequency of intra-cluster similarity anomalies, and topological structure change rate. The abnormal operating conditions include the magnitude of the drop in online rate, the number of failed Dianhong API calls, and firmware version incompatibility indicators. The service transmission characteristics include the duration of collection delay exceeding the threshold, the frequency of datagram transport layer security authentication failures, and the gradient change in packet loss rate.
6. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 5, characterized in that, Step S30 also includes feature standardization processing, which normalizes the link quality features, the topology association features, the operating condition anomaly features, and the service transmission features to the [0,1] interval.
7. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 1, characterized in that, In step S40, the fault modes in the preset rule base include power line link interference fault mode, radio frequency signal attenuation fault mode, topology assignment error fault mode, power line adaptation fault mode, hardware fault mode, protocol incompatibility fault mode, and data transmission timeout fault mode. The machine learning model uses the random forest algorithm, is trained based on historical fault data, and supports automatic classification of root cause types and output of confidence scores.
8. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 7, characterized in that, Step S40 also includes fusion decision, which includes weighted fusion of rule reasoning and machine learning results. When the overall confidence level is ≥85%, the final root cause is output. When the overall confidence level is <85%, the feature weights and rule base are updated and reanalyzed.
9. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 1, characterized in that, Step S50 also includes dynamic optimization, which includes pushing the result if the source tracing verification is successful, and updating the feature weights and rule base on the cloud side if the source tracing verification fails, and re-executing steps S30 to S50.
10. The method for fault tracing and root cause analysis of a dual-mode communication module according to claim 1, characterized in that, In step S60, the root cause analysis report includes the fault type, confidence level, propagation path, and impact on services.
Citation Information
Patent Citations
Distributed cloud edge collaborative task management system and method
CN117933909A