A method and system for fault root cause localization based on automated analysis
By identifying abnormal ports in the fiber optic network, generating a standardized status field mapping table and a global identity identifier, the problem of alarm attribution confusion caused by differences in optical network terminal report formats is solved, and accurate location and automated analysis of fiber optic network faults are achieved.
Patent Information
- Application Number
- CN202511565438.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing technologies in fiber optic networks suffer from inconsistent alarm information attribution due to differences in reporting formats of optical network terminals from different manufacturers. This makes it difficult to accurately associate faults with optical network terminals, thus affecting fault location efficiency.
By identifying abnormal ports in passive optical networks, a standardized status field mapping table is generated. Clustering algorithms are used to match device feature vectors, generating globally unique identifiers. The time series of address conflicts are parsed to accurately associate alarm information with specific optical network terminals.
It enables precise location of the root cause of the fault, reduces manual intervention and maintenance time, improves the automation and accuracy of fault location, and solves the problem of confused alarm attribution caused by duplicate device names or logical address conflicts.
Smart Images

Figure CN121037200B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, and in particular to a fault root cause positioning method and system based on automated analysis. BACKGROUND
[0002] With the wide application of fiber networks, accurately positioning the fault root cause in the operation and maintenance of communication networks has become a key field to ensure network service quality and user network experience. Among them, the popularity of fiber-to-the-home technology in fiber networks makes the optical network terminal used for fiber access, i.e., the optical network terminal, an important role in the fiber network, so timely identification of abnormal optical network terminals in the fiber network can effectively improve the reliability of the fiber network.
[0003] When the prior art identifies abnormal optical network terminals, it often relies on standardized device status reports to identify the status of each optical network terminal as abnormal, but the above identification method has significant limitations when dealing with differences between multiple vendors' devices. Among them, the report formats of optical network terminals from different vendors often differ greatly, and using a unified standardized device status report for abnormal identification can easily cause confusion in the ownership of alarm information, which in turn makes it difficult to accurately associate faults with optical network terminals when processing alarm information. Under the same passive optical network port, optical network terminals from different vendors may use different status report formats, such as the use of non-standard encoding for the status field of some devices, or the absence of key parameters, which makes it difficult to handle reports of different formats uniformly. Especially when there are optical network terminals with the same name or duplicate IP addresses in the fiber network, the blurring of device identity identification makes it difficult to accurately determine the ownership of the alarm, which in turn affects the rapid positioning of the faulty optical network terminal. For example, in an actual scenario, the operation and maintenance personnel found that two optical network terminals under a certain passive optical network port were in conflict due to IP address conflict, and mistakenly attributed the alarm information of one device to another device, resulting in a time-consuming fault location of several hours, which seriously affected the operation and maintenance efficiency. SUMMARY
[0004] The present application discloses a fault root cause positioning method and system based on automated analysis, which is used to improve the accuracy of fault positioning.
[0005] In order to achieve the above purpose, the present application discloses a fault root cause positioning method based on automated analysis, comprising:
[0006] identifying an abnormal port according to the port state data of each port in the passive optical network, and extracting a fault port state data set and a plurality of optical network terminals according to the fault correlation of each port and the abnormal port;
[0007] According to the fault port state data set, field characteristics of each device manufacturer are identified, a standardized state field mapping table is constructed in combination with a preset field mapping configuration, and device state data of each optical network terminal is converted into standard state data;
[0008] According to the standard state data and a preset clustering algorithm, a device feature vector of each optical network terminal is generated, and a device name of the device feature vector is matched;
[0009] According to the device name, address port information of each optical network terminal is obtained, and an identity of each optical network terminal is generated;
[0010] According to the identity, a logical address of each optical network terminal and a conflict logical address in the logical address are obtained, and an allocation time sequence data set of each conflict logical address is generated by querying a preset address mapping record;
[0011] According to an alarm logical address of network alarm information, the logical address and the allocation time sequence data set, a fault optical network terminal is determined from a plurality of optical network terminals.
[0012] The application discloses a fault root cause positioning method based on automatic analysis. First, the abnormal state of the port is identified to only position the optical network terminal connected with the abnormal port, which greatly reduces the workload of abnormal port device positioning and improves the efficiency of fault positioning. Then, a standardized state field mapping table is constructed according to the port state data set, which unifies the formats of heterogeneous device state data of multiple manufacturers, eliminates the parsing errors caused by field differences, and lays a foundation for subsequent accurate analysis. Secondly, a globally unique device identity is generated and the time sequence of address conflicts is analyzed, which completely solves the problem of alarm attribution confusion caused by device name conflicts or logical address conflicts, realizes the accurate association of alarm information to specific optical network terminals, and significantly improves the automation degree and accuracy of fault root cause positioning, reduces manual intervention and operation time.
[0013] As a preferred example, the abnormal port is identified according to the port state data of each port in the passive optical network, and a plurality of optical network terminals are extracted according to the fault correlation of each port with the abnormal port, including:
[0014] According to a preset network management protocol, binary port state data corresponding to each port in the passive optical network is collected;
[0015] parsing the binary port status data according to a preset data format specification to convert the binary port status data into structured port status data; wherein the structured port status data comprises optical power values of the ports, device coding information of the ports, and collection time stamps;
[0016] obtaining a comparison result of the optical power values and a preset optical power threshold value to identify abnormal ports from the plurality of ports according to the comparison result;
[0017] obtaining a historical optical power value sequence of the abnormal ports within a preset collection period, and obtaining an optical power decay rate by dividing a power difference value of adjacent time points by a time interval according to the historical optical power value sequence;
[0018] when the optical power decay rate is greater than a preset rate threshold value, calculating a Pearson correlation coefficient of each of the ports other than the abnormal port with the optical power decay rate according to the structured port status data;
[0019] dividing each of the ports other than the abnormal port according to the Pearson correlation coefficient and a preset coefficient threshold value to obtain a plurality of fault-related ports;
[0020] obtaining a fault port group according to the plurality of fault-related ports and the abnormal port;
[0021] obtaining a fault port status data set corresponding to the fault port group and a plurality of optical network terminals associated therewith according to the structured port status data of the abnormal port and each of the fault-related ports;
[0022] In the above scheme, the optical power decay rate and the Pearson correlation coefficient of other ports are calculated to perform port grouping, which not only identifies obvious abnormal ports, but also finds related ports with performance degradation due to associated faults, thereby achieving earlier and more comprehensive identification of potential faults and associated faults, expanding the range of fault detection and improving the accuracy of fault location.
[0023] As a preferred example, the standardized status field mapping table is constructed according to field characteristics of each device manufacturer identified from the fault port status data set and in combination with a preset field mapping configuration to convert device status data of each optical network terminal into standard status data, comprising:
[0024] reading an original report structure of the fault port status data set to identify a manufacturer identifier corresponding to each device manufacturer, and a field name and a data format corresponding to the manufacturer identifier;
[0025] According to the field name and the data format, a state code, a field type and a numerical content corresponding to each of the device manufacturers are extracted by matching a predefined state field feature mode;
[0026] A field mapping configuration corresponding to the manufacturer identification and the field type is called to construct a mapping relationship from the field name to a standard field corresponding to each of the device manufacturers; when the field type matches a debugging information feature, a redundancy mark is set for the field type and a service field of the field type and a corresponding mapping rule are determined;
[0027] The state code is converted according to the service field, the mapping rule and the mapping relationship, and the order and format of each field are reorganized and a debugging information field with the redundancy mark is removed according to a predefined standard data structure definition, to obtain a standardized state field mapping table; the standardized state field mapping table includes a unified field name and a unified data format;
[0028] A device report of each of the optical network terminals is obtained, and a plurality of the device reports are batch-converted according to the standardized state field mapping table, to obtain a standard state data corresponding to each of the optical network terminals.
[0029] The above scheme defines the identification of multi-manufacturer field features, mapping rules and the removal process of redundant debugging information in detail, not only realizes efficient and accurate processing of non-standard codes and redundant fields, ensures that the converted standard data only contains high-quality and high-value service information, greatly improves the accuracy and processing efficiency of subsequent device state data analysis, directly solves the technical problem that different state report formats cannot be uniformly processed, and then uses a report in a unified format for fault positioning, improving the accuracy of positioning.
[0030] As a preferred example, the generation of a device feature vector of each of the optical network terminals according to the standard state data and a preset clustering algorithm, and the matching of a device name of the device feature vector, includes:
[0031] A plurality of optical power values, a plurality of bit error rate values and a plurality of collection time points of each of the optical network terminals are read according to the standard state data, to obtain an optical power mean value, a bit error rate mean value, an optical power variance and a bit error rate variance according to the optical power values, the bit error rate values and a collection time length of the plurality of collection time points;
[0032] A multi-dimensional feature data of each of the optical network terminals is constructed according to the optical power mean value, the optical power variance, the bit error rate mean value and the bit error rate variance;
[0033] According to the multi-dimensional feature data and a preset clustering algorithm, a device feature vector set marked with clustering attribution is obtained;
[0034] Each cluster in the device feature vector set is acquired, and the maximum optical power and the minimum optical power of all the optical network terminals in the cluster are counted to form an optical power range and calculate a bit error rate distribution interval of the bit error rate value;
[0035] According to the optical power range, the bit error rate distribution interval and a preset device performance specification table, the device performance corresponding to each cluster is determined, and the device performance is taken as a preliminary identity of the optical network terminal;
[0036] According to the preliminary identity of the device, preset device configuration data is searched to obtain a device name corresponding to the optical network terminal.
[0037] In the above scheme, the device running features such as the mean and variance of the optical power and the bit error rate of the device are statistically analyzed by the clustering algorithm to match the device name, so that the device type and identity can be intelligently inferred through the behavior characteristics even if the device report cannot directly provide an accurate name or the name is wrong, thereby enhancing the reliability of fault positioning in the information missing or error scenario and improving the accuracy of fault positioning.
[0038] As a preferred example, the address port information of each optical network terminal is acquired according to the device name to generate the identity of each optical network terminal, including:
[0039] The occurrence frequency of each device name is counted, and when the occurrence frequency is greater than a preset frequency threshold, the optical network terminals corresponding to the device name are determined as the name conflict device;
[0040] The device name is taken as the identity of each optical network terminal that is not determined as the name conflict device, and the hardware address and the port number value corresponding to each name conflict device are acquired according to the standard state data to extract the manufacturer identification code and the device serial number in the hardware address;
[0041] The manufacturer identification code, the device serial number and the port number value are combined according to a preset format to obtain a suffix identification, and the suffix identification is added to the device name to obtain the identity of each name conflict device.
[0042] The scheme creates a suffix identification of the fusion vendor identification, the device serial number and the port number for the device with the same name, so that each optical network terminal in the network has a globally unique and easily identifiable identity, which fundamentally eliminates the identity uncertainty caused by the same name, and provides a unique basis for subsequent accurate alarm attribution and fault location.
[0043] As a preferred example, the logical address of each optical network terminal is obtained according to the identity identification, and the allocation time sequence data set of each conflict logical address is generated by querying the preset address mapping record, including:
[0044] According to the identity identification, the logical address of each optical network terminal is obtained, and the occurrence frequency of each logical address is counted, so that when the occurrence frequency of the logical address is greater than a preset frequency threshold, the logical address is determined as a conflict logical address;
[0045] A plurality of conflict devices using the conflict logical address are determined, a plurality of mapping records of the conflict logical address to the hardware address are obtained by querying a preset address resolution protocol, and the creation timestamp and the update timestamp of each mapping record are extracted to obtain an address mapping data set containing the conflict logical address, the hardware address and the timestamp;
[0046] According to the timestamp in the address mapping data set, the plurality of mapping records are arranged in ascending order, so that the address attribution determination result corresponding to the conflict logical address is determined by comparing the time when each hardware address first establishes a mapping with the conflict logical address;
[0047] According to the address attribution determination result, the time point when each conflict device obtains and loses the conflict logical address is extracted from the plurality of mapping records;
[0048] According to the time point, the allocation event of the conflict logical address of each conflict device is time-sequenced to generate an allocation time sequence data set corresponding to the conflict logical address; wherein the allocation time sequence data set includes the time point when each conflict device obtains and loses the conflict logical address and the hardware address of each conflict device.
[0049] In the above scheme, the detailed time sequence data set of the logical address conflict is generated by querying the historical changes of the address mapping record, so that it can be clearly traced back which device occupies a conflict logical address at different time points from the time dimension, thereby accurately matching the alarm occurrence time with the occupied device of the logical address, successfully solving the technical problem of alarm misattribution caused by logical address conflict, and improving the accuracy of fault location.
[0050] As a preferred example, the alarm logic address according to the network alarm information, the logic address, and the allocation time sequence data set determine a faulty optical network terminal from a plurality of the optical network terminals, comprising:
[0051] Extracting an alarm occurrence timestamp, an alarm source logic address, and an alarm type field of the network alarm information;
[0052] According to the logic address of the optical network terminal not determined as the conflict device, the conflict logic address, and the allocation time sequence data set, querying the corresponding faulty optical network terminal of the network alarm information in combination with a preset event window centered on the alarm occurrence timestamp.
[0053] The above scheme uses a time window to match alarm events and device events, which can accurately associate abstract alarm events and specific device state change events on a timeline, thereby reliably locking the real faulty device causing the alarm, reducing the false positive and false negative rates, and further improving the accuracy of fault positioning.
[0054] As a preferred example, the address port information of each optical network terminal is obtained according to the device name to generate an identity of each optical network terminal, further comprising:
[0055] When a plurality of the devices with the same name use the same hardware address and the same port number value, the historical port number value and the historical hardware address of each of the devices with the same name are extracted in timestamp order to obtain connection trajectory data corresponding to each of the devices with the same name; wherein the connection trajectory data includes timestamp data, historical port number value, and historical hardware address;
[0056] The continuous timestamp data is divided into a plurality of time windows to obtain migration path records of each of the devices with the same name in each of the time windows; according to a port number change sequence and a hardware address replacement sequence in the migration path records, a change event chain is formed, and optical power values and bit error rate values before and after each change event are extracted as operating characteristics of the devices with the same name;
[0057] The timestamp data and the hardware address replacement sequence are associated with the operating characteristics to obtain a device identity traceability chain corresponding to each of the devices with the same name, respectively; a power change sequence is obtained by subtracting optical power values at adjacent time points in the device identity traceability chain, and a Pearson correlation coefficient of the power change sequence and a preset standard power change sequence is obtained to confirm a real hardware address of each of the devices with the same name according to the Pearson correlation coefficient;
[0058] The Hungarian algorithm is used to perform bipartite graph optimal matching between the set of real hardware addresses and the set of currently recorded port numbers, to generate a new address-port mapping relationship table;
[0059] According to the new address-port mapping relationship table, the hardware address and the port number value corresponding to each of the devices with the same name are obtained.
[0060] In the above scheme, the real hardware address is traced by analyzing the historical connection trajectory and running characteristic change of the device with the same name, so that even in the extremely complex scene where the hardware address and the port number are maliciously tampered with or repeated, the real identity of the device can be distinguished by analyzing its historical behavior characteristics, greatly enhancing the ability to resist abnormalities and data fraud, improving the accuracy of identity recognition, and further improving the accuracy of fault location.
[0061] As a preferred example, the alarm logic address according to the network alarm information, the logical address and the allocation time sequence data set determine the faulty optical network terminal from a plurality of optical network terminals, further comprising:
[0062] Obtain the optical power value sequence and the bit error rate value sequence before and after the alarm occurrence timestamp from the standard state data of the faulty optical network terminal, and form comprehensive time sequence data containing alarm events and device parameter changes through timestamp alignment;
[0063] According to the comprehensive time sequence data, the difference between the average optical power in a preset time period before the alarm occurrence timestamp and the average optical power in a preset time period after the alarm occurrence timestamp is calculated as the optical power change amplitude;
[0064] According to the comprehensive time sequence data, the difference between the average bit error rate in a preset time period before the alarm occurrence timestamp and the average bit error rate in a preset time period after the alarm occurrence timestamp is calculated as the bit error rate change amplitude;
[0065] When the optical power change amplitude or the bit error rate change amplitude exceeds a preset amplitude threshold, the time sequence correlation mode of the network alarm information and the optical power change amplitude or the bit error rate change amplitude is obtained; wherein, the time sequence correlation mode includes the change event of optical power and the change time of bit error rate;
[0066] By comparing the change time with the alarm time, the type of the fault root cause in the faulty optical network terminal is obtained; wherein, the type includes optical power anomaly, bit error rate exceeding standard and passive optical network topology configuration error.
[0067] In the above scheme, the type of root cause is determined by analyzing the statistical variation amplitude of optical power and bit error rate before and after the alarm, which not only locates the fault equipment, but also further automatically diagnoses the root cause type of the fault, realizes the leap from locating where the fault to diagnosing why the fault, provides a direct repair direction for the operation and maintenance personnel, and greatly improves the operation and maintenance efficiency.
[0068] In another aspect, the application discloses a fault root cause positioning system based on automatic analysis, comprising a device positioning module, a data conversion module, a feature extraction module, an identification matching module, an address conflict module and a fault positioning module.
[0069] The device positioning module is used to identify an abnormal port according to port state data of each port in a passive optical network, and extract a fault port state data set and a plurality of optical network terminals according to the fault correlation of each port and the abnormal port.
[0070] The data conversion module is used to identify the field characteristics of each device manufacturer according to the fault port state data set, and construct a standardized state field mapping table combined with a preset field mapping configuration, so as to convert the device state data of each optical network terminal into standard state data.
[0071] The feature extraction module is used to generate a device feature vector of each optical network terminal according to the standard state data and a preset clustering algorithm, and match the device name of the device feature vector.
[0072] The identification matching module is used to obtain address port information of each optical network terminal according to the device name, so as to generate an identity of each optical network terminal.
[0073] The address conflict module is used to obtain a logical address of each optical network terminal and a conflict logical address in the logical address according to the identity, and generate an allocation time sequence data set of each conflict logical address by querying a preset address mapping record.
[0074] The fault positioning module is used to determine a fault optical network terminal from a plurality of optical network terminals according to an alarm logical address of network alarm information, the logical address and the allocation time sequence data set.
[0075] The application discloses a fault root cause positioning system based on automatic analysis, which firstly identifies the abnormal state of a port to perform abnormal positioning on an optical network terminal connected with the abnormal port only, greatly reduces the workload of abnormal port device positioning, and improves the efficiency of fault positioning. BRIEF DESCRIPTION OF DRAWINGS
[0076] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following embodiments are only some of the embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0077] Figure 1 is a flowchart of a fault root cause positioning method based on automatic analysis disclosed by an embodiment of the present application;
[0078] Figure 2 is a structural schematic diagram of a fault root cause positioning system based on automatic analysis disclosed by an embodiment of the present application;
[0079] Figure 3 is a flowchart of a fault root cause positioning method based on automatic analysis disclosed by another embodiment of the present application;
[0080] Figure 4 is a construction flowchart of a device logical address allocation time sequence data set disclosed by another embodiment of the present application. DETAILED DESCRIPTION
[0081] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the scope of protection of the present application.
[0082] Embodiment one
[0083] Reference Figure 1To solve the technical problem that the prior art cannot accurately and quickly locate the abnormal optical network terminal in the optical fiber network, the embodiment provides a fault root cause positioning method based on automatic analysis, which comprises the following steps:
[0084] Step 101: identifying an abnormal port according to the port state data of each port in the passive optical network, and extracting a fault port state data set and a plurality of optical network terminals according to the fault correlation of each port and the abnormal port.
[0085] In the embodiment, the step mainly comprises the following steps: collecting binary port state data corresponding to each port in the passive optical network according to a preset network management protocol; analyzing the binary port state data according to a preset data format specification to convert the binary port state data into structured port state data; wherein the structured port state data comprises an optical power value of the port, device coding information of the port and a collection time stamp; obtaining a comparison result of the optical power value and a preset optical power threshold to identify an abnormal port from a plurality of ports according to the comparison result; obtaining a historical optical power value sequence of the abnormal port within a preset collection period, and obtaining an optical power decay rate obtained by dividing a power difference value of adjacent time points by a time interval according to the historical optical power value sequence; when the optical power decay rate is greater than a preset rate threshold, calculating a Pearson correlation coefficient of each of the ports other than the abnormal port and the optical power decay rate according to the structured port state data; dividing the ports other than the abnormal port according to the Pearson correlation coefficient and a preset coefficient threshold to obtain a plurality of fault-related ports; obtaining a fault port group according to a plurality of the fault-related ports and the abnormal port; obtaining a fault port state data set corresponding to the fault port group and a plurality of optical network terminals associated therewith according to the structured port state data of the abnormal port and each of the fault-related ports.
[0086] In the embodiment, the above-mentioned step performs port grouping by calculating the optical power decay rate and the Pearson correlation coefficient thereof and other ports, which not only identifies the obvious abnormal port, but also finds the related port with performance degradation due to the correlation fault, realizes earlier and more comprehensive identification of potential faults and correlation faults, thereby expanding the range of fault detection and improving the accuracy of fault location.
[0087] Step 102: identifying the field characteristics of each device manufacturer according to the fault port state data set and constructing a standardized state field mapping table in combination with a preset field mapping configuration to convert the device state data of each optical network terminal into standard state data.
[0088] In the embodiment, the step mainly includes: reading the original report structure of the fault port state data set to identify the manufacturer identifier corresponding to each device manufacturer, and the field name and data format corresponding to the manufacturer identifier; according to the field name and the data format, extracting the state code, field type and numerical content corresponding to each device manufacturer by matching the pre-defined state field feature mode; calling the field mapping configuration corresponding to the manufacturer identifier and the field type to construct the mapping relationship of the field name to the standard field corresponding to each device manufacturer; when the field type matches the debugging information feature, setting a redundancy mark for the field type and determining the business field of the field type and the mapping rule corresponding thereto; converting the state code according to the business field, the mapping rule and the mapping relationship, and reorganizing the order and format of each field according to the pre-set standard data structure definition and eliminating the debugging information field with the redundancy mark to obtain a standardized state field mapping table; wherein the standardized state field mapping table includes a unified field name and a unified data format; obtaining the device report of each optical network terminal, and batch converting multiple device reports according to the standardized state field mapping table to obtain the standard state data corresponding to each optical network terminal.
[0089] In the embodiment, the above step defines the identification of multi-manufacturer field features, mapping rules and the elimination process of redundant debugging information in detail, which not only realizes efficient and accurate processing of non-standard codes and redundant fields, ensures that the converted standard data only contains high-quality and high-value business information, greatly improves the accuracy and processing efficiency of subsequent device state data analysis, directly solves the technical problem that different state report formats cannot be uniformly processed, and then uses the report in a unified format for fault positioning to improve the accuracy of positioning.
[0090] Step 103: generating a device feature vector of each optical network terminal according to the standard state data and a pre-set clustering algorithm, and matching the device name of the device feature vector.
[0091] In the embodiment, the step mainly includes: reading multiple optical power values, multiple bit error rate values and multiple collection time points of each optical network terminal according to the standard state data, obtaining optical power mean value, bit error rate mean value, optical power variance and bit error rate variance according to the optical power values, the bit error rate values and the collection time length of the multiple collection time points; constructing multiple-dimensional feature data of each optical network terminal according to the optical power mean value, the optical power variance, the bit error rate mean value and the bit error rate variance; obtaining a device feature vector set marked with cluster attribution according to the multiple-dimensional feature data and a preset clustering algorithm; obtaining each cluster in the device feature vector set and counting the maximum optical power value and the minimum optical power value of all the optical network terminals in the cluster to form an optical power range and calculate a bit error rate distribution interval of the bit error rate values; determining the device performance corresponding to each cluster according to the optical power range, the bit error rate distribution interval and a preset device performance specification table, and taking the device performance as a device preliminary identity of the optical network terminal; and searching a preset device configuration data according to the device preliminary identity to obtain a device name corresponding to the optical network terminal.
[0092] In the embodiment, the above step matches the device name by statistical analysis of the device running features such as the mean value and variance of the optical power and the bit error rate of the device through the clustering algorithm, so that the device type and identity can be intelligently inferred through the behavior characteristics even if the device report cannot directly provide an accurate name or the name is wrong, thereby enhancing the reliability of fault positioning in the information missing or error scenario and improving the accuracy of fault positioning.
[0093] Step 104: obtaining address port information of each optical network terminal according to the device name to generate an identity of each optical network terminal.
[0094] In the embodiment, the step mainly includes: counting the occurrence times of each device name, determining multiple optical network terminals corresponding to the device name as a name conflict device when the occurrence times are greater than a preset number threshold; taking the device name as a device identity of each optical network terminal that is not determined as the name conflict device, and obtaining a hardware address and a port number value corresponding to each name conflict device according to the standard state data, to extract a manufacturer identification code and a device serial number in the hardware address; combining the manufacturer identification code, the device serial number and the port number value according to a preset format to obtain a suffix identifier, and adding the suffix identifier after the device name to obtain a device identity of each name conflict device.
[0095] In this embodiment, the step further comprises: when multiple said devices with the same name use the same said hardware address and the same said port number value, extracting the historical port number value and the historical hardware address of each said device with the same name in chronological order to obtain the connection track data corresponding to each said device with the same name; wherein the connection track data comprises timestamp data, historical port number value and historical hardware address; dividing the continuous said timestamp data into multiple time windows to obtain the migration path record of each said device with the same name in each said time window; forming a change event chain according to the port number change sequence and the hardware address replacement sequence in the migration path record, and extracting the optical power value and the bit error rate value before and after each change event as the running feature of the said device with the same name; associating the said timestamp data and the said hardware address replacement sequence with the said running feature to obtain the device identity traceability chain corresponding to each said device with the same name respectively; obtaining the power change sequence by subtracting the optical power values at adjacent time points in the said device identity traceability chain, and obtaining the Pearson correlation coefficient of the power change sequence and a preset standard power change sequence to confirm the real hardware address of each said device with the same name according to the Pearson correlation coefficient; using the Hungarian algorithm to perform bipartite graph optimal matching between the set of said real hardware addresses and the set of currently recorded port numbers to generate a new address-port mapping relationship table; and obtaining the said hardware address and the said port number value corresponding to each said device with the same name according to the new address-port mapping relationship table.
[0096] In this embodiment, the above step creates a suffix identification that fuses the manufacturer identification, device serial number and port number for the device with the same name, so that each optical network terminal in the network has a globally unique and easily identifiable identity, which fundamentally eliminates the identity uncertainty caused by name duplication and provides a unique basis for subsequent accurate alarm attribution and fault location. In the case of hardware address and port number, the real hardware address is traced by analyzing the historical connection track and running feature change of the device with the same name, so that even in the extremely complex scenario where the hardware address and port number are maliciously tampered with or repeated, the real identity of the device can be identified by analyzing its historical behavior characteristics, greatly enhancing the ability to resist abnormalities and data fraud, improving the accuracy of identity recognition, and thus improving the accuracy of fault location.
[0097] Step 105: obtaining the logical address of each said optical network terminal and the conflict logical address in the said logical address according to the said identity, and generating the allocation time sequence data set of each said conflict logical address by querying a preset address mapping record.
[0098] In this embodiment, the step mainly includes: obtaining a logical address of each of the optical network terminals according to the device identity, and counting the occurrence times of each of the logical addresses, so as to determine that the logical address is a conflict logical address when the occurrence times of the logical address are greater than a preset number threshold; determining a plurality of conflict devices using the conflict logical address, and obtaining a plurality of mapping records of the conflict logical address to a hardware address by querying a preset address resolution protocol, and extracting a creation timestamp and an update timestamp of each of the mapping records to obtain an address mapping dataset containing the conflict logical address, the hardware address and the timestamp; arranging the plurality of mapping records in ascending order according to the timestamp in the address mapping dataset, so as to determine an address ownership determination result corresponding to the conflict logical address by comparing the time when each hardware address first establishes a mapping with the conflict logical address; extracting the time point when each of the conflict devices obtains and loses the conflict logical address from each of the mapping records according to the address ownership determination result; and arranging the allocation event of the conflict logical address of each of the conflict devices in time according to the time point to generate an allocation time sequence dataset corresponding to the conflict logical address; wherein the allocation time sequence dataset includes the time point when each of the conflict devices obtains and loses the conflict logical address and the hardware address of each of the conflict devices.
[0099] In this embodiment, the above-mentioned step generates a detailed time sequence dataset of logical address conflict by querying the historical changes of address mapping records, so that it can clearly trace which device occupies a conflict logical address at different time points from the time dimension, thereby accurately matching the alarm occurrence time with the occupied device of the logical address, successfully solving the technical problem of alarm misattribution caused by logical address conflict, and improving the accuracy of fault location.
[0100] Step 106: determining a fault optical network terminal from the plurality of optical network terminals according to the alarm logical address of the network alarm information, the logical address and the allocation time sequence dataset.
[0101] In this embodiment, the step mainly includes: extracting the alarm occurrence timestamp, the alarm source logical address and the alarm type field of the network alarm information; and determining a fault optical network terminal corresponding to the network alarm information by querying a preset event window in combination with the alarm occurrence timestamp as the center, the logical address of the optical network terminal which is not determined as the conflict device, the conflict logical address and the allocation time sequence dataset.
[0102] In the embodiment, the step further comprises: obtaining a sequence of optical power values and a sequence of bit error rate values before and after the alarm occurrence timestamp from the standard state data of the faulty optical network terminal, and forming comprehensive timing data containing the alarm event and the device parameter change through timestamp alignment; calculating a difference between the average optical power in a preset time period before the alarm occurrence timestamp and the average optical power in a preset time period after the alarm occurrence timestamp as the optical power change amplitude according to the comprehensive timing data; calculating a difference between the average bit error rate in a preset time period before the alarm occurrence timestamp and the average bit error rate in a preset time period after the alarm occurrence timestamp as the bit error rate change amplitude according to the comprehensive timing data; when the optical power change amplitude or the bit error rate change amplitude exceeds a preset amplitude threshold, obtaining a timing correlation mode of the network alarm information and the optical power change amplitude or the bit error rate change amplitude; wherein the timing correlation mode comprises a change time of the optical power and a change time of the bit error rate; by comparing the change times with the alarm time, the type of the fault root cause in the faulty optical network terminal is obtained; wherein the type comprises optical power abnormality, bit error rate exceeding the standard and topology configuration error of the passive optical network.
[0103] In the embodiment, the above-mentioned step uses a time window to match the alarm event and the device event, which can accurately associate the abstract alarm event with the specific device state change event on the time line, thereby reliably locking the real faulty device causing the alarm, reducing the false positive and false negative rates, and further improving the accuracy of fault positioning. After positioning the specific faulty device, the type of the root cause is determined by analyzing the statistical change amplitudes of the optical power and the bit error rate before and after the alarm, which not only locates the faulty device, but also further automatically diagnoses the root cause type of the fault, realizes the leap from locating where the fault to diagnosing why the fault, provides a direct repair direction for the operation and maintenance personnel, and greatly improves the operation and maintenance efficiency.
[0104] On the other hand, with reference to Figure 2 The embodiment further discloses a fault root cause positioning system based on automatic analysis, mainly comprising a device positioning module 201, a data conversion module 202, a feature extraction module 203, an identification matching module 204, an address conflict module 205 and a fault positioning module 206.
[0105] The device positioning module 201 is used for identifying an abnormal port according to the port state data of each port in the passive optical network, and extracting a fault port state data set and a plurality of optical network terminals according to the fault correlation of each port with the abnormal port.
[0106] The data conversion module 202 is configured to identify the field characteristics of each device manufacturer according to the fault port state data set, and construct a standardized state field mapping table in combination with a preset field mapping configuration, so as to convert the device state data of each optical network terminal into standard state data.
[0107] The feature extraction module 203 is configured to generate a device feature vector of each optical network terminal according to the standard state data and a preset clustering algorithm, and match the device name of the device feature vector.
[0108] The identification matching module 204 is configured to obtain the address port information of each network terminal according to the device name, so as to generate the identity of each optical network terminal.
[0109] The address conflict module 205 is configured to obtain the logical address of each optical network terminal and the conflict logical address in the logical address according to the identity, and generate an allocation time sequence data set of each conflict logical address by querying a preset address mapping record.
[0110] The fault positioning module 206 is configured to determine a fault optical network terminal from a plurality of optical network terminals according to an alarm logical address of network alarm information, the logical address, and the allocation time sequence data set.
[0111] The embodiment discloses a fault root cause positioning method and system based on automated analysis. First, the abnormal state of the port is identified to only position the optical network terminal connected to the abnormal port, which greatly reduces the workload of abnormal port device positioning and improves the efficiency of fault positioning. Then, a standardized state field mapping table is constructed according to the port state data set, which unifies the format of heterogeneous device state data of multiple manufacturers, eliminates parsing errors caused by field differences, and lays a foundation for subsequent accurate analysis. Second, a globally unique device identity is generated and the time sequence of address conflicts is analyzed, which completely solves the problem of alarm attribution confusion caused by device name duplication or logical address conflict, and realizes accurate association of alarm information to a specific optical network terminal, thereby significantly improving the automation and accuracy of fault root cause positioning and reducing manual intervention and operation time.
[0112] Embodiment Two
[0113] Reference Figure 3 To solve the problems of non-uniform data format of multiple manufacturers' devices, ambiguous device identity, and complex fault positioning in the prior art, the embodiment provides a fault root cause positioning method based on automated analysis, which not only accurately positions the fault optical network terminal that appears abnormal but also accurately identifies the fault type of the fault optical network terminal, thereby improving the accuracy of fault root cause analysis. Specifically, the specific steps of the fault root cause positioning method are as follows:
[0114] Step 301: Obtain the port state data corresponding to each port in the passive optical network, and group the ports according to the state data and a preset fault correlation analysis method to obtain a set of port state data of a fault port group and a plurality of optical network terminals connected to each fault port.
[0115] In this embodiment, this step mainly includes: collecting original port state data corresponding to each port in the passive optical network according to a preset network management protocol; wherein the original port state data is a binary data stream that has not been processed; parsing the original port state data according to a preset data format specification to convert the binary port state data into structured port state data; wherein the structured port state data includes an optical power value, port coding information and a collection timestamp corresponding to each port; wherein when the optical power value is lower than a preset power threshold, the port is marked as an abnormal port, and an alarm level of the abnormal port is obtained according to a comparison result of the optical power value and the power threshold; for the abnormal port and its alarm level, according to a historical optical power value sequence of the abnormal port within a preset collection period, the optical power decay rate of the abnormal port is obtained by dividing the power difference between adjacent time points by the time interval, and when the optical power decay rate exceeds a preset rate threshold, the abnormal type of the abnormal port is determined as a rapid degradation fault type; when the rapid degradation fault type is determined, according to the port coding information of the abnormal port, a plurality of other ports connected to the abnormal port except the abnormal port are extracted; the port state data of each of the other ports is obtained, and the Pearson correlation coefficient of the optical power decay rate of each of the other ports and the abnormal port is calculated according to the port state data; when the Pearson correlation coefficient is greater than a preset coefficient threshold, the other port and the abnormal port are divided into the same port group, to obtain a fault port group grouped according to fault correlation and structured port state data corresponding to each fault port in the fault port group; the structured port state data corresponding to each fault port is integrated, the original data format and manufacturer characteristic information are retained, to obtain a set of port state data including an original state field, a timestamp and a unit of measurement, and a plurality of optical network terminals associated with each fault port.
[0116] Specifically, in the optical source network, the optical network terminal is deployed at the user end to convert optical signals into electrical signals and provide Internet access for users, and the ports in the optical source network can connect a plurality of optical network terminals through a splitter or the like to send optical signals on an optical fiber to a plurality of optical network terminals connected to the port through the port.
[0117] Preferably, in the process of fault positioning, an automated fault root cause positioning system can be constructed for automated processing of fault positioning. For example, a CI / CD automated fault root cause positioning framework can be built, and data acquisition modules, data processing modules, fault analysis modules, fault positioning output modules, etc. can be deployed in the framework, and the framework can be connected to the interface of the management platform of the passive optical network, the fault rule library and the early warning threshold parameters can be initialized, and the basic environment construction of the fault root cause positioning system can be completed. In this regard, the fault root cause positioning system can be built using a microservice architecture, and each functional module in the system can be independently deployed and communicated through an API interface. In order to improve the performance of the fault root cause positioning system, the system can be deployed on a cloud platform or a local server cluster, and containerization technology can be used to achieve rapid deployment and elastic expansion. The data acquisition module in the fault root cause positioning system is responsible for communication with network devices, the data processing module performs format unification and standardization, the fault analysis module executes root cause positioning algorithms, and the fault positioning output module pushes fault positioning information.
[0118] After the system is built, the system can collect port state data corresponding to each port in the passive optical network through a predetermined network management protocol, such as the SNMP protocol (Simple Network Management Protocol). When collecting data through the SNMP protocol, the SNMP protocol accesses the management information base of each device in the managed optical network through a periodic polling mechanism to obtain the port state data of each port from the management information base. The original port state data is transmitted in binary stream form, and after the binary port state data is parsed according to the predetermined data format specification, it is converted into readable structured port state data. The optical power value of the port in the structured port state data reflects the intensity of the optical signal during transmission, and when the optical fiber ages or the connector is contaminated, the optical power will gradually decrease. When identifying abnormal attenuation of the optical power at a predetermined power threshold, the signal reception sensitivity of the optical module in the port can be determined according to the predetermined power threshold to ensure that the signal quality meets the bit error rate requirement. The record of the alarm level of the abnormal port provides an important basis for subsequent fault severity judgment. It should be noted that the calculation of the optical power attenuation rate is crucial for predicting the development trend of the fault. By extracting the optical power values at consecutive time points to form time series data, the instantaneous attenuation rate can be obtained by dividing the power difference between adjacent sampling points by the time interval. This calculation method can reflect the dynamic process of optical path degradation, and rapid degradation usually indicates the risk of optical fiber breakage or connector failure. By reading the coding information of the port, the fault location can be accurately positioned, and accurate fault positioning information can be provided for maintenance personnel.
[0119] Specifically, the Pearson correlation coefficient plays an important role in port failure correlation analysis. When the optical power attenuations of multiple ports present similar change trends, the correlation coefficient is close to 1, indicating that these ports may be affected by the same failure source; for example, when the optical power attenuations of multiple ports present similar change trends, the correlation coefficient is close to 1, indicating that these ports may be affected by the same failure source, and the correlation grouping can identify the common cause failure mode.
[0120] In some embodiments of the present embodiment, the data format unification processing ensures the compatibility of data of different equipment manufacturers. In the present embodiment, various time representations are converted into the unified YYYY-MM-DDTHH:mm:ss format by using a unified time stamp such as the ISO8601 time stamp standard, eliminating the confusion caused by time zone differences. The IEC61850 specification formulates a unified data model for power communication systems, and the state field coding system can accurately describe the equipment operating state. At this time, the original data format of each equipment manufacturer is retained to provide original information for subsequent standardization processing.
[0121] Step 302: identifying the field characteristics of each equipment manufacturer according to the port state data set and constructing a standardized state field mapping table in combination with a pre-set field mapping configuration, so as to convert the equipment state data of each optical network terminal into standard state data according to the standardized state field mapping table.
[0122] In the embodiment, the step mainly includes: reading the original report structure of the port state dataset to identify the manufacturer identifier corresponding to each device manufacturer, and the field name and data format corresponding to the manufacturer identifier; according to the field name and the data format, extracting the state code, field type and numerical content corresponding to each device manufacturer by matching the pre-defined state field feature mode to obtain an original field set containing manufacturer features; calling the field mapping configuration corresponding to each device manufacturer from a pre-set configuration library according to the manufacturer identifier and the field type in the original field set to construct the mapping relationship of the field name to the standard field corresponding to each device manufacturer; wherein, if the field type matches the debugging information feature, the field type is set as a redundant mark and the business field to be retained in the field corresponding to the field type and the mapping rule thereof are determined; according to the business field and the mapping rule, converting the state code in the original field set according to the mapping relationship and reorganizing the field order and format according to the pre-set standard data structure definition, eliminating the debugging information field with the redundant mark, and generating a standardized state field mapping table containing uniform field names and data formats; performing batch conversion processing on the device report uploaded by each optical network terminal through the standardized state field mapping table, converting the field name specific to each device manufacturer in the device report into a standard field name, and converting the data representation form according to the format specification defined in the mapping table to obtain the standard state data corresponding to each optical network terminal.
[0123] In some embodiments of the present embodiment, the original field set is output according to the constructed CI / CD automatic fault root cause positioning system, and the field mapping configuration is called from the pre-set CI / CD configuration library to generate a standardized state field mapping table containing uniform field names and data formats. Then, the device reports uploaded by multiple optical network terminals are collected through the pre-set CI / CD pipeline in the CI / CD automatic fault root cause positioning system, and the device reports are processed to obtain the standard state data corresponding to each optical network terminal.
[0124] Specifically, the fault root cause positioning system can directly read the original report structure from the port state data set to parse the report format of each port; wherein, the identification process of the original report structure involves in-depth parsing of the output format of different optical network terminals. Different optical network terminals of different equipment manufacturers use their own proprietary formats when reporting state information, such as different optical network terminals may use "optical-power" or "rx_power" to represent optical power. The pre-defined state field feature mode identifies these differentiated expressions through keyword matching and context analysis; for example, when a value followed by a "dBm" unit is detected, it is determined that the field is an optical power related parameter, and when "alarm", "warning" and other keywords are found, it is identified as an alarm state field. It should be noted that the CI / CD configuration library plays a core role in the entire standardization process. The configuration library stores the field mapping rules of each manufacturer's equipment, which are accumulated through long-term operation and maintenance experience. The field mapping configuration includes source field name, target standard field name, data type conversion rule and value domain mapping relationship. The identification basis of the debug information feature includes the field name containing keywords such as "debug" and "test", or the field value being temporary data such as process ID and memory address. This identification mechanism ensures that only business data that is actually valuable for fault analysis is retained. The generation process of the standardized state field mapping table embodies the core concept of data standardization. The mapping table uses a key-value pair structure, with each manufacturer's field corresponding to a standard field, and necessary conversion rules are recorded. For example, a certain equipment manufacturer uses 0 and 1 to represent port status, while the standard format requires the use of "UP" and "DOWN", and the mapping table contains such value conversion rules, which are ensured by the pre-defined standard data structure definition to ensure that the data of different equipment manufacturers has consistent semantic expression after conversion.
[0125] In some embodiments of the present embodiment, the CI / CD pipeline integrates automatic data collection and conversion functions. When the state data of a new optical network terminal enters the pipeline, the system automatically identifies the data source manufacturer and calls the corresponding mapping rules for conversion. Batch conversion processing uses a streaming processing method to avoid memory pressure caused by simultaneous processing of a large amount of data. During the conversion process, the system records the conversion log of each field, which facilitates subsequent auditing and problem tracking.
[0126] Step 303: According to a pre-set clustering algorithm, a device feature vector of each optical network terminal is extracted from the standard state data, and a preliminary identity of each optical network terminal is generated according to the optical power and the bit error rate in the device feature vector.
[0127] In the embodiment, the step mainly includes: reading the optical power value and the bit error rate value corresponding to each collection time from the standard state data respectively; generating optical power value sequence and bit error rate value sequence according to the collection time in time stamp order respectively; calculating the optical power mean and the bit error rate mean by accumulating the optical power value and the bit error rate value of each time point respectively divided by the total number of time points; calculating the variance by the square sum of the value and the mean difference of each time point divided by the total number of time points, obtaining four-dimensional feature data including optical power mean, optical power variance, bit error rate mean and bit error rate variance. According to the four-dimensional feature data, a data matrix is constructed and a clustering algorithm is used to set the clustering number to the preset device type quantity. The square sum of the distance of each data point to the clustering center is minimized by iteratively updating the clustering center position. The association relationship formed by each data point and its clustering center is the device feature vector of the optical network terminal; wherein the association relationship is the combination of the data point coordinates of each optical network terminal and the clustering center coordinates to which it belongs, obtaining the device feature vector marked with clustering attribution. For each cluster in the device feature vector, the maximum and minimum optical power of all optical network terminals in the cluster are counted to form the optical power range, and the bit error rate distribution interval is calculated in the same way. According to the optical power range and the bit error rate distribution interval, a preset device performance specification table is matched to determine the device model or performance level corresponding to each cluster as the preliminary identity of each optical network terminal.
[0128] In some embodiments of the present embodiment, the construction process of the four-dimensional feature data embodies the extraction of statistical characteristics of time series data. The statistical characteristics of optical power and bit error rate are processed respectively after sorting by timestamp to obtain the mean and variance. Among them, the mean of optical power reflects the average transmission performance of the optical network terminal in the collection period. When the optical fiber ages or the joint loosens, the mean will continuously decrease. The variance represents the stability of the optical power. The variance of the normal optical network terminal is small, while the variance of the optical network terminal with intermittent faults is large. The statistical characteristics of the bit error rate have similar significance. The mean reflects the overall level of transmission quality, and the variance reflects the degree of quality fluctuation. These four dimensions depict the running characteristics of the optical network terminal from different angles, providing a comprehensive data basis for subsequent clustering analysis. It should be noted that the application of the clustering algorithm, such as the K-means clustering algorithm, in device classification is based on the principle that similar devices have similar performance parameters. The clustering algorithm makes the same type of optical network terminal gather together and the different types of optical network terminal separate from each other through iterative optimization. The setting of the number of clusters K is usually based on the number of device models actually deployed in the network. For example, if an operator network deploys three different specifications of ONU devices, K is set to 3. The initial clustering center can be randomly selected or set based on empirical values. After multiple iterations, each clustering center tends to be stable, representing the typical characteristics of the device of that class. The minimization process of the distance square sum ensures the tightness and separation of the clustering results. Specifically, the formation of the device feature vector set marks the completion of the conversion from raw data to structured features. Each device is no longer an isolated data point, but a feature vector with a clustering label, which represents the similarity of the device. For example, the feature vector of a certain device is classified into the second cluster, which means that the device has similar optical power and bit error rate characteristics as the second type of device. This clustering relationship provides an important basis for subsequent identity recognition.
[0129] In some embodiments of the present embodiment, the device performance specification table is a pre-established reference standard that contains the standard performance parameter range of each type of optical network terminal. For example, the optical power range of high-end optical network terminals is usually between -15 and -20 dBm, and the bit error rate is less than 10 to the power of -9; the optical power range of medium-end devices is between -20 and -25 dBm, and the bit error rate is around 10 to the power of -8; the optical power range of low-end devices is between -25 and -28 dBm, and the bit error rate can reach 10 to the power of -7. By matching the statistical characteristics of the clustering results with the reference table, the device model or performance level corresponding to each cluster can be inferred.
[0130] Step 304: obtaining the device name of each optical network terminal according to the preliminary identity, and dividing each optical network terminal into non-repeated device and repeated device according to the device name and extracting the hardware address and port number of the repeated device, and generating the device identity of each optical network terminal by combining the device name.
[0131] In the embodiment, the step mainly includes: querying the corresponding device name in the preset device configuration data according to the preliminary identity; counting the occurrence times of each device name in the current passive optical network, and determining that the optical network terminal corresponding to the device name exists identity conflict if the device name occurs multiple times, and generating a repeated device list corresponding to the device name. For each repeated device in the repeated device list, reading the hardware address and port number signal from the device state data set, extracting the manufacturer identification code and device serial number in the hardware address to combine the port number information to obtain the device distinguishing attribute set containing the address feature and the port number; combining the address feature and the port number in the device distinguishing attribute set according to the preset format and adding it after the device name to constitute the suffix identification of the device name, so as to realize the secondary distinction of the repeated device according to the suffix identification and the device name, and generate the device identity of each optical network terminal. Preferably, the device configuration data is updated according to the device identity, and a mapping relationship library of device name to identity is established, and when subsequent data flows in, the device is matched according to the mapping relationship library to solve the phenomenon of device identity recognition ambiguity.
[0132] It should be noted that when multiple said devices with the same name use the same said hardware address and the same said port number value, at this time it is determined that the suffix identifier exists conflict, at this time the historical port number value and the historical hardware address of each said device with the same name can be extracted according to the timestamp order to obtain the connection track data corresponding to each said device with the same name; wherein, the connection track data includes timestamp data, historical port number value and historical hardware address; the continuous said timestamp data is divided into multiple time windows to obtain the migration path record of each said device with the same name in each said time window; according to the port number change sequence and the hardware address replacement sequence in the migration path record, a change event chain is formed, and the optical power value and the bit error rate value before and after each change event are extracted as the running characteristics of the said device with the same name; the timestamp data and the hardware address replacement sequence are associated with the running characteristics to obtain the device identity identification trace chain corresponding to each said device with the same name respectively; according to the power change sequence obtained by subtracting the optical power values of adjacent time points in the device identity identification trace chain, and the Pearson correlation coefficient of the power change sequence and the preset standard power change sequence, the real hardware address of each said device with the same name is confirmed according to the Pearson correlation coefficient; the set of confirmed real hardware addresses and the set of currently recorded port numbers are bipartite graph optimal matched by using the Hungarian algorithm to generate a new address port mapping relationship table; according to the new address port mapping relationship table, the hardware address and the port number value corresponding to each said device with the same name are obtained.
[0133] In some embodiments of the present embodiment, when multiple optical network terminals in a passive optical network are configured with the same device name, it can cause network management confusion and difficulty in fault location. Device configuration data is usually stored in the database of the network management system corresponding to the passive optical network, which contains basic information of the device such as name, model, installation location, etc. By traversing the database and counting the frequency of occurrence of each device name, the existence of the conflicting device group can be quickly found. It should be noted that the hardware address as the hardware identifier of the optical network terminal has global uniqueness. The hardware consists of two parts, the manufacturer identification code and the device serial number. The former is uniformly assigned to each manufacturer by the Institute of Electrical and Electronics Engineers, and the latter is assigned by the manufacturer to ensure uniqueness. The port number reflects the physical connection position of the optical network terminal. In the passive optical network, each port corresponds to a specific user access point. The combination of these two attributes forms a composite identifier that contains both device hardware characteristics and network location information, greatly reducing the possibility of identifier conflict. Specifically, the combination method of the preset format determines the readability and practicality of the final identifier. A common format is to connect the device name, the last six bits of the hardware address, and the port information with an underscore, such as ONU_A1B2C3_01, where ONU is the device name, A1B2C3 is the hardware address feature, and 01 is the port information. This format not only retains the original device name for easy identification by maintenance personnel, but also achieves uniqueness through additional information. The extraction of the hardware address usually selects the last few bits because the manufacturer identification code in front is often the same in the same network, while the sequence number has better discrimination. Preferably, the updating of the device configuration data in the CI / CD pipeline in the system embodies the advantages of automated operation. The mapping relationship library is stored in a key-value pair structure, with the key being the original device identifier information such as the address and port combination, and the value being the generated unique identity. When new monitoring data or alarm information arrives, the system first extracts the device identifier information from the data, and then looks up the corresponding unique identifier in the mapping relationship library, thereby accurately locating the specific device.
[0134] It should be noted that when multiple devices with the same name use the same combination of hardware address and port information as the device identity, it is determined that there is a combination conflict. When the combination of hardware address and port information causes abnormal device identity, the device record with identity conflict is retrieved from the standard state data, the historical port information sequence and hardware address sequence of each device are extracted in chronological order, the connection trajectory data containing timestamp, port information and hardware address are constructed, the continuous timestamp data is divided into multiple time windows, and the migration path record of the device in each time window is obtained. According to the port information change sequence and the hardware address replacement sequence in the migration path record, the change event chain is arranged in chronological order, the optical power value and the bit error rate value before and after each change event are extracted as the device operation characteristics, the time and location change information are associated with the operation characteristics, and the complete device identity trace chain is obtained. The power change sequence is obtained by subtracting the optical power values at adjacent time points in the trace chain, the sequence is calculated with the power change mode of the historical normal device by Pearson correlation coefficient, and if the correlation coefficient exceeds the preset threshold, it is confirmed that the real physical location of the device has not changed, and the port number drift phenomenon and the media access control address conflict source are identified. The Hungarian algorithm is used to perform bipartite optimal matching between the set of confirmed device real physical locations and the set of recorded address port combinations, the input of the algorithm is the matching cost matrix between the two sets, and the output is a one-to-one matching scheme with the minimum total cost. A new address port mapping relationship table is generated as a device identity repair scheme according to the matching result, and when a new conflict is detected, the above process is automatically triggered to realize dynamic correction.
[0135] In some embodiments of the present embodiment, the construction process of the connection trajectory data reveals the essence of the dynamic change of network topology. When the network is expanded or fault is repaired, the device can be migrated to different ports, and the hardware address can also be changed in some cases, such as replacing the network card or resetting the device. By extracting these change information in timestamp order, the moving trajectory of the device in the network is formed. The system divides the continuous timestamp sequence into fixed-length time windows (such as 24 hours per window), analyzes the port information transformation and hardware address change of the device in each time window, and obtains the migration path of the device in each time window. The historical port number sequence is the port number change sequence, and the hardware address sequence is the address replacement sequence. For example, a certain optical network terminal has experienced three port changes in a month, from port 1 to port 5, and then to port 8, and each migration will leave a timestamp record in the data set. Such trajectory data provides an important basis for subsequent anomaly detection. It should be noted that the establishment of the device identity traceability chain is based on the event-driven data correlation principle. The change event chain not only records the change of location information, but more importantly, captures the running characteristics of the device before and after the change. As key indicators of device running status, optical power and bit error rate have a relatively stable numerical range under normal circumstances. When the physical location of the device does not change and only the configuration information is incorrect, the optical power change pattern will remain consistent. The continuity of this feature is an important basis for identifying the true device identity.
[0136] Specifically, the application of Pearson correlation coefficient in power change pattern matching embodies the value of statistical methods in network management. The calculation of correlation coefficient is based on the ratio of the covariance of two sequences to their respective standard deviations, with a value range of -1 to 1. When the calculation result is close to 1, it indicates that the two power change sequences are highly similar, meaning that they are likely to be the same device. Through this method, the system can identify the port information drift phenomenon (device does not move but port number record changes) and the root cause of hardware address conflict (multiple ports incorrectly use the same hardware address). For example, if the historical power change sequence of a certain device is a gradual decay trend, and if a device with a questionable identity also presents the same decay pattern with a correlation coefficient of 0.9 or higher, it can be determined as the same device. The application of the Hungarian algorithm solves the optimal matching problem of devices and ports. This algorithm is derived from the bipartite graph matching theory in graph theory, which finds the perfect matching with the smallest total cost by constructing a cost matrix. In the device identity repair scenario, the set of true physical locations of the devices contains the actual deployment locations of all devices in the network, and each device occupies an element in the set. The cost can be defined as the difference between the device characteristics and the historical data of the port. The algorithm iteratively optimizes through the augmentation path, and finally finds the most suitable port allocation for each device. The design of the automatic trigger mechanism enables the system to immediately start the repair process when a new identity conflict is detected, without the need for manual intervention.
[0137] Step 305: obtaining the logical address of each optical network terminal according to the device identity and identifying the conflict logical address, and generating the allocation time sequence dataset corresponding to the conflict logical address by querying the mapping record of the conflict logical address and the hardware address.
[0138] In the embodiment, the step mainly includes: identifying multiple optical network terminals using the same logical address according to the device identity to form a logical address conflict device list; counting the occurrence times of each logical address according to the logical address in the logical address conflict device list, marking the logical address as a conflict logical address if the logical address occurs multiple times to obtain a conflict logical address device list. According to the conflict logical address in the conflict logical address device list, querying the preset address resolution protocol to obtain the mapping record of the conflict logical address to the hardware address, extracting the creation timestamp and the update timestamp of each mapping record to form an address mapping dataset containing the conflict logical address, the hardware address and the timestamp. Arranging all the mapping records of the same conflict logical address in ascending order of the timestamp in the address mapping dataset by the timestamp information, comparing the time of the first mapping of each hardware address to the conflict logical address, and determining the optical network terminal that establishes the earliest mapping as the actual owner of the conflict logical address to obtain the conflict logical address ownership determination result. Using the conflict logical address ownership determination result, extracting the time point of each device obtaining and losing the conflict logical address from the historical record of the address resolution protocol, arranging the conflict logical address allocation event sequence in time order to record the start timestamp and the end timestamp of each allocation, arranging the conflict logical address allocation event of each device in time order to record the start time and the end time of each network management protocol lease, and generating an allocation time sequence dataset containing the hardware address of the device, the conflict logical address, and the allocation start and end time.
[0139] In some embodiments of the present embodiment, the construction process of the allocation time sequence dataset can refer to Figure 4 . As Figure 4As shown, first, all device identities are scanned, and the logical address used by each device is extracted. The frequency of occurrence of each logical address is counted through a hash table. When a logical address is used by multiple devices with different hardware addresses, these devices are added to the logical address conflict device list. The detection process of logical address conflict involves a comprehensive scan of the network resource allocation state. In an optical network environment, when multiple optical network terminals are incorrectly configured with the same logical address, it will cause abnormal network communication. This conflict is common in manual configuration scenarios or address pool confusion after network management protocol server, such as DHCP server, fails to recover. By traversing the device identity and counting the frequency of occurrence of each logical address, the logical address with conflict can be quickly located. For example, when it is found that the logical address 192.168.1.100 is used by three different optical network terminals at the same time, the three devices are immediately marked as a conflict device group. It should be noted that the address resolution protocol, such as ARP table, plays a key role in solving the logical address attribution problem. Its core function is to establish the mapping relationship between logical address and hardware address. Whenever a device needs to communicate with other devices, it will request the hardware address corresponding to the target logical address through the address resolution protocol. This mapping relationship will be cached in the address resolution protocol. The address resolution protocol not only records the current mapping relationship, but also saves the creation and update timestamps of each record. These time information becomes an important basis for determining the real attribution of the logical address. Specifically, the process of timestamp sorting and comparison embodies the application of time sequence analysis in network management. When multiple hardware addresses have established mapping with the same logical address, by comparing the first creation time of each mapping record, it can be determined which device first obtains the use right of the logical address. This time-based priority determination method conforms to the basic principle of network address allocation. For example, if the hardware address AA:BB:CC:DD:EE:FF first establishes mapping with the logical address 192.168.1.100 at 9 am, and another hardware address establishes mapping at 10 am, the former is identified as the legitimate user of the logical address.
[0140] In some embodiments of the present embodiment, the construction process of the allocation time sequence data set corresponding to the conflict logical address needs to deeply mine the historical records of the address resolution protocol. The system identifies the time nodes at which each device obtains and releases the logical address by analyzing the change log of the address resolution protocol. The start time of allocation is defined as the time when a certain hardware address first establishes mapping with a specific logical address, and the end time is the time when the mapping relationship is covered by a new hardware address-logical address pair or actively deleted. This time sequence data not only records the historical track of address allocation, but also reveals the usage mode of address resources in the network.
[0141] Step 306: determining the fault optical network terminal corresponding to the network alarm information according to the alarm logic address, the timestamp sequence of the network alarm information, and the logical address and the allocation time sequence dataset, by time window matching and device identification association.
[0142] In the embodiment, the step mainly includes: obtaining network alarm information and extracting the alarm occurrence timestamp, alarm source logical address and alarm type field of the network alarm information; querying the home optical network terminal corresponding to the logical address at the alarm timestamp according to the logical address and the allocation time sequence dataset, to obtain preliminary matching data containing the network alarm information and the home optical network terminal. When the home optical network terminal belongs to the optical network terminal of the allocation time sequence dataset, a time window of a preset time length before and after the alarm timestamp is set for the network alarm information in the preliminary matching data, and the logical address allocation in the window period is searched in the allocation time sequence dataset. If the alarm source logical address is only allocated to a single device within the window, a unique match is determined, and a deterministic corresponding relationship between the network alarm information and the fault optical network terminal is obtained. Preferably, after the fault optical network terminal is determined, the information number of the network alarm information, the device identity of the fault optical network terminal, the alarm occurrence timestamp, the alarm source logical address and the window matching uniqueness flag are combined and stored to construct a mapping record in the form of key-value pair, and an alarm attribution relationship table containing complete alarm traceability information is generated. Preferably, the historical alarm attribution relationship table is periodically read, a preset proportion of historical alarm attribution records is extracted, whether the same device generates the same alarm again is compared, the accuracy rate is calculated by dividing the number of successful matches by the total number of samples, and if the accuracy rate is lower than a preset threshold, the time window parameters are adjusted and the matching process is re-executed.
[0143] In some embodiments of the present embodiment, the network management interface serves as the collection portal of alarm information, receiving alarm messages from network devices in real time through SNMP Trap or Syslog protocol. These alarm information contains rich metadata, such as the precise timestamp of alarm occurrence, the logical address of the device triggering the alarm, the severity of the alarm and the specific fault type. The logical address of the alarm source is the key information, which points to the identification of the device generating the alarm at the network layer. However, due to the dynamic allocation characteristics of the logical address, the same logical address may correspond to different physical devices at different times, which requires accurate matching combined with the time dimension. It should be noted that the setting of the time window is crucial for accurate matching. The time window is a time range centered on the alarm occurrence time, usually set to 30 minutes or 1 hour before and after. The selection of this range is based on the typical length of the logical address lease and the frequency of device restart in the network. Within this window, the system queries the logical allocation history record to determine whether the logical address of the alarm source has multiple allocation conditions. If a logical address is always allocated to the same device within the entire time window, the alarm can be determined with high confidence. This time constraint mechanism effectively avoids false alarms caused by logical address reallocation.
[0144] Preferably, the construction of the alarm attribution table adopts a structured data organization method. Each record contains an alarm number as the primary key, a device unique identifier for precise positioning of the physical device, an alarm occurrence time record event timing, a logical address preserving network layer information, and a window matching uniqueness flag indicating the matching confidence. When the flag is true, it means that the logical address corresponds to only one device within the time window; when it is false, further correlation analysis is needed. This design not only ensures the integrity of the data, but also provides a basis for subsequent verification.
[0145] In some embodiments of the present embodiment, the periodic verification mechanism of the CI / CD pipeline embodies the application of the continuous integration idea in network management. The system automatically triggers the verification task at fixed time intervals, such as every hour or every day. The verification process selects a certain proportion of historical alarm records by random sampling, and then traces the same type of alarms generated by these devices in the subsequent time period. The accuracy is calculated based on the simple proportion relationship: the number of correctly matched alarms divided by the total number of samples. When the accuracy falls below the threshold, the system will automatically adjust the time window size or update the matching rules.
[0146] Step 307: Correlation analysis and pattern recognition are performed on the alarm attribution table to obtain the correlation pattern between the alarm information and the optical power change and the bit error rate change of the faulty optical network terminal. The type of fault root cause is determined by comparing and analyzing the alarm occurrence time and the device parameter abnormal time.
[0147] In the embodiment, the step mainly includes: obtaining the optical power value sequence and the bit error rate value sequence before and after the alarm occurrence timestamp from the standardized state field data set of the fault optical network terminal, and forming comprehensive timing data containing alarm events and device parameter changes through timestamp alignment; calculating the difference between the optical power mean value within a preset time length before the alarm occurrence timestamp and the optical power mean value within a preset time length after the alarm occurrence timestamp as the optical power change amplitude according to the comprehensive timing data; calculating the difference between the bit error rate mean value within a preset time length before the alarm occurrence timestamp and the bit error rate mean value within a preset time length after the alarm occurrence timestamp as the bit error rate change amplitude according to the comprehensive timing data; when the optical power change amplitude or the bit error rate change amplitude exceeds a preset amplitude threshold, obtaining the timing correlation mode of the network alarm information and the optical power change amplitude or the bit error rate change amplitude; wherein the timing correlation mode includes the change time of the optical power and the change time of the bit error rate; by comparing the change time with the alarm time, the type of the fault root cause in the fault optical network terminal is obtained; wherein the type includes optical power abnormality, bit error rate exceeding standard and passive optical network topology configuration error.
[0148] In some embodiments of the present embodiment, the construction process of the comprehensive timing data embodies the importance of multi-source data fusion. The comprehensive timing data usually contains alarm number, occurrence time, alarm level and type code and other information, wherein the alarm type code follows the industry standard definition, such as 01 representing optical power alarm, 02 representing bit error rate alarm, and 03 representing connection interruption alarm. The device parameter data comes from continuous performance monitoring, and the optical power and bit error rate values are collected once every fixed time interval. Through the timestamp alignment technology, the discrete alarm events and the continuous parameter change curve are associated to form a complete fault evolution view. This timing alignment avoids the analysis deviation caused by different data collection frequencies. It should be noted that the parameter change amplitude is calculated by using the time window mean value comparison method.
[0149] In some embodiments of the present application, all optical power sampling values within 30 minutes before the alarm occurs are selected, and the arithmetic mean value thereof is calculated as the reference power; sampling values within 30 minutes after the alarm occurs are selected to calculate the average power, and the difference between the two is the power change amplitude. This mean calculation method can effectively filter the interference of transient fluctuations and reflect the overall trend of parameter changes. The same principle is used for the calculation of the bit error rate, but since the bit error rate itself is a ratio value, its change amplitude is usually expressed in multiples, such as from 10 to the power of -9 to 10 to the power of -7, with a change amplitude of 100 times. Specifically, the recognition of the time sequence correlation mode is based on the time characteristics of the causal relationship. In optical fiber communication, optical power attenuation is usually a gradual process, and it takes a certain time from the beginning of the decline to trigger the alarm threshold. This time difference reflects the development speed of the fault. For example, the power loss caused by fiber bending may gradually intensify within a few hours, while the loss caused by loose connectors may increase sharply within a few minutes. By analyzing the time relationship between parameter changes and alarms, the physical cause of the fault can be inferred. When the optical power decreases more than 15 minutes before the alarm, it usually indicates slow degradation; if it occurs almost simultaneously, it may be a sudden fault.
[0150] In some embodiments of the present application, the type determination logic of the fault root cause is based on the typical characteristics of network faults. As the most common cause of failure, the determination of optical power anomaly is based on a power drop of more than 3 dB and a time lead of more than 15 minutes before the alarm. Bit error rate exceeding the standard usually accompanies a sharp deterioration of data transmission quality, and its characteristic is that the bit error rate crosses multiple orders of magnitude within a short period of time. Topology configuration error is manifested as normal parameters but communication interruption, in which case the change record of device configuration information needs to be checked.
[0151] Step 308: Obtain device performance deviation data according to the type determination result of the fault root cause, determine the fault impact area, and establish a fault early warning mechanism based on the performance deviation data.
[0152] In the embodiment, the step mainly includes: according to the type determination result of the fault root cause, extracting the current optical power value and the bit error rate value of the fault optical network terminal, subtracting the average value of the historical normal operation of the fault optical network terminal to obtain the deviation value, dividing the historical average value by the deviation value and multiplying by 100% to obtain the performance deviation percentage, and querying the downstream connection relationship of the device from the network configuration data to obtain the fault feature set containing the performance deviation data and the downstream device identifier. Through the downstream device identifier in the fault feature set, the current performance parameters of each downstream device are read, the performance deviation percentage is calculated, and if the deviation exceeds the preset threshold, the device is included in the fault influence area. The total number of affected devices and the number of ports in the area are counted, and the fault impact evaluation data containing the impact range and the deviation degree are generated. Using the fault impact evaluation data, the performance deviation values and the network level positions of the devices in the fault impact area are compared, the device with the maximum deviation value and in the upstream position of other fault devices in the network topology is determined as the fault source, the hardware address of the device is extracted, and the fault root cause type determined before is combined to obtain the fault positioning result containing the hardware address of the fault device and the specific fault type. According to the historical performance deviation change rule of the fault device identified in the fault positioning result, the current deviation value is multiplied by a preset coefficient as a warning threshold, and when the performance deviation of any device reaches the corresponding warning threshold in subsequent monitoring, a warning signal is automatically triggered, and a fault warning mechanism based on performance deviation data is established.
[0153] In some embodiments of the present embodiment, the calculation of the performance deviation percentage reflects the importance of the relative change amount. The average value of the historical normal operation is obtained by long-term data accumulation, and the average value is usually calculated based on the data of more than 30 days of stable operation of the device. For example, the average normal optical power of a certain optical network terminal is -20 dBm, the current detection value is -26 dBm, the deviation value is -6 dBm, and the deviation percentage is 30%. This percentage representation eliminates the influence of the reference value difference of different devices, making the performance degradation of different types of devices comparable. The error rate deviation calculation needs special attention, as it is in exponential form, so the deviation is usually calculated after taking the logarithm. It should be noted that the determination of the fault influence area is based on the tree topology characteristics of the optical network. In a passive optical network, one port is connected to multiple optical splitters, and each splitter is connected to multiple network element devices. When an upstream device fails, all downstream devices will be affected. The degree of influence decreases with the level of the topology, and the directly connected devices are most severely affected, while the indirectly connected devices are relatively less affected. By checking the performance deviation of the downstream devices level by level, the influence range of the fault can be accurately outlined. The statistics of the number of ports provides a quantitative index for evaluating the business impact. Specifically, the logic of fault source positioning is based on the directionality principle of fault propagation. In optical fiber communication, signals pass through multiple levels of splitting from the port to each optical network terminal, and the failure of an upstream device will propagate downward along the optical path. By comparing the performance deviation values and topology positions of the devices in the fault area, the source of the fault can be traced back. For example, when a certain optical splitter fails, all the optical network terminals connected to it will experience performance degradation, but the performance deviation of the splitter is usually the largest. At the same time, the optical network terminals connected to other splitters at the same level as the splitter perform normally, which further confirms the location of the fault source. The hardware address as the unique identifier of the device ensures the accuracy of the fault location.
[0154] In some embodiments of the present embodiment, the establishment of the early warning mechanism adopts a dynamic threshold setting method. The selection of the preset coefficient is based on the statistical law of the fault development speed, and the coefficient of the fast degradation type fault is set to 0.5, that is, when the deviation reaches half of the current value, the early warning is triggered; the coefficient of the slow degradation type fault is set to 0.7, leaving more processing time. This differentiated early warning strategy avoids frequent false positives and ensures timely detection of potential faults. The automatic triggering of the early warning signal is realized by continuous monitoring, and the system periodically compares the real-time performance data with the early warning threshold, and generates early warning information as soon as it exceeds.
[0155] The embodiment provides a fault root cause positioning method based on automatic analysis, solves the problems of non-uniform device data formats of multiple manufacturers, device identity ambiguity and complex fault positioning. In the method, a standardized data processing flow is constructed, device state data is extracted and analyzed, a uniform format data set is generated, a K-means clustering algorithm is used to generate a device feature vector, a unique device identity is realized by combining a hardware address and a port number, and the identity ambiguity problem is solved. At the same time, the logical address duplication problem is solved by address resolution protocol query and time series analysis, a device logical address allocation time series data set is generated, and combined with network management system alarm information, the fault root cause and the influence area are accurately positioned through time window matching and correlation analysis. The method provided in the embodiment establishes a fault early warning mechanism through continuous monitoring and performance deviation analysis, and significantly improves the fault positioning efficiency and accuracy, and optimizes the automation level of network operation and maintenance.
[0156] The above specific embodiments further illustrate the purpose, technical solutions and advantages of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the protection scope of the present application. It is particularly pointed out that any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A fault root cause localization method based on automated analysis, characterized in that, include: Abnormal ports are identified based on the port status data of each port in the passive optical network, and fault port status datasets and multiple optical network terminals are extracted based on the fault correlation between each port and the abnormal ports. Based on the fault port status dataset, the field characteristics of each equipment manufacturer are identified and a standardized status field mapping table is constructed in combination with the preset field mapping configuration, so as to convert the equipment status data of each optical network terminal into standard status data. Based on the standard state data and a preset clustering algorithm, a device feature vector is generated for each optical network terminal, and the device name is matched with the device feature vector. The address and port information of each optical network terminal are obtained according to the device name to generate an identity identifier for each optical network terminal. The logical address of each optical network terminal and the conflicting logical address in the logical address are obtained according to the identity identifier. The allocation time series dataset of each conflicting logical address is generated by querying the preset address mapping record. The faulty optical network terminal is determined from among the multiple optical network terminals based on the alarm logical address of the network alarm information, the logical address, and the allocated time series dataset.
2. The fault root cause localization method based on automated analysis according to claim 1, characterized in that, The step of identifying abnormal ports based on port status data of each port in the passive optical network, and extracting fault port status datasets and multiple optical network terminals based on the fault correlation between each port and the abnormal ports, includes: Collect binary port status data corresponding to each port in the passive optical network according to the preset network management protocol; The binary port status data is parsed according to a preset data format specification to convert it into structured port status data; wherein, the structured port status data includes the optical power value of the port, the device encoding information of the port, and the acquisition timestamp; The optical power value is compared with a preset optical power threshold to identify abnormal ports from among the multiple ports based on the comparison result. Obtain the historical optical power value sequence of the abnormal port within a preset acquisition period, and obtain the optical power attenuation rate by dividing the power difference between adjacent time points by the time interval based on the historical optical power value sequence. When the optical power attenuation rate is greater than a preset rate threshold, Pearson correlation coefficients between each of the ports other than the abnormal port and the optical power attenuation rate are calculated based on the structured port status data. Based on the Pearson correlation coefficient and a preset coefficient threshold, all ports other than the abnormal port are divided to obtain multiple fault-related ports; A fault port group is obtained based on multiple fault-related ports and abnormal ports; Based on the abnormal port and the structured port status data of each of the fault-related ports, the fault port status dataset corresponding to the fault port group and its associated multiple optical network terminals are obtained.
3. The fault root cause localization method based on automated analysis according to claim 2, characterized in that, The step of identifying field features of each equipment manufacturer based on the fault port status dataset and constructing a standardized status field mapping table in conjunction with a preset field mapping configuration to convert the device status data of each optical network terminal into standard status data includes: Read the original report structure of the fault port status dataset to identify the manufacturer identifier for each device manufacturer and the field name and data format corresponding to the manufacturer identifier; Based on the field name and the data format, the status code, field type and numerical content corresponding to each device manufacturer are extracted by matching predefined status field feature patterns. The system invokes the manufacturer identifier and the field mapping configuration corresponding to the field type to construct the mapping relationship between the field name and the standard field for each of the device manufacturers; wherein, when the field type matches the debugging information feature, a redundancy flag is set for the field type and the business field of the field type and its corresponding mapping rule are determined; The status code is converted according to the business fields, the mapping rules, and the mapping relationship. The order and format of each field are reorganized according to the preset standard data structure definition, and the debugging information fields with the redundant markers are removed to obtain a standardized status field mapping table. The standardized status field mapping table contains unified field names and unified data formats. The device report of each optical network terminal is obtained, and the multiple device reports are batch converted according to the standardized status field mapping table to obtain the standard status data corresponding to each optical network terminal.
4. The fault root cause localization method based on automated analysis according to claim 3, characterized in that, The step of generating device feature vectors for each optical network terminal based on the standard state data and a preset clustering algorithm, and matching the device name with the device feature vectors, includes: Based on the standard state data, read multiple optical power values, multiple bit error rate values, and multiple acquisition times for each optical network terminal, and obtain the average optical power, average bit error rate, optical power variance, and bit error rate variance based on the optical power values, the bit error rate values, and the acquisition duration of the multiple acquisition times. Multidimensional feature data for each optical network terminal is constructed based on the mean optical power, the variance of optical power, the mean bit error rate, and the variance of bit error rate. Based on the multidimensional feature data and the preset clustering algorithm, a set of device feature vectors labeled with cluster affiliation is obtained; Obtain each cluster in the device feature vector set and count the maximum and minimum optical power values of all optical network terminals in the cluster to form an optical power range, and calculate the bit error rate distribution interval of the bit error rate value; Based on the optical power range, the bit error rate distribution interval, and the preset equipment performance specification comparison table, the equipment performance corresponding to each cluster is determined, and the equipment performance is used as the initial equipment identification of the optical network terminal. Based on the device's preliminary identification, the preset device configuration data is searched to obtain the device name corresponding to the optical network terminal.
5. The fault root cause localization method based on automated analysis according to claim 4, characterized in that, The step of obtaining the address and port information of each optical network terminal based on the device name to generate an identity identifier for each optical network terminal includes: The occurrence count of each device name is counted. When the occurrence count exceeds a preset threshold, multiple optical network terminals corresponding to that device name are determined to be devices with the same name. The device name is used as the identity identifier of each optical network terminal that is not identified as a device with the same name. The hardware address and port number value corresponding to each device with the same name are obtained according to the standard status data, so as to extract the manufacturer identification code and device serial number from the hardware address. The manufacturer identification code, the device serial number, and the port number are combined according to a preset format to obtain a suffix identifier. The suffix identifier is then added to the device name to obtain the identity identifier of each device with the same name.
6. The fault root cause localization method based on automated analysis according to claim 5, characterized in that, The step of obtaining the logical addresses of each optical network terminal and the conflicting logical addresses among the logical addresses based on the identity identifier, and generating the allocation time series dataset of each conflicting logical address by querying a preset address mapping record, includes: The logical address of each optical network terminal is obtained based on the identity identifier, and the occurrence frequency of each logical address is counted. When the occurrence frequency of a logical address is greater than a preset threshold, the logical address is determined to be a conflicting logical address. Multiple conflicting devices using the conflicting logical address are identified, and multiple mapping records from the conflicting logical address to the hardware address are obtained by querying a preset address resolution protocol. The creation timestamp and update timestamp of each mapping record are extracted to obtain an address mapping dataset containing the conflicting logical address, hardware address and timestamp. The multiple mapping records are sorted in ascending order according to the timestamps in the address mapping dataset, so as to determine the address ownership result corresponding to the conflicting logical address by comparing the time when each hardware address first establishes a mapping with the conflicting logical address; Based on the address attribution determination result, extract the time points when each conflicting device obtains and loses the conflicting logical address from multiple mapping records; Based on the time points, the allocation events of the conflict logical address of each conflicting device are sorted by time to generate an allocation time series dataset corresponding to the conflict logical address; wherein, the allocation time series dataset includes the time points when each conflicting device obtains and loses the conflict logical address and the hardware address of each conflicting device.
7. The fault root cause localization method based on automated analysis according to claim 6, characterized in that, The step of determining the faulty optical network terminal from multiple optical network terminals based on the alarm logical address of the network alarm information, the logical address, and the allocated time series dataset includes: Extract the alarm occurrence timestamp, alarm source logical address, and alarm type fields from the network alarm information; Based on the logical address of the optical network terminal that was not identified as the conflicting device, the conflicting logical address, and the allocated time series dataset, the faulty optical network terminal corresponding to the network alarm information is queried using the alarm occurrence timestamp as the center and combined with a preset event window.
8. The fault root cause localization method based on automated analysis according to claim 5, characterized in that, The step of obtaining the address and port information of each optical network terminal based on the device name to generate an identity identifier for each optical network terminal further includes: When multiple devices with the same name use the same hardware address and the same port number value, the historical port number value and historical hardware address of each device with the same name are extracted in timestamp order to obtain the connection trajectory data corresponding to each device with the same name; wherein, the connection trajectory data includes timestamp data, historical port number value and historical hardware address; The continuous timestamp data is divided into multiple time windows to obtain the migration path record of each device with the same name in each time window; a change event chain is formed according to the port number change sequence and hardware address replacement sequence in the migration path record, and the optical power value and bit error rate value before and after each change event are extracted as the operating characteristics of the device with the same name. The timestamp data and the hardware address change sequence are associated with the operating characteristics to obtain the device identity traceability chain corresponding to each of the duplicate-named devices; the power change sequence is obtained by subtracting the optical power values of adjacent time points in the device identity traceability chain, and the Pearson correlation coefficient between the power change sequence and the preset standard power change sequence is obtained, so as to confirm the real hardware address of each of the duplicate-named devices based on the Pearson correlation coefficient; The Hungarian algorithm is used to perform bipartite graph optimal matching between the set of real hardware addresses and the set of currently recorded port numbers to generate a new address port mapping table. The hardware address and port number value corresponding to each device with the same name are obtained according to the new address and port mapping table.
9. A fault root cause localization method based on automated analysis according to claim 7, characterized in that, The step of determining the faulty optical network terminal from multiple optical network terminals based on the alarm logical address of the network alarm information, the logical address, and the allocated time series dataset further includes: The optical power value sequence and bit error rate value sequence before and after the alarm occurrence timestamp are obtained from the standard status data of the faulty optical network terminal, and comprehensive time-series data containing alarm events and changes in equipment parameters are formed by aligning the timestamps. The difference between the average optical power within a preset time period before the alarm occurrence timestamp and the average optical power within a preset time period after the alarm occurrence timestamp is calculated based on the comprehensive time series data as the optical power change amplitude. The difference between the average bit error rate within a preset time period before the alarm occurrence timestamp and the average bit error rate within a preset time period after the alarm occurrence timestamp is calculated based on the comprehensive time series data as the bit error rate change range. When the amplitude of the optical power change or the amplitude of the bit error rate change exceeds a preset amplitude threshold, the time-series correlation pattern between the network alarm information and the amplitude of the optical power change or the amplitude of the bit error rate change is obtained; wherein, the time-series correlation pattern includes the optical power change event and the bit error rate change time. By comparing the change time with the alarm time, the type of fault root cause in the faulty optical network terminal is obtained; wherein, the type includes abnormal optical power, excessive bit error rate, and passive optical network topology configuration error.
10. A fault root cause localization system based on automated analysis, characterized in that, It includes a device positioning module, a data conversion module, a feature extraction module, an identifier matching module, an address conflict module, and a fault location module; The device positioning module is used to identify abnormal ports based on the port status data of each port in the passive optical network, and to extract the fault port status dataset and multiple optical network terminals based on the fault correlation between each port and the abnormal port. The data conversion module is used to identify the field characteristics of each equipment manufacturer based on the fault port status dataset and construct a standardized status field mapping table in combination with the preset field mapping configuration, so as to convert the equipment status data of each optical network terminal into standard status data. The feature extraction module is used to generate device feature vectors for each optical network terminal based on the standard state data and a preset clustering algorithm, and to match the device name of the device feature vectors. The identifier matching module is used to obtain the address port information of each optical network terminal according to the device name, so as to generate an identity identifier for each optical network terminal. The address conflict module is used to obtain the logical address of each optical network terminal and the conflicting logical address in the logical address according to the identity identifier, and generate the allocation time series dataset of each conflicting logical address by querying the preset address mapping record; The fault location module is used to determine the faulty optical network terminal from multiple optical network terminals based on the alarm logical address of the network alarm information, the logical address, and the allocated time series dataset.
Citation Information
Patent Citations
Fault data analysis method and device, equipment and storage medium
CN117076166A
Network security monitoring system
CN120074962A