Big data-based cable fault determination method and system
By constructing a cable fault feature gene library and generating a fault association network topology, combined with real-time data updates, the problems of insufficient accuracy and efficiency in traditional cable fault location methods are solved, achieving high-precision and efficient fault location and ensuring the stable operation of the power system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG QUANXING YINQIAO OPTICAL & ELECTRIC CABLE SCI & TECH DEV
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional cable fault location methods cannot fully consider the characteristics of the cable laying environment and historical fault characteristics, resulting in insufficient location accuracy and efficiency, and failing to meet the high precision and high efficiency requirements of modern power systems.
A cable fault feature gene library is constructed, which includes cable operating parameter features, laying environment features and historical fault features. A fault association network topology is generated, and the topology attributes are updated in real time. A spatiotemporal coupled positioning network is constructed to accurately map fault sections.
It enables dynamic simulation and in-depth analysis of cable fault states, improves the accuracy and efficiency of fault location, shortens fault repair time, reduces power outage losses, and ensures the safe and stable operation of the power system.
Smart Images

Figure CN122109713A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, and more specifically, to a method and system for determining cable faults based on big data. Background Technology
[0002] In power systems, cables serve as a crucial carrier for electrical energy transmission, and their safe and stable operation is of paramount importance. However, due to the complex and diverse environments in which cables are laid, including underground, underwater, and at high altitudes, and the long-term exposure to high voltage, high current, and external environmental factors, cable faults occur frequently. Accurately and quickly locating cable faults is of utmost importance for timely fault repair, restoration of power supply, reduction of power outage losses, and ensuring the safe and reliable operation of the power system.
[0003] Currently, traditional cable fault location methods mainly rely on traveling wave methods and impedance methods. The traveling wave method determines the fault location by detecting the propagation time and velocity of the traveling wave generated by the fault in the cable. However, the traveling wave is easily affected by factors such as cable branches and joints during propagation, leading to a decrease in location accuracy. The impedance method calculates the fault location based on the relationship between cable parameters such as resistance and inductance and the fault distance. However, cable parameters change with environmental factors such as temperature and humidity, thus affecting the accuracy of location. Furthermore, these traditional methods mostly consider only the electrical parameters of the cable, neglecting multi-source information such as the characteristics of the cable's laying environment and historical fault characteristics. They cannot comprehensively and accurately reflect the actual operating state of the cable, making it difficult to meet the high-precision and high-efficiency requirements of modern power systems for cable fault location. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a cable fault location method based on big data, the method comprising: A cable fault feature gene library is constructed, which contains fault feature gene units extracted from multi-source cable operation data. The fault feature gene units are associated with cable operation parameter characteristics, laying environment characteristics and historical fault characteristics, and each fault feature gene unit has a unique cable section identifier. Based on the cable fault feature gene library, a fault association network topology is generated. The fault association network topology uses fault feature gene units as nodes and the association strength between different fault feature gene units as links. The attributes of the links include the association duration and the association triggering conditions. The node and link attributes in the fault-related network topology are updated based on the real-time collected cable operation data, the fault-related network topology is evolved, and the topology evolution trajectory is recorded. By combining cable laying path data with the topology evolution trajectory, a spatiotemporal coupled positioning network is constructed. The spatiotemporal coupled positioning network maps the evolution nodes of the fault association network topology and the correspondence between the physical laying sections of the cable. Based on the spatiotemporal coupling positioning network, the physical laying section where the cable fault is located is determined, a cable fault location instruction containing section identifier and fault characteristic gene unit association information is generated, and the cable fault location instruction is transmitted to the cable operation and maintenance terminal.
[0005] In another aspect, embodiments of the present invention also provide a cable fault location system based on big data, including a processor and a machine-readable storage medium connected to the processor. The machine-readable storage medium is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the above-described method.
[0006] Based on the above, this invention constructs a cable fault feature gene library containing cable operating parameter characteristics, laying environment characteristics, and historical fault characteristics. A fault association network topology generated based on this library uses fault feature gene units as nodes and the correlation strength between different nodes as links. It also records the association duration and triggering conditions of each link in detail. This allows for the characterization of complex relationships between fault features, enabling dynamic simulation and in-depth analysis of cable fault states. The node and link attributes in the fault association network topology are updated based on real-time collected cable operating data, and the topology evolution trajectory is recorded. This allows fault location to track changes in cable operating states in real time and promptly identify potential fault hazards. A spatiotemporal coupled positioning network, constructed by combining cable laying path data and topology evolution trajectory, accurately maps the evolution nodes of the fault association network topology to the physical laying sections of the cable. Finally, based on the spatiotemporal coupled positioning network, the physical laying section where the cable fault is located is determined, and a cable fault location command containing section identifiers and fault characteristic gene unit association information is generated and transmitted to the cable operation and maintenance terminal. This provides operation and maintenance personnel with detailed and accurate fault information, greatly improving the accuracy and efficiency of fault location, effectively shortening fault repair time, reducing power outage losses, and ensuring the safe and stable operation of the power system. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of the execution flow of the cable fault location method based on big data provided in the embodiments of the present invention.
[0008] Figure 2 This is a schematic diagram of exemplary hardware and software components of the cable fault location system based on big data provided in an embodiment of the present invention. Detailed Implementation
[0009] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a cable fault location method based on big data according to an embodiment of the present invention. The following is a detailed description of the cable fault location method based on big data.
[0010] Step S110: Construct a cable fault feature gene library, which contains fault feature gene units extracted from multi-source cable operation data. The fault feature gene units are associated with cable operation parameter characteristics, laying environment characteristics and historical fault characteristics, and each fault feature gene unit has a unique cable section identifier.
[0011] This embodiment uses urban underground cable networks as an application scenario. In this scenario, cables are mostly laid along urban main roads and underground pipe networks in residential areas. They are prone to complex faults due to multiple factors such as traffic loads, soil corrosion, construction disturbances, and long-term operational losses. When constructing a cable fault characteristic gene library, it is necessary to integrate the multi-dimensional characteristics of such complex faults and transform scattered operational, environmental, and historical fault data into structured gene units.
[0012] Step S111: Obtain multi-source cable operation data, which includes real-time cable operation parameter data, cable laying environment data, and historical cable fault case data.
[0013] In urban underground cable networks, real-time cable operating parameters are collected through monitoring devices deployed on the cable itself and its auxiliary equipment. Cable body monitoring includes conductor temperature distribution data collected via built-in fiber optic sensors; partial discharge quantity and phase data collected via online partial discharge monitoring devices; insulation volume resistivity and surface resistivity data collected via insulation resistance testers; three-phase current imbalance and harmonic content data collected via current transformers; and temperature data of cable terminations and pressure data of intermediate joints. Cable laying environment data is obtained through monitoring equipment buried around the cable trenches, covering soil pH, redox potential, and moisture content data; surrounding groundwater level changes; electromagnetic radiation intensity data from nearby high-voltage lines; and vibration frequency and amplitude data from road vehicle loads. The historical cable fault case data comes from the archived records of the cable operation and maintenance management platform. It includes detailed information on complex faults that have occurred in the past five years, such as insulation breakdown faults caused by insulation aging and increased partial discharge, conductor damage faults caused by external disturbances and material defects, and sheath failure faults caused by the combined effects of soil corrosion and electrochemical aging. Each case records the precise time of the fault, the cable section involved, the parameter change curves during the fault development process, the characteristics of the fault point found on site, and the repair plan.
[0014] Step S112: Extract parameter segments reflecting the cable's operating status from the real-time operating parameter data of the cable. The parameter segments include conductor temperature change segments, insulation layer resistance change segments, and current fluctuation segments. Mark each parameter segment as an operating characteristic segment.
[0015] For the collected real-time cable operating parameter data, the data is first cleaned to remove outliers caused by sensor drift and missing values due to communication interruptions. For conductor temperature distribution data, continuous temperature change sequences are extracted at preset time intervals as conductor temperature change segments. These segments include not only the temperature values at each monitoring point on the conductor but also the temperature gradient trend and the duration of high temperatures. For example, segments showing a section of cable where the conductor temperature gradually rises from the normal range during the morning peak load, with a significantly higher rate of temperature increase at local points than in other areas. For insulation resistance data, continuous data sequences showing a simultaneous decrease in both volume resistivity and surface resistivity of the insulation layer are extracted as insulation resistance change segments. These segments must demonstrate the magnitude and rate of resistance decrease and whether they are accompanied by fluctuations. For example, segments showing a slow decrease in insulation resistance followed by an intermittent sharp drop due to partial discharge. For current data, continuous sequences where the three-phase current imbalance exceeds the normal range or the harmonic content is abnormally high are extracted as current fluctuation segments. These segments include the fluctuation frequency, fluctuation range, and correlation information with load changes. For example, a segment showing a sudden current surge and a synchronous increase in third harmonic content in a cable section during industrial load startup. Each extracted parameter segment is labeled as an operational characteristic segment, and its acquisition time is associated with the location of the corresponding cable monitoring point.
[0016] Step S113: Extract environmental segments that affect the cable condition from the cable laying environment data. The environmental segments include environmental temperature change segments, soil moisture change segments, and surrounding electromagnetic interference change segments. Mark each environmental segment as an environmental feature segment.
[0017] After preprocessing the cable laying environment data, environmental segments affecting cable conditions are extracted. For ambient temperature data, continuous temperature change sequences at different depths above and below ground are extracted as ambient temperature change segments. These segments must cover sudden temperature changes during seasonal transitions and abnormal temperature fluctuations under extreme weather conditions, such as segments where underground soil temperature drops rapidly and remains below surface temperature after summer rainstorms. For soil moisture data, continuous data sequences where soil moisture content exceeds a critical value or where the rate of moisture change is abnormal are extracted as soil moisture change segments. These segments include the range and duration of the moisture increase and its correlation with groundwater level changes, such as segments where soil moisture rises sharply in a short period due to underground pipe leaks and spreads to the cable laying area. For surrounding electromagnetic interference data, continuous sequences where electromagnetic radiation intensity exceeds a safety threshold or where the interference frequency is close to the cable operating frequency are extracted as surrounding electromagnetic interference change segments. These segments reflect the intensity changes, duration, and impact on cable signal transmission of the interference source, such as segments where the electromagnetic radiation intensity around the cable fluctuates periodically after a newly built substation is put into operation. Each extracted environmental segment is marked as an environmental feature segment, and its corresponding spatial range and duration are recorded.
[0018] Step S114: Extract the operational feature fragments and environmental feature fragments corresponding to the time of the fault from the cable historical fault case data, and bind the fault type information with the corresponding operational feature fragments and environmental feature fragments to form a fault feature combination.
[0019] By analyzing historical cable fault case data, for each complex fault, operational and environmental characteristic segments before and during the fault's development were identified. Taking insulation breakdown faults caused by insulation aging accompanied by increased partial discharge as an example, conductor temperature changes (indicating persistently high local temperatures), insulation resistance changes (indicating a slow decrease in resistance), and partial discharge quantity changes (indicating a gradual increase in discharge quantity) were extracted over the three months prior to the fault. Simultaneously, soil moisture changes (indicating prolonged high soil humidity) and soil pH changes (indicating increased soil acidity) were also extracted during the corresponding period. The fault type information of "insulation breakdown caused by insulation aging accompanied by increased partial discharge" was then linked to the extracted operational and environmental characteristic segments. Taking conductor damage faults caused by external disturbances and material defects as another example, current fluctuations (indicating a sudden drop in current accompanied by abnormal harmonics), vibration frequency changes (indicating strong vibrations from road construction), and soil compaction changes during the corresponding period (indicating increased soil density due to construction) were extracted. This fault type was then linked to these characteristic segments, forming multiple sets of fault characteristic combinations.
[0020] Step S115: Encode the operational feature fragments and environmental feature fragments in each fault feature combination to generate a gene coding fragment containing operational coding sequences and environmental coding sequences.
[0021] An encoding algorithm is used to encode the operational and environmental feature segments in each fault feature combination. For operational feature segments, different encoding rules are set according to feature type: conductor temperature change segments are divided into different encoding intervals according to temperature range, rate of change, and duration, converting continuous temperature data into discrete base sequences; insulation layer resistance change segments are encoded according to resistance decrease amplitude, rate of change, and fluctuation frequency, generating corresponding base sequences; partial discharge quantity change segments are encoded according to discharge quantity level and discharge phase distribution characteristics. The encoding results of each operational feature segment are concatenated in a preset order to form an operational encoding sequence. For environmental feature segments, encoding is also performed according to feature type: environmental temperature change segments are encoded according to temperature change range and fluctuation frequency; soil moisture change segments are encoded according to moisture value and rate of change; electromagnetic interference change segments are encoded according to interference intensity and frequency characteristics. The encoding results of each environmental feature segment are concatenated to form an environmental encoding sequence. The operational encoding sequence and the environmental encoding sequence are combined to generate gene encoding segments, each corresponding to a specific fault feature combination.
[0022] Step S116: Associate the gene coding fragment with the corresponding cable segment identifier and fault type information to form a fault feature gene unit. The structure of the fault feature gene unit includes a segment identifier field, a gene coding field, and a fault type field.
[0023] Each gene coding fragment is matched with its corresponding cable segment identifier, which is determined by the segment coding rules of the urban cable network and includes the area code, line number, and segment sequence number. The gene coding fragment, cable segment identifier, and corresponding fault type information are then linked and integrated to form a fault feature gene unit. This fault feature gene unit is stored in a structured data format, where the segment identifier field records the specific cable segment code, the gene coding field stores the gene coding fragment composed of the operational coding sequence and the environmental coding sequence, and the fault type field specifies the name of the corresponding complex fault type and a brief description of the fault. For example, the segment identifier field of a fault feature gene unit might be "Eastern City Area - 10kV East Trunk Line - 08 Segment," the gene coding field might be a combination of the operational coding sequence and the environmental coding sequence, and the fault type field might be "Insulation breakdown caused by intensified partial discharge due to insulation aging" along with a brief description of the fault.
[0024] Step S117: Collect all fault feature gene units, establish association rules between units, the association rules are determined based on the similarity of gene coding fragments of different fault feature gene units, integrate the association rules with the fault feature gene units to form a cable fault feature gene library.
[0025] All generated fault feature gene units are collected, and sequence alignment is used to calculate the similarity of gene coding fragments between different units. For any two fault feature gene units, the base sequences in their gene coding fragments are compared position by position, and the proportion of identical bases is counted to determine the similarity. When the similarity reaches a preset threshold, the two units are considered to be associated. For example, the gene unit corresponding to "insulation breakdown caused by intensified partial discharge due to insulation aging" and the gene unit corresponding to "sheath aging caused by soil acid corrosion" are associated because of the high similarity of soil environmental feature codes in their gene coding fragments. Based on the similarity analysis results between all units, association rules are established to clarify which fault feature gene units are associated and the degree of association. All fault feature gene units and the established association rules are integrated and stored in a database to form a cable fault feature gene library. This cable fault feature gene library supports retrieval and updating by cable section, fault type, gene coding fragment, and other dimensions.
[0026] Step S120: Generate a fault association network topology based on the cable fault feature gene library. The fault association network topology uses fault feature gene units as nodes and the association strength between different fault feature gene units as links. The attributes of the links include the association duration and the association triggering conditions.
[0027] In urban underground cable network scenarios, a fault association network topology is generated based on a constructed cable fault feature gene library. This topology visually displays the relationships between different complex fault features. This process requires transforming fault feature gene units in the gene library into network nodes and constructing links and defining link attributes based on the degree of association between units.
[0028] Step S121: Extract all fault feature gene units from the cable fault feature gene library, and use each fault feature gene unit as the initial node of the fault association network topology. The attributes of each initial node include segment identifier, gene coding sequence and fault type.
[0029] All fault feature gene units are extracted in batches from the cable fault feature gene library, with each unit serving as an initial node in the fault association network topology. Attribute information is configured for each initial node, including a segment identifier attribute recording the cable segment code corresponding to the node, such as "Chengxi Area - 10kV West Trunk Line - Segment 12"; a gene coding sequence attribute storing the operational and environmental coding sequences corresponding to the node; and a fault type attribute specifying the complex fault type corresponding to the node, such as "conductor damage caused by external disturbance combined with material defects" or "sheath failure caused by the combined effects of soil corrosion and electrochemical aging." This attribute information is bound to the initial nodes to ensure that the characteristics of each node are identifiable.
[0030] Step S122: Calculate the gene coding sequence similarity between any two initial nodes. The gene coding sequence similarity is determined based on the proportion of identical bases in the coding fragment. The higher the proportion of identical bases, the higher the gene coding sequence similarity.
[0031] For all extracted initial nodes, the similarity between any two nodes is calculated. During the calculation, the gene coding sequences of the two nodes are divided into segments of equal length. Each segment is compared to check for identical bases at corresponding positions. The number of identical bases in each segment is counted, and the proportion of identical bases in each segment relative to the total number of bases in that segment is calculated. Finally, the average of these proportions is taken as the overall similarity between the gene coding sequences of the two nodes. For example, the gene coding sequences of node A and node B are divided into multiple segments of a preset length. After comparison, the first segment has a certain proportion of identical bases, the second segment has a different proportion, and so on, averaging the proportions of all segments to obtain the similarity. If the two nodes' gene coding sequences have a very high proportion of identical bases in the environmental feature coding portion, but there are some differences in the operational feature coding portion, the overall similarity will still be determined based on the combined proportions; the higher the proportion of identical bases, the greater the final similarity value.
[0032] Step S123: Determine the association strength between two initial nodes based on the similarity of the gene coding sequence. The association strength is positively correlated with the similarity of the gene coding sequence. When the similarity of the gene coding sequence reaches a preset association threshold, a link is established between the two initial nodes.
[0033] The association threshold is determined based on the statistical patterns of faults in urban underground cable networks and expert experience. The calculated similarity of the gene coding sequences of any two initial nodes is compared with the preset association threshold. If the similarity is greater than or equal to the threshold, the two nodes are considered associated. The association strength is directly determined based on the similarity; the higher the similarity, the stronger the association. For example, if the gene coding sequence similarity between nodes C and D reaches a certain value and this value meets the preset association threshold, then the association strength between them is set according to this similarity value, and a link is established between the two nodes to represent their association relationship. If the similarity between nodes E and F does not reach the threshold, no link is established.
[0034] Step S124: Record the time point when the two initial nodes of the established link first appeared to be associated and the time point when they last appeared to be associated. Calculate the difference between the two time points and determine the association duration attribute of the link.
[0035] For two initial nodes in an established link, the time points when they were first identified as having a link and the last time they were identified as having a link are extracted from historical cable fault case data and the association records in the gene pool. The time difference between the last association time point and the first association time point is the association duration attribute of the link. For example, if node G and node H first became associated on a certain day, month, and year, and last became associated on a different day, month, and year, the difference between the two time points is the association duration of the link, and this duration information is bound to the corresponding link.
[0036] Step S125: Analyze the fault types corresponding to the two initial nodes of the established link, determine the environmental conditions and operating parameter conditions that trigger the simultaneous occurrence of the two fault types, and integrate the environmental conditions and operating parameter conditions that trigger the simultaneous occurrence of the two fault types into the associated trigger condition attribute of the link.
[0037] For the fault types corresponding to the two initial nodes in establishing the link, this study analyzes the common conditions under which these two fault types occur simultaneously in historical fault cases. Taking the fault of "insulation breakdown caused by insulation aging accompanied by aggravated partial discharge" corresponding to node I and the fault of "sheath damage caused by soil corrosion" corresponding to node J as examples, a review of historical data reveals that when the soil pH value is below a certain range, the soil moisture content is above a certain range (environmental conditions), and the cable conductor is consistently at over 80% of its rated load and the insulation resistance is below a certain value (operating parameter conditions), both faults are prone to occur simultaneously. Integrating these environmental conditions and operating parameter conditions forms the associated triggering condition attribute of the link, clarifying under what circumstances the fault types corresponding to the two nodes will occur simultaneously.
[0038] Step S126: Label all established links with attributes, including association strength, association duration and association triggering conditions.
[0039] The algorithm iterates through all established links in the fault association network topology, labeling each link with its calculated association strength, duration, and triggering conditions. The labeling is structured to ensure that the attribute information of each link is clearly visible. For example, a link might be labeled with the following: association strength as a value determined based on similarity, association duration as the calculated time length, and triggering conditions as a combination of integrated environmental and operational parameters. This allows for direct access to the link's attributes during subsequent analysis.
[0040] Step S127: Divide the nodes into hierarchical levels according to the cable segment identifier of the initial node. Nodes with the same cable segment identifier are grouped into the same level. Nodes in different levels establish cross-level links based on the association strength, forming a fault association network topology that includes node hierarchy and link attributes.
[0041] Cable segment identifiers of all initial nodes are extracted, and nodes with the same identifier are grouped into the same level, with each level corresponding to a specific segment in the urban underground cable network. For example, all nodes with the segment identifier "Chengnan Area - 10kV South Trunk Line - 05 Segment" are grouped into the same level, while nodes with the identifier "Chengnan Area - 10kV South Trunk Line - 06 Segment" are grouped into another level. For nodes at different levels, if the similarity of their gene coding sequences reaches a preset association threshold, cross-level links are established based on the association strength. Through hierarchical division and cross-level link construction, a fault association network topology containing node hierarchical distribution and link attribute information is finally formed. This fault association network topology can intuitively reflect the fault characteristic association of different cable segments.
[0042] Step S130: Update the node attributes and link attributes in the fault-related network topology based on the real-time collected cable operation data, evolve the fault-related network topology, and record the topology evolution trajectory.
[0043] In the actual operation of urban underground cable networks, the node and link attributes of fault-related network topology are dynamically updated by real-time collected operation data, so that the topology can reflect the current operating status of the cable and the changes in fault association, while recording the evolution process of the topology.
[0044] Step S131: Collect cable operation data in real time. The real-time collected cable operation data includes the current cable conductor temperature data, the current insulation layer resistance data, and the current ambient temperature data.
[0045] A real-time monitoring system deployed throughout the city's underground cable network continuously collects cable operation data. This includes real-time acquisition of conductor temperature data via fiber optic temperature sensors installed at various monitoring points, covering instantaneous temperature values at different locations on the conductor; insulation resistance data collected by online insulation resistance monitoring devices, including instantaneous values of volume resistivity and surface resistivity of the insulation layer; and ambient temperature data collected by temperature sensors in the cable laying environment, encompassing instantaneous ambient temperatures at both the surface and underground cable laying locations. The collected data is transmitted in real-time to the data processing center via a wireless communication module, ensuring data timeliness.
[0046] Step S132: Extract the current operating feature segment from the real-time collected cable operation data, encode the current operating feature segment, and generate the current gene coding sequence.
[0047] The cable operation data transmitted in real time to the data processing center is preprocessed to remove noise. Current operation feature segments are then extracted, including current conductor temperature change segments (reflecting recent conductor temperature fluctuations) and current insulation resistance change segments (reflecting recent insulation resistance trends). Using the same encoding rules as in step S115, these current operation feature segments are encoded to generate a current operation encoding sequence. Simultaneously, current environmental feature segments are extracted and encoded based on real-time collected environmental data (such as current soil moisture and current electromagnetic interference intensity) to form a current environment encoding sequence. The current operation encoding sequence and the current environment encoding sequence are combined to generate the current gene encoding sequence.
[0048] Step S133: Compare the current gene coding sequence with the gene coding sequences of each node in the fault association network topology, and calculate the relative change rate of similarity between the current gene coding sequence and the gene coding sequences of each node.
[0049] The generated current gene coding sequence is compared one by one with the gene coding sequences of all nodes in the fault association network topology. During the comparison, the same segmented comparison method as in step S122 is used. First, the current gene coding sequence and the node gene coding sequences are segmented into segments of the same length. The proportion of identical bases in each segment is calculated, and the average of the proportions of each segment is taken as the current similarity between the two. Simultaneously, historical similarity data between the gene coding sequence and the current gene coding sequence in the node's historical records are retrieved, and the relative rate of change of similarity is calculated. The relative rate of change of similarity is calculated by subtracting the most recent historical similarity from the current similarity, and then dividing the difference by the most recent historical similarity. For example, if the historical similarity of a node in the last comparison was one value, and the current similarity is another value, the relative rate of change of similarity for that node is obtained by subtracting the historical similarity from the current similarity and dividing by the historical similarity. This calculation quantifies the degree of change in the similarity between the current gene coding sequence and the gene coding sequences of each node.
[0050] Step S134: Update the gene coding field of the corresponding node according to the relative change rate of similarity. If the relative change rate of similarity exceeds the preset update threshold, replace the original gene coding sequence of the node with the current gene coding sequence, and at the same time update the fault type field of the node to the fault type corresponding to the current gene coding sequence.
[0051] The preset update threshold is set based on the fault evolution patterns of urban underground cable networks. The calculated relative change rate of similarity for each node is compared with the preset update threshold. If the relative change rate of similarity for a node exceeds the threshold, it indicates a significant change in the operating status of the cable section corresponding to that node, requiring an update to the node's gene coding field. Specifically, the original gene coding sequence of the node is replaced with the current gene coding sequence. Based on the operating and environmental characteristics corresponding to the current gene coding sequence, the corresponding fault type is determined, and the node's fault type field is updated. For example, if the original fault type field for a node is "sheath failure caused by the combined effects of soil corrosion and electrochemical aging," and the relative change rate of similarity exceeds the update threshold, after replacing the gene coding sequence, the fault type field is updated to "sheath failure accompanied by local insulation damage" based on the new sequence characteristics. If the relative change rate of similarity for a node does not exceed the update threshold, the original gene coding field and fault type field of the node remain unchanged.
[0052] Step S135: Calculate the gene coding sequence similarity between the updated node and other nodes, and adjust the association strength attribute of the link between nodes according to the new gene coding sequence similarity. If the new gene coding sequence similarity is higher than the gene coding sequence similarity corresponding to the original association strength, the association strength of the link is increased. If the new gene coding sequence similarity is lower than the gene coding sequence similarity corresponding to the original association strength, the association strength of the link is decreased.
[0053] For nodes whose attributes have been updated, the genetic code sequence similarity between them and all other nodes in the fault association network topology is recalculated. The calculation process still uses segmented comparison and averaging to obtain new genetic code sequence similarities. The new similarity is compared with the genetic code sequence similarity corresponding to the original link strength between this node and other nodes, and the link strength attribute is adjusted based on the comparison result. If the new similarity is higher than the original similarity, it indicates a stronger association between the fault features of the two nodes, and the link strength is increased proportionally to the difference between the new and original similarities. If the new similarity is lower than the original similarity, it indicates a weaker association, and the link strength is decreased proportionally to the difference. For example, if the original link strength similarity between an updated node and node K has a certain value, and the newly calculated similarity has a higher value, then the link strength is increased proportionally to the difference between the two values.
[0054] Step S1351: Extract the gene coding sequence of the updated node, and split the gene coding sequence of the updated node into multiple coding segments, each coding segment corresponding to a running parameter feature or an environmental feature.
[0055] The complete gene coding sequence of the updated node is extracted and divided into multiple coding segments according to the feature type classification rules at the time of encoding. Each coding segment corresponds to a specific operational parameter feature or environmental feature. For example, the gene coding sequence can be divided into conductor temperature coding segments, insulation layer resistance coding segments, partial discharge quantity coding segments, soil moisture coding segments, electromagnetic interference coding segments, etc. During the division, it is necessary to ensure that the length of each coding segment is consistent with the length of the feature fragment at the time of encoding, and that the feature correspondence of each segment is accurate, without overlap or confusion.
[0056] Step S1352: Extract the gene coding sequences of other nodes in the fault-related network topology that have link connections with the updated node, and also split the gene coding sequences of other nodes in the fault-related network topology that have link connections with the updated node into the same number and type of coding segments.
[0057] From the fault-related network topology, identify all other nodes that have established links with the updated node, and extract their gene coding sequences. Following the same splitting rules and feature correspondences as in step S1351, also split the gene coding sequences of these other nodes into multiple coding segments, ensuring that the number, type, and corresponding features of each segment are consistent with those of the updated node. For example, if the updated node is split into 5 coding segments, each corresponding to one of 5 features, then the gene coding sequences of the other nodes must also be split into 5 coding segments of the same feature type.
[0058] Step S1353: Calculate the similarity between the updated node and the corresponding coded segments of other nodes to obtain multiple segment similarity values. Each segment similarity value is determined based on the proportion of identical bases in the corresponding coded segment.
[0059] For each updated node and its corresponding coded segments with other nodes, similarity is calculated pairwise. During the calculation, the number of identical bases in each corresponding segment is counted, and then this number is divided by the total number of bases in that segment. The resulting ratio is the similarity value of that corresponding coded segment. For example, the conductor temperature coded segment of the updated node is compared with the conductor temperature coded segment of node M. The number of identical bases is counted and divided by the total segment length to obtain the conductor temperature segment similarity value. The similarity values for the insulation resistance coded segment, partial discharge quantity coded segment, etc., are calculated in the same way, ultimately yielding multiple segment similarity values.
[0060] Step S1354: Assign a weight coefficient to each segment similarity value based on the feature importance corresponding to different encoded segments. The higher the feature importance, the larger the weight coefficient.
[0061] By analyzing the factors influencing faults in urban underground cable networks, the importance of features corresponding to different coded segments was determined. For example, insulation resistance and partial discharge characteristics have a high impact on fault occurrence, and their corresponding coded segments are assigned a high importance level; conductor temperature and soil moisture characteristics have a moderate impact, and are assigned a medium importance level; electromagnetic interference characteristics have a low impact, and are assigned a low importance level. Weight coefficients are assigned to the similarity values of each segment according to their importance level: high-importance segments are assigned larger weight coefficients, medium-importance segments are assigned medium weight coefficients, and low-importance segments are assigned smaller weight coefficients, with the sum of the weight coefficients of all segments being 1.
[0062] Step S1355: After standardizing the similarity value of each coding segment, calculate the weighted sum based on the weight coefficients to obtain the comprehensive gene coding sequence similarity.
[0063] The calculated similarity values of multiple sub-segments are standardized to convert them to the same numerical range, eliminating potential biases caused by differences in calculation dimensions. After standardization, each sub-segment similarity value is multiplied by its corresponding weight coefficient to obtain a weighted similarity value. The weighted similarity values of all sub-segments are then summed to obtain the overall gene coding sequence similarity between the updated node and other nodes.
[0064] Step S1356: Compare the comprehensive gene coding sequence similarity with the gene coding sequence similarity corresponding to the original association strength, and calculate the similarity difference.
[0065] Retrieve the gene coding sequence similarity corresponding to the association strength of the original links between the updated node and other nodes, and compare it with the newly calculated comprehensive gene coding sequence similarity. Subtract the gene coding sequence similarity corresponding to the original association strength from the comprehensive gene coding sequence similarity; the result is the similarity difference. If the comprehensive similarity is higher than the original similarity, the difference is positive; if the comprehensive similarity is lower than the original similarity, the difference is negative; if the two are equal, the difference is zero.
[0066] Step S1357: If the similarity difference is positive, increase the association strength attribute of the link according to the proportion of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength. The increase proportion is positively correlated with the proportion of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength.
[0067] When the similarity difference is positive, it indicates that the updated node has a stronger association with the fault characteristics of other nodes. The ratio of the similarity difference to the gene coding sequence similarity corresponding to the original association strength is calculated; this ratio represents the increase in the link's association strength. For example, if the similarity difference is one value and the original similarity is another value, the ratio is the increase ratio. The link's association strength attribute is increased proportionally; the larger the increase ratio, the greater the increase in association strength.
[0068] Step S1358: If the similarity difference is negative, reduce the association strength attribute of the link according to the ratio of the absolute value of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength. The reduction ratio is positively correlated with the ratio of the absolute value of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength.
[0069] When the similarity difference is negative, it indicates that the correlation between the updated node and other nodes' fault characteristics has weakened. The ratio of the absolute value of the similarity difference to the gene coding sequence similarity corresponding to the original association strength is calculated; this ratio represents the reduction in the link's association strength. For example, if the absolute value of the similarity difference is one value and the original similarity is another value, the ratio of the two is the reduction ratio. The link's association strength attribute is adjusted downwards according to this ratio; the larger the reduction ratio, the greater the decrease in association strength.
[0070] Step S1359: If the updated node and other nodes originally had no connection links, and the similarity of the comprehensive gene coding sequence reaches a preset association threshold, then a new link is established between the updated node and other nodes, and the value corresponding to the similarity of the comprehensive gene coding sequence is determined as the initial association strength attribute of the new link.
[0071] For the updated node and other nodes in the network topology that were not originally connected by links, if the similarity of their comprehensive gene coding sequences reaches a preset association threshold, it indicates that the fault characteristics of the two nodes are significantly related, and a new link needs to be established between them. The initial association strength attribute of the new link is directly set to the value corresponding to the comprehensive gene coding sequence similarity. At the same time, the association duration and association triggering condition attributes of the new link will be supplemented in subsequent steps S124 and S125.
[0072] Step S136: Monitor the association duration attribute of the link. If the updated node and other nodes still meet the association triggering conditions, extend the association duration. If the updated node and other nodes do not meet the association triggering conditions, terminate the existence of the link and record the termination time.
[0073] Continuously monitor whether the association trigger conditions between the updated node and other nodes are met. Using real-time collected operational and environmental data, determine if the current operational parameters and environmental conditions are consistent with the link association trigger conditions. If the association trigger conditions are still met, it indicates that the association relationship between the link continues to exist. The association duration is extended to the current time, i.e., the last association time is updated to the current time, and the association duration is recalculated. If the association trigger conditions are no longer met, it indicates that the association relationship has been interrupted. The existence of the link is terminated, and the specific time of link termination is recorded. Simultaneously, the link is deleted from the faulty association network topology.
[0074] Step S137: When the node attribute update or link attribute adjustment reaches the preset evolution threshold, record the node distribution, number of links and link attributes of the current fault-related network topology to form an evolution state snapshot.
[0075] The preset evolution threshold can be set as the number of node attribute updates, the magnitude of link attribute adjustments, or the proportion of changes in the number of links. The system monitors changes in the fault-related network topology in real time. When the number of node attribute updates reaches a preset number, the magnitude of link association strength adjustments exceeds a preset range, or the proportion of increases or decreases in the number of links reaches a set value, the evolution threshold is determined to have been reached. At this point, the system records the distribution location and attributes of all nodes in the current network topology, the number of all links, their connection relationships, and the association strength, duration, and triggering conditions of each link. This information is then integrated into an evolution state snapshot and stored in a structured file format.
[0076] Step S138: Arrange all evolution state snapshots in chronological order, and mark the real-time acquisition time point and updated node and link information corresponding to each evolution state snapshot to form the evolution trajectory of the fault-related network topology.
[0077] All generated evolutionary state snapshots are sorted sequentially according to their corresponding real-time acquisition time points to form a time series. Each evolutionary state snapshot is labeled with its corresponding acquisition time point, along with the identifiers and update content of nodes whose attributes were updated, and the identifiers and adjustment content of links whose attributes were adjusted, created, or deleted. This time-sorted and labeled sequence of evolutionary state snapshots is defined as the evolution trajectory of the fault-related network topology. This trajectory fully records the process of network topology changes over time, and the evolution process can be replayed chronologically using visualization tools.
[0078] Step S140: Combine the cable laying path data with the topology evolution trajectory to construct a spatiotemporal coupled positioning network. The spatiotemporal coupled positioning network maps the evolution nodes of the fault association network topology to the correspondence between the physical laying sections of the cable.
[0079] In the context of urban underground cable networks, cable laying path data reflects the physical spatial distribution of cables, while topology evolution trajectory reflects the temporal variation pattern of fault characteristics. Combining the two to construct a spatiotemporal coupled positioning network can realize the spatiotemporal correlation between fault characteristics and physical location.
[0080] Step S141: Obtain cable laying path data, which includes the location coordinates of the physical laying sections of the cable, the length of each section, and the connection relationship between the sections.
[0081] Cable laying path data is obtained through the urban cable operation and maintenance management platform. This data is compiled from engineering records during cable laying and inspection updates during subsequent operation and maintenance. The location coordinates of the physical laying sections are represented by latitude and longitude coordinates, including the starting and ending latitude and longitude coordinates of each section; the section length is the straight-line distance between the starting and ending points; the connection relationships between sections record the coordinates of the connection points between adjacent physical laying sections and the connection method, such as direct connection or connection via intermediate joints. In addition, the cable laying path data also includes auxiliary information such as the material type and laying method (e.g., direct burial, duct laying) for each physical laying section.
[0082] Step S142: Extract all evolution state snapshots from the topology evolution trajectory. Each evolution state snapshot contains the updated fault feature gene unit node and the cable segment identifier corresponding to the updated fault feature gene unit node.
[0083] The algorithm iterates through all evolutionary state snapshots in the topology evolution trajectory, extracting information on all fault feature gene unit nodes contained in each snapshot, with a focus on extracting the cable segment identifiers of each node. Since each evolutionary state snapshot records the network topology state at a specific point in time, the extracted fault feature gene unit nodes are the nodes after attribute updates at that point in time, and their corresponding cable segment identifiers clearly identify the cable segment to which the node belongs. An association index is established between each evolutionary state snapshot and the extracted nodes and their corresponding cable segment identifiers to facilitate subsequent matching operations.
[0084] Step S143: Match the cable segment identifier in each evolution state snapshot with the physical laying segment in the cable laying path data to determine the location coordinates and segment length of the physical laying segment corresponding to each fault feature gene unit node.
[0085] Based on the cable segment identifier of each node in each evolutionary state snapshot, the corresponding physical laying segment is searched in the cable laying path data. By accurately matching the segment identifier, the physical laying segment to which each fault characteristic gene unit node belongs is determined, and then the start coordinates, end coordinates, and length of the physical laying segment are extracted. For example, if the cable segment identifier of a node is "Chengdong Area - 10kV East Trunk Line - 08 Segment", after finding the physical laying segment with the same identifier in the laying path data, its corresponding latitude and longitude coordinates and length information are extracted to complete the mapping between the node and the physical segment.
[0086] Step S1431: Extract cable segment identifiers from all fault characteristic gene unit nodes from the evolution state snapshot to form a segment identifier list.
[0087] For a single evolutionary state snapshot, all fault characteristic gene unit nodes are traversed, and the cable segment identifier field of each node is extracted. These identifiers are then arranged in any order to form a segment identifier list corresponding to that snapshot. If multiple nodes in a snapshot belong to the same cable segment, their corresponding segment identifiers appear repeatedly in the list to reflect the number of fault characteristic nodes in that segment at that point in time.
[0088] Step S1432: Extract the identifier of the physical laying section and its corresponding location coordinates and section length from the cable laying path data, and construct a physical section information table. The fields of the physical section information table include physical section identifier, starting position coordinates, ending position coordinates and section length.
[0089] The cable laying path data is parsed to extract the identifiers of all physical laying sections, as well as the start and end coordinates and section length for each identifier. This information is then organized by field to construct a physical section information table. The physical section identifier field serves as the primary key, ensuring that each identifier uniquely corresponds to one record in the table. The start and end coordinate fields store latitude and longitude data, and the section length field stores distance data.
[0090] Step S1433: Traverse each cable segment identifier in the segment identifier list and search for the physical segment identifier that matches the cable segment identifier in the physical segment information table.
[0091] Each cable segment identifier in the segment identifier list is read one by one. Using this identifier as the query condition, a search is performed in the physical segment information table to find if there is a physical segment identifier that is exactly the same. The search uses an exact match method. If a matching identifier is found, the physical laying segment corresponding to the cable segment identifier is determined; if no matching identifier is found, the cable segment identifier is marked as an abnormal identifier, and subsequent manual verification and correction are required.
[0092] Step S1434: After finding the physical segment identifier that matches the cable segment identifier, extract the start position coordinates, end position coordinates and segment length corresponding to the physical segment identifier.
[0093] Once a physical segment identifier matching the cable segment identifier in the physical segment information table is found, the start and end coordinates, as well as the segment length, are extracted from the record corresponding to that physical segment identifier. This data is temporarily stored and bound to the corresponding cable segment identifier to ensure accurate data attribution.
[0094] Step S1435: Calculate the midpoint coordinates of the physical laying section, wherein the midpoint coordinates are the average of the starting coordinates and the ending coordinates of the physical laying section.
[0095] Based on the starting and ending coordinates of the physical laying section, the midpoint coordinates are calculated. For latitude and longitude coordinates, the average of the starting and ending longitudes is calculated as the midpoint longitude, and the average of the starting and ending latitudes is calculated as the midpoint latitude. The midpoint longitude and midpoint latitude together form the midpoint coordinates of the physical laying section. For example, if the starting coordinates are (longitude 1, latitude 1) and the ending coordinates are (longitude 2, latitude 2), then the midpoint longitude is (longitude 1 + longitude 2) / 2, and the midpoint latitude is (latitude 1 + latitude 2) / 2. The combination of these two coordinates gives the midpoint coordinates.
[0096] Step S1436: Bind the midpoint coordinates, the starting coordinates of the physical laying section, the ending coordinates of the physical laying section, and the length of the physical laying section to the corresponding fault feature gene unit node to form a target mapping relationship.
[0097] The calculated midpoint coordinates, along with the extracted start and end coordinates and segment length, are associated and bound to the attribute information of the corresponding fault feature gene unit nodes to form a target mapping relationship. This target mapping relationship clarifies the physical spatial range and center position of the laying segment corresponding to each fault feature gene unit node.
[0098] Step S1437: Repeat steps S1433 to S1436 for all cable segment identifiers in the segment identifier list. That is, traverse each cable segment identifier in turn, find the matching physical segment identifier in the physical segment information table, extract the corresponding location coordinates and segment length, calculate the midpoint location coordinates, and bind the above spatial information with the corresponding fault feature gene unit node to form a target mapping relationship until all identifiers in the segment identifier list have been processed, ensuring that each fault feature gene unit node can uniquely correspond to the physical laying segment and its spatial parameters in the cable laying path data.
[0099] Step S144: Establish a two-dimensional coordinate system with time as the vertical axis and the location coordinates of the cable laying path as the horizontal axis. Mark the time point corresponding to each evolution state snapshot and the location coordinates of the fault feature gene unit node at that time point in the two-dimensional coordinate system.
[0100] The construction rules for the two-dimensional coordinate system are determined, with the vertical axis representing the time dimension and the horizontal axis representing the spatial dimension, i.e., the position coordinates of the cable laying path. The time range of the vertical axis must cover the start and end times of the topology evolution trajectory, and the position coordinate range of the horizontal axis must cover the minimum and maximum position coordinates of the cable laying path. Time scales are marked on the vertical axis, with the scale interval set according to the acquisition frequency of the evolution state snapshots; the higher the acquisition frequency, the smaller the time scale interval. Position coordinate scales are marked on the horizontal axis, with the scale interval set according to the total length of the cable laying path; the longer the total length, the larger the position coordinate scale interval. For each evolution state snapshot, its corresponding acquisition time point is extracted and mapped to the corresponding scale position on the vertical axis. Simultaneously, the midpoint coordinates of all fault feature gene unit nodes in that snapshot are extracted and mapped to the corresponding scale position on the horizontal axis. Within the two-dimensional coordinate system, each fault feature gene unit node corresponds to a specific coordinate point, with the vertical axis coordinate being the snapshot acquisition time point and the horizontal axis coordinate being the node's midpoint coordinate.
[0101] Step S1441: Determine the vertical axis range and horizontal axis range of the two-dimensional coordinate system. The vertical axis range covers the starting time point to the ending time point of the topology evolution trajectory, and the horizontal axis range covers the minimum position coordinates to the maximum position coordinates of the cable laying path.
[0102] The earliest snapshot of the evolutionary state is extracted from the topological evolution trajectory and used as the starting point of the vertical axis, while the latest snapshot is extracted and used as the ending point. The range of the vertical axis is defined by these two points. Similarly, the starting and ending coordinates of all physical laying sections are extracted from the cable laying path data. The smallest coordinate is selected as the starting point of the horizontal axis, and the largest coordinate as the ending point. The range of the horizontal axis is defined by these two coordinates. To ensure the coordinate system can fully accommodate all data points, a certain amount of redundancy space is reserved on both the vertical and horizontal axes. The size of the redundancy space is determined based on the distribution density of the data points.
[0103] Step S1442: Mark the time scale on the vertical axis of the two-dimensional coordinate system. The interval of the time scale is determined based on the acquisition interval of the evolution state snapshot. The shorter the acquisition interval, the smaller the time scale interval.
[0104] The acquisition interval for snapshots of evolutionary states in the statistical topological evolution trajectory is determined. If the acquisition interval is a fixed value, the time scale interval on the vertical axis is set according to this fixed value; if the acquisition interval is not fixed, the average of all acquisition intervals is taken as the time scale interval. Starting from the initial time point, time nodes are sequentially marked on the vertical axis according to the set time scale interval, with specific time information such as year, month, day, hour, and minute noted next to each time node. Important time nodes, such as those where the attributes of fault characteristic gene unit nodes change significantly, can be highlighted using a special marking style for rapid identification in subsequent analysis.
[0105] Step S1443: Mark the position coordinate scale on the horizontal axis of the two-dimensional coordinate system. The interval of the position coordinate scale is determined based on the total length of the cable laying path. The longer the total length, the larger the interval of the position coordinate scale.
[0106] Calculate the total length of the cable laying path, which is the distance from the starting point (minimum position coordinate) to the ending point (maximum position coordinate) on the horizontal axis. Set the scale interval for the horizontal axis position coordinates based on the total length; use a larger interval for a longer total length and a smaller interval for a shorter total length, ensuring a moderate number of markings on the horizontal axis—neither too many causing congestion nor too few causing insufficient accuracy. Starting from the initial position coordinate, mark the position coordinate nodes sequentially on the horizontal axis according to the set scale intervals, with the specific coordinate value next to each node. Simultaneously, label the corresponding cable segment below the horizontal axis for each position coordinate interval, clearly defining the correspondence between the coordinates and the actual cable segments.
[0107] Step S1444: Extract the acquisition time point of each evolution state snapshot from the topological evolution trajectory, and map the acquisition time point to the corresponding scale position of the vertical axis of the two-dimensional coordinate system.
[0108] Traverse each evolutionary state snapshot in the topological evolution trajectory and read the acquisition time point recorded in its metadata. According to the time scale labeling rules of the vertical axis, determine the corresponding scale position of the acquisition time point on the vertical axis. If the acquisition time point falls exactly on a labeled time scale, then directly use that scale position as the mapping point; if the acquisition time point is between two adjacent time scales, then determine its precise mapping position on the vertical axis through linear interpolation and make a temporary mark at that position.
[0109] Step S1445: Extract the midpoint coordinates of all fault feature gene unit nodes in each evolutionary state snapshot, and map the midpoint coordinates to the corresponding scale position of the horizontal axis of the two-dimensional coordinate system.
[0110] For each evolutionary state snapshot, the midpoint coordinates of all fault feature gene unit nodes are extracted from the target mapping relationship bound to them. According to the position coordinate scale labeling rules of the horizontal axis, the corresponding scale position of each midpoint coordinate on the horizontal axis is determined. If the midpoint coordinate coincides with a certain labeled position coordinate scale, then that scale position is directly used as the mapping point; if it is located between two adjacent position coordinate scales, then its precise mapping position on the horizontal axis is determined by linear interpolation, and a temporary mark is also made.
[0111] Step S1446: In a two-dimensional coordinate system, each fault feature gene unit node corresponds to a coordinate point. The vertical axis coordinate of this coordinate point is the acquisition time point of the evolution state snapshot, and the horizontal axis coordinate of this coordinate point is the midpoint position coordinate of the fault feature gene unit node.
[0112] By combining the vertical and horizontal axis mapping positions of each fault feature gene unit node, a unique coordinate point is determined in the two-dimensional coordinate system. This coordinate point represents the node's position in the spatiotemporal dimension. For example, the acquisition time point of a snapshot of an evolutionary state is mapped to a certain position on the vertical axis, and the midpoint coordinate of a fault feature gene unit node in that snapshot is mapped to a certain position on the horizontal axis. The intersection of these two positions is the coordinate point corresponding to the node. This operation is repeated for all fault feature gene unit nodes in each evolutionary state snapshot to complete the labeling of all nodes in the two-dimensional coordinate system.
[0113] Step S1447: Mark each coordinate point with a marker symbol of a preset shape. The color of the marker symbol is determined based on the fault type of the fault feature gene unit node. Different fault types correspond to different colors. The size of the marker symbol is determined based on the association strength of the fault feature gene unit node. The higher the association strength, the larger the marker symbol.
[0114] Multiple marker symbols, such as circles, squares, and triangles, are pre-defined, allowing selection of the appropriate symbol shape based on the cable line type to which the fault characteristic gene unit node belongs. Different colors are assigned to different fault types; for example, red corresponds to "insulation breakdown caused by insulation aging accompanied by increased partial discharge," blue to "conductor damage caused by external disturbance combined with material defects," and yellow to "sheath failure caused by the combined effects of soil corrosion and electrochemical aging," ensuring clear color differentiation and preventing confusion. The size of the marker symbols is set according to the correlation strength of the fault characteristic gene unit nodes, with the largest symbol corresponding to the highest correlation strength, the smallest symbol to the lowest correlation strength, and a medium-sized symbol corresponding to intermediate correlation strength. The symbol size visually reflects the differences in node correlation strength. The pre-defined marker symbols, colors, and sizes are applied to the corresponding coordinate points in a two-dimensional coordinate system, completing the visual annotation of the coordinate points.
[0115] Step S145: Connect the position coordinates of the same fault feature gene unit node at different time points to form the spatiotemporal trajectory line of the fault feature gene unit node. The thickness of the spatiotemporal trajectory line is determined based on the association strength attribute of the fault feature gene unit node. The higher the association strength, the thicker the trajectory line.
[0116] In a two-dimensional coordinate system, all coordinate points belonging to the same fault feature gene unit node are selected. These coordinate points correspond to the acquisition time points of snapshots of different evolutionary states. These coordinate points are connected sequentially with straight lines in chronological order, forming a continuous line. This line is the spatiotemporal trajectory of the fault feature gene unit node. The direction of the trajectory reflects the changes in the fault feature corresponding to the node in the time dimension and its distribution in the spatial dimension. The association strength attribute of the fault feature gene unit node is extracted. The thickness parameter of the spatiotemporal trajectory line is set according to the magnitude of the association strength; the larger the association strength value, the wider the trajectory line; the smaller the association strength value, the narrower the trajectory line. By observing the changes in the thickness of the trajectory line, the dynamic changes in the association strength of the fault feature of the node can be intuitively judged.
[0117] Step S146: Mark the connection relationship of each physical laying section in the two-dimensional coordinate system, and connect the position coordinates of adjacent physical laying sections with line segments to form the spatial baseline of the cable laying path.
[0118] The starting and ending coordinates of each physical laying segment, as well as the connection information between adjacent physical laying segments, are extracted from the cable laying path data. On the horizontal axis of the two-dimensional coordinate system, the points corresponding to the starting and ending coordinates of each physical laying segment are located, and these two points are connected by a line segment, which represents the spatial position of the physical laying segment. Based on the connection relationship between adjacent physical laying segments, the point corresponding to the ending coordinate of the previous physical laying segment is connected to the point corresponding to the starting coordinate of the next adjacent physical laying segment by a line segment, and so on, until all physical laying segments are connected according to the connection relationship, forming a complete spatial baseline for the cable laying path. The spatial baseline uses a fixed color and thickness to clearly distinguish it from the spatiotemporal trajectory lines of the fault characteristic gene unit nodes, facilitating clear identification of the cable's physical laying path.
[0119] Step S147: Overlay the spatiotemporal trajectory lines with the spatial baseline, label the fault type of the fault feature gene unit node corresponding to each spatiotemporal trajectory line, and form a spatiotemporal coupled positioning network containing time dimension, spatial dimension and fault information.
[0120] The spatiotemporal trajectories of all completed fault feature gene unit nodes are overlaid with the spatial baseline of the cable laying path in the same two-dimensional coordinate system. This ensures a complete correspondence between the positional coordinates of the spatiotemporal trajectories and the spatial baseline, achieving precise matching between the spatiotemporal information of fault features and the physical path of the cable. At the start, end, and key turning points of each spatiotemporal trajectory line, the fault type name of the corresponding fault feature gene unit node is labeled. The color of the labeled text matches the color of the corresponding marker symbol for the trajectory line, facilitating rapid association between the trajectory line and the fault type. After overlay, the resulting spatiotemporally coupled positioning network simultaneously contains fault evolution information in the time dimension, cable location information in the spatial dimension, and specific fault type information. This can be displayed through a visual interface, supporting intuitive analysis of the spatiotemporal distribution and evolution process of fault features.
[0121] Step S150: Determine the physical laying section where the cable fault is located based on the spatiotemporal coupling positioning network, generate a cable fault location instruction containing section identifier and fault characteristic gene unit association information, and transmit the cable fault location instruction to the cable operation and maintenance terminal.
[0122] Based on the spatiotemporal and fault information integrated by the spatiotemporal coupled positioning network, the physical laying section where the fault is located is accurately located by analyzing the clustering and evolution trend of fault features and their correspondence with the physical sections of the cable. Then, standardized fault location instructions are generated and transmitted to the operation and maintenance terminal in a timely manner to guide the on-site operation and maintenance work.
[0123] Step S151: Analyze the spatiotemporal trajectory lines in the spatiotemporal coupled positioning network, and identify the clustering areas where the spatiotemporal trajectory lines are clustered. The clustering area is the overlapping area of multiple spatiotemporal trajectory lines within the same horizontal axis coordinate range and similar vertical axis coordinate range.
[0124] A comprehensive analysis of all spatiotemporal trajectories in the spatiotemporal coupled positioning network was conducted, focusing on the spatial and temporal distribution characteristics of the trajectories. Detection revealed that when multiple spatiotemporal trajectories have similar or overlapping position coordinates on the horizontal axis and similar time coordinates on the vertical axis, these trajectories form overlapping or intersecting areas, which are termed spatiotemporal trajectory clusters. The formation of clusters typically indicates that multiple or repeated fault characteristics have occurred in a specific cable physical section within a particular time frame, representing a high-risk area for fault occurrence. During the identification process, the horizontal axis coordinate range, the vertical axis time range, and the number of spatiotemporal trajectories contained within the clustered area must be recorded.
[0125] For example, step S1511: Divide the two-dimensional coordinate system of the spatiotemporal coupled positioning network into multiple grid cells, each grid cell corresponding to a fixed range of horizontal axis coordinates and a fixed range of vertical axis coordinates.
[0126] Based on the horizontal axis coordinate range and vertical axis time range of the two-dimensional coordinate system of the spatiotemporal coupled positioning network, the size parameters of the grid cells are set. The horizontal axis coordinate range and vertical axis time range of the grid cells are determined according to the data density. The higher the data density, the smaller the grid cells are to ensure recognition accuracy; the lower the data density, the larger the grid cells are to improve recognition efficiency. According to the set size parameters, the entire two-dimensional coordinate system is uniformly divided into multiple non-overlapping grid cells that completely cover the coordinate system range. Each grid cell has a unique number, and the numbering rule is arranged sequentially from left to right and from top to bottom. Simultaneously, the specific horizontal axis coordinate start value, horizontal axis coordinate end value, vertical axis time start value, and vertical axis time end value corresponding to each grid cell are recorded.
[0127] Step S1512: Count the number of spatiotemporal trajectory line segments contained in each grid cell, wherein the spatiotemporal trajectory line segment is the part of the spatiotemporal trajectory line that passes through the grid cell.
[0128] Traverse each grid cell and identify all spatiotemporal trajectory lines passing through that grid cell. The portion of each spatiotemporal trajectory line that crosses a grid cell is considered a spatiotemporal trajectory line segment. For a spatiotemporal trajectory line entirely within a single grid cell, the entire trajectory line is considered a single trajectory line segment; for a spatiotemporal trajectory line spanning multiple grid cells, the portion of the trajectory line within each grid cell is considered an independent trajectory line segment. Count the number of all spatiotemporal trajectory line segments within each grid cell and store the results in association with the grid cell number.
[0129] Step S1513: Calculate the normalized ratio of the number of spatiotemporal trajectory line segments in each grid cell to the time span and spatial span of the grid cell, and use it as the trajectory line density index.
[0130] Extract the time span (i.e., the difference between the end value and the start value of the time axis on the vertical axis) and the spatial span (i.e., the difference between the end value and the start value of the horizontal axis on the horizontal axis) of each grid cell. Divide the number of spatiotemporal trajectory line segments counted within the grid cell by the time span and the spatial span respectively to obtain two intermediate ratios. Then normalize these two intermediate ratios, converting them to a numerical range of 0 to 1. Take the average of the two normalized intermediate ratios; the result is the trajectory line density index of the grid cell, which reflects the density of spatiotemporal trajectory lines per unit time and unit space.
[0131] Step S1514: Set a trajectory line density threshold, filter out grid cells whose trajectory line density exceeds the trajectory line density threshold, and mark them as target grid cells.
[0132] Based on historical fault statistics and maintenance experience of the urban underground cable network, a trajectory line density threshold is set. This threshold must comprehensively consider the spatiotemporal distribution density of trajectory lines under normal operating conditions to ensure that normal trajectory line distributions are not misjudged as clusters, nor are genuine fault clusters overlooked. The trajectory line density index of each grid cell is compared with the preset threshold. If the trajectory line density index of a grid cell exceeds the threshold, the grid cell is marked as a target grid cell, indicating a dense distribution of trajectory lines within that cell.
[0133] Step S1515: Analyze the spatial distribution of the target mesh cells and group adjacent target mesh cells into the same mesh cluster. Adjacent means that the mesh cells have a common boundary in the horizontal axis direction or a common boundary in the vertical axis direction.
[0134] Spatial distribution analysis is performed on all marked target mesh cells to determine whether adjacency exists between them. Two target mesh cells are considered adjacent if they share a common boundary along the horizontal axis or the vertical axis. Adjacent target mesh cells are grouped into a single mesh cluster, while individual target mesh cells not adjacent to any other target mesh cells are treated as separate mesh clusters. Each mesh cluster is assigned a unique cluster identifier to distinguish different clustering region candidates.
[0135] Step S1516: Calculate the horizontal axis span and vertical axis span of each grid cluster. The horizontal axis span is the difference between the maximum horizontal axis coordinate of the grid cell in the grid cluster and the minimum horizontal axis coordinate of the grid cell in the grid cluster. The vertical axis span is the difference between the maximum vertical axis coordinate of the grid cell in the grid cluster and the minimum vertical axis coordinate of the grid cell in the grid cluster.
[0136] For each grid cluster, iterate through all its grid cells and extract the starting and ending values of the horizontal axis coordinate, as well as the starting and ending values of the vertical axis time. Find the maximum horizontal axis coordinate value (i.e., the maximum value among all horizontal axis coordinate ending values) and the minimum horizontal axis coordinate value (i.e., the minimum value among all horizontal axis coordinate starting values) of all grid cells within the cluster; the difference between these two values is the horizontal axis span of the grid cluster. Similarly, find the maximum vertical axis time value (i.e., the maximum value among all vertical axis time ending values) and the minimum vertical axis time value (i.e., the minimum value among all vertical axis time starting values) of all grid cells within the cluster; the difference between these two values is the vertical axis span of the grid cluster.
[0137] Step S1517: Select grid clusters whose horizontal axis span and vertical axis span are both less than the span threshold after normalization, and mark them as candidate clusters.
[0138] The horizontal and vertical spans of each grid cluster are normalized to a range of 0 to 1, the same as the trajectory line density index. A preset span threshold is set based on the length of the cable's physical laying section and the typical time period of fault evolution to exclude grid clusters with excessively large spans that do not have practical fault aggregation significance. The normalized horizontal and vertical spans are compared with the span threshold. If both are less than the threshold, the grid cluster is marked as a candidate cluster, indicating that the trajectory line aggregation within the cluster has a clear spatiotemporal range and may correspond to a specific fault area.
[0139] Step S1518: Extract the fault feature gene unit nodes corresponding to all spatiotemporal trajectory line segments in the candidate cluster, and count the fault type consistency of the fault feature gene unit nodes. The fault type consistency is the proportion of the number of fault feature gene unit nodes with the same fault type to the total number of all fault feature gene unit nodes in the candidate cluster, and the total number of nodes in the candidate cluster reaches the minimum statistical sample size requirement.
[0140] From all grid cells contained in the candidate cluster, extract the identifier and fault type information of the fault feature gene unit node corresponding to each spatiotemporal trajectory line segment. Count the total number of fault feature gene unit nodes within the candidate cluster. If the total number does not meet the preset minimum statistical sample size requirement, the candidate cluster is excluded, as a small sample size may lead to unrepresentative analysis results. For candidate clusters that meet the minimum statistical sample size requirement, count the number of nodes belonging to the same fault type. Divide this number by the total number of nodes; the resulting ratio represents the fault type consistency.
[0141] Step S1519: Select candidate clusters whose fault type consistency exceeds the consistency threshold. The range of the grid cells corresponding to the candidate clusters is taken as the region where the spatiotemporal trajectory lines appear to cluster.
[0142] A consistency threshold is set, which is determined based on the accuracy requirements of fault diagnosis and is typically set to a high value to ensure strong consistency of fault characteristics within the clustered area. The fault type consistency of each candidate cluster is compared with the consistency threshold. If the fault type consistency of a candidate cluster is greater than the consistency threshold, the spatiotemporal trajectory line cluster within that cluster is determined to have a clear fault orientation. The horizontal axis coordinate range and vertical axis time range of all grid cells corresponding to that candidate cluster are integrated to determine the region where the spatiotemporal trajectory lines cluster, and the boundary coordinates and included grid cell numbers of this clustered region are recorded. If the fault type consistency of a candidate cluster does not exceed the consistency threshold, it indicates that the fault characteristics within that cluster are relatively dispersed and do not have clear fault location value, and are therefore excluded.
[0143] Step S152: Extract the fault feature gene unit nodes corresponding to all spatiotemporal trajectory lines within the aggregation area, and collect the fault type, correlation strength, and corresponding physical laying section identifier of the fault feature gene unit nodes.
[0144] The algorithm iterates through all grid cells within the spatiotemporal trajectory line cluster area, tracing the corresponding fault feature gene unit node from the spatiotemporal trajectory line segments contained in each grid cell. For each extracted fault feature gene unit node, it retrieves its fault type information, correlation strength attribute, and physical laying section identifier bound through target mapping relationship from the cable fault feature gene library and fault association network topology. This information is then summarized and organized to form a cluster area node information table. This table includes fields such as node identifier, fault type, correlation strength value, and physical laying section identifier, ensuring that the key information of each node is complete and searchable. For example, a certain cluster area may contain multiple nodes corresponding to the fault "insulation breakdown caused by insulation aging accompanied by increased partial discharge," with different correlation strength values and corresponding physical laying section identifiers concentrated in "Chengdong Area - 10kV East Trunk Line - 08 Section" and "Chengdong Area - 10kV East Trunk Line - 09 Section."
[0145] Step S153: Count the frequency of occurrence of each fault type in the cluster area, and determine the fault type with the highest frequency as the main fault type.
[0146] Based on the node information table of the clustered region, the fault types of all fault characteristic gene unit nodes recorded therein are classified and statistically analyzed. Fault type names are grouped, and the number of nodes in each group is calculated; this number represents the frequency of occurrence of the corresponding fault type. All fault types are sorted from highest to lowest frequency, and the fault type with the highest frequency is selected as the primary fault type for that clustered region. If multiple fault types have the same highest frequency, the sum of the node association strengths corresponding to these fault types is further compared, and the fault type with the largest sum of association strengths is determined as the primary fault type. For example, in a certain clustered region, the fault "insulation breakdown caused by intensified partial discharge due to insulation aging" has the largest number of nodes, and its sum of association strengths is higher than other fault types; therefore, it is determined as the primary fault type.
[0147] Step S154: Calculate the ratio of the number of nodes in each physical laying section within the aggregation area to the total number of nodes in the aggregation area, and determine the physical laying section with the highest ratio as the candidate fault section.
[0148] First, the total number of fault characteristic gene unit nodes within the cluster area is counted, i.e., the total number of records in the cluster area node information table. Then, the clusters are grouped according to physical laying segment identifiers, and the number of nodes corresponding to each physical laying segment is calculated. The number of nodes in each physical laying segment is divided by the total number of nodes to obtain the node proportion of that physical laying segment. All physical laying segments are sorted from highest to lowest according to their node proportion, and the physical laying segment with the highest proportion is selected as the fault candidate segment. If multiple physical laying segments have the same and highest node proportion, the average correlation strength of the nodes within these segments is compared, and the physical laying segment with the highest average correlation strength is determined as the fault candidate segment. For example, if the total number of nodes in the cluster area is several, and the node proportion corresponding to "Chengdong Area - 10kV East Trunk Line - 08 Section" is the highest, then it is determined as the fault candidate segment.
[0149] Step S155: Extract the starting position coordinates, ending position coordinates, and length of the fault candidate segment, and simultaneously extract the gene coding sequences of all fault feature gene unit nodes within the fault candidate segment.
[0150] Based on the identified fault candidate segments, the corresponding record is searched in the physical segment information table of the cable laying path data to extract the start and end coordinates and segment length data of the fault candidate segment. Simultaneously, nodes whose physical laying segment identifiers match the fault candidate segment identifiers are selected from the clustered area node information table. The gene coding sequences of these nodes, including operational and environmental coding sequences, are retrieved from the cable fault feature gene library. The extracted spatial parameters of the fault candidate segments are then associated and stored with the node gene coding sequences.
[0151] Step S156: Integrate the main fault type, the location information of the fault candidate segment, the segment identifier of the fault candidate segment, and the gene coding sequence of the corresponding fault feature gene unit node to form core fault location information.
[0152] The identified main fault type names, the starting and ending coordinates of the candidate fault segments, the segment length, the segment identifier, and the gene coding sequences of all fault characteristic gene unit nodes within the candidate fault segments are integrated according to a preset structured format. The integrated core fault location information must contain a clear hierarchical relationship, such as being divided into three parts: fault type information, segment spatial information, and node coding information. Each part contains specific parameter content. For example, the fault type information section records "insulation breakdown caused by insulation aging accompanied by increased partial discharge," the segment spatial information section records the starting coordinates, ending coordinates, and length of "Chengdong Area - 10kV East Trunk Line - 08 Segment," and the node coding information section records the operational coding sequence and environmental coding sequence of each node within the segment.
[0153] Step S157: Perform structured processing on the fault location core information, and encapsulate the structured fault location core information into a cable fault location instruction.
[0154] The core fault location information is structured using a standardized instruction format, such as XML or JSON, defining uniform tags and data types for each information field to ensure the standardization and readability of the instruction content. After structuring, the structured core fault location information is encapsulated into cable fault location instructions according to the communication protocol of the cable maintenance system. During encapsulation, instruction header information needs to be added, including instruction number, generation time, and data checksum. The data checksum is used by the receiving end to verify the integrity and accuracy of the instruction. For example, the instruction header contains a unique instruction number, the specific time the instruction was generated, and a checksum calculated using a preset algorithm; the instruction body contains the structured core fault location information.
[0155] Step S158: Transmit the cable fault location command to the cable maintenance terminal.
[0156] The packaged cable fault location command is transmitted to the cable maintenance terminal via wired or wireless communication networks. An encrypted transmission protocol is used during transmission to encrypt the command data, preventing the data from being stolen or tampered with during transmission. Upon receiving the command, the cable maintenance terminal verifies its integrity using the data checksum in the command header. After successful verification, it parses the core fault location information in the command body and displays the main fault types, the location of candidate fault sections, and their corresponding genetic sequences on the terminal interface, providing maintenance personnel with clear fault location information and troubleshooting directions.
[0157] Figure 2 The illustration shows exemplary hardware and software components of a big data-based cable fault location system 100 that can implement the ideas of this application, according to some embodiments of this application. For example, a processor 120 can be used in the big data-based cable fault location system 100 and to perform the functions in this application.
[0158] The big data-based cable fault location system 100 can be a general-purpose server or a special-purpose server; both can be used to implement the big data-based cable fault location method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the load.
[0159] For example, the big data-based cable fault location system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the big data-based cable fault location system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to the program instructions. The big data-based cable fault location system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0160] For ease of explanation, only one processor is described in the big data-based cable fault location system 100. However, it should be noted that the big data-based cable fault location system 100 of this application may also include multiple processors, and therefore the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the big data-based cable fault location system 100 performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0161] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned cable fault location method based on big data is implemented.
[0162] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A cable fault location method based on big data, characterized in that, The method includes: A cable fault feature gene library is constructed, which contains fault feature gene units extracted from multi-source cable operation data. The fault feature gene units are associated with cable operation parameter characteristics, laying environment characteristics and historical fault characteristics, and each fault feature gene unit has a unique cable section identifier. Based on the cable fault feature gene library, a fault association network topology is generated. The fault association network topology uses fault feature gene units as nodes and the association strength between different fault feature gene units as links. The attributes of the links include the association duration and the association triggering conditions. The node and link attributes in the fault-related network topology are updated based on the real-time collected cable operation data, the fault-related network topology is evolved, and the topology evolution trajectory is recorded. By combining cable laying path data with the topology evolution trajectory, a spatiotemporal coupled positioning network is constructed. The spatiotemporal coupled positioning network maps the evolution nodes of the fault association network topology and the correspondence between the physical laying sections of the cable. Based on the spatiotemporal coupling positioning network, the physical laying section where the cable fault is located is determined, a cable fault location instruction containing section identifier and fault characteristic gene unit association information is generated, and the cable fault location instruction is transmitted to the cable operation and maintenance terminal.
2. The cable fault location method based on big data according to claim 1, characterized in that, The construction of the cable fault feature gene library includes: Acquire multi-source cable operation data, which includes real-time cable operation parameter data, cable laying environment data, and historical cable fault case data. Extract parameter segments reflecting the cable's operating status from the real-time operating parameter data of the cable. The parameter segments include conductor temperature change segments, insulation layer resistance change segments, and current fluctuation segments. Each parameter segment is marked as an operating characteristic segment. Environmental segments affecting the cable condition are extracted from the cable laying environment data. These environmental segments include environmental temperature change segments, soil moisture change segments, and surrounding electromagnetic interference change segments. Each environmental segment is marked as an environmental feature segment. Extract the operational and environmental feature segments corresponding to the time of the fault from the historical cable fault case data, and bind the fault type information with the corresponding operational and environmental feature segments to form a fault feature combination; The operational and environmental feature segments in each fault feature combination are encoded to generate gene coding segments containing operational and environmental coding sequences; The gene coding fragment is associated with the corresponding cable segment identifier and fault type information to form a fault feature gene unit. The structure of the fault feature gene unit includes a segment identifier field, a gene coding field, and a fault type field. Collect all fault feature gene units, establish association rules between units, the association rules are determined based on the similarity of gene coding fragments of different fault feature gene units, and integrate the association rules with the fault feature gene units to form a cable fault feature gene library.
3. The cable fault location method based on big data according to claim 1, characterized in that, The generation of the fault association network topology based on the cable fault feature gene library includes: All fault feature gene units are extracted from the cable fault feature gene library. Each fault feature gene unit is used as the initial node of the fault association network topology. The attributes of each initial node include segment identifier, gene coding sequence and fault type. Calculate the gene coding sequence similarity between any two initial nodes. The gene coding sequence similarity is determined based on the proportion of identical bases in the coding fragments. The higher the proportion of identical bases, the higher the gene coding sequence similarity. The association strength between two initial nodes is determined based on the similarity of the gene coding sequence. The association strength is positively correlated with the similarity of the gene coding sequence. When the similarity of the gene coding sequence reaches a preset association threshold, a link is established between the two initial nodes. Record the time point when the two initial nodes of the established link first become associated and the time point when they last become associated. Calculate the difference between the two time points and determine the link's association duration attribute. Analyze the fault types corresponding to the two initial nodes of the link, determine the environmental conditions and operating parameters that trigger the simultaneous occurrence of the two fault types, and integrate the environmental conditions and operating parameters that trigger the simultaneous occurrence of the two fault types into the associated trigger condition attribute of the link. All established links are labeled with attributes, including association strength, association duration, and association triggering conditions; The nodes are hierarchically divided according to the cable segment identifier of the initial node. Nodes with the same cable segment identifier are grouped into the same level. Nodes in different levels establish cross-level links based on the association strength, forming a fault association network topology that includes node hierarchy and link attributes.
4. The cable fault location method based on big data according to claim 1, characterized in that, The process of updating node and link attributes in the fault-related network topology based on real-time collected cable operation data, evolving the fault-related network topology, and recording the topology evolution trajectory includes: Real-time acquisition of cable operation data, including current cable conductor temperature data, current insulation layer resistance data, and current ambient temperature data; Extract the current operating feature segment from the real-time collected cable operation data, encode the current operating feature segment, and generate the current gene coding sequence; The current gene coding sequence is compared with the gene coding sequences of each node in the fault association network topology, and the relative change rate of similarity between the current gene coding sequence and the gene coding sequences of each node is calculated. The gene coding field of the corresponding node is updated according to the relative change rate of similarity. If the relative change rate of similarity exceeds the preset update threshold, the original gene coding sequence of the node is replaced with the current gene coding sequence, and the fault type field of the node is updated to the fault type corresponding to the current gene coding sequence. Calculate the gene coding sequence similarity between the updated node and other nodes, and adjust the association strength attribute of the link between nodes according to the new gene coding sequence similarity. If the new gene coding sequence similarity is higher than the gene coding sequence similarity corresponding to the original association strength, the association strength of the link is increased. If the new gene coding sequence similarity is lower than the gene coding sequence similarity corresponding to the original association strength, the association strength of the link is decreased. Monitor the association duration attribute of the link. If the updated node and other nodes still meet the association triggering conditions, extend the association duration. If the updated node and other nodes do not meet the association triggering conditions, terminate the link and record the termination time. When node attribute updates or link attribute adjustments reach a preset evolution threshold, record the node distribution, number of links, and link attributes of the current fault-related network topology to form an evolution state snapshot. Arrange all evolution state snapshots in chronological order, and label the real-time acquisition time point and updated node and link information corresponding to each evolution state snapshot to form the evolution trajectory of the fault-related network topology.
5. The cable fault location method based on big data according to claim 4, characterized in that, The calculation of the gene coding sequence similarity between the updated node and other nodes, and the adjustment of the association strength attribute of the links between nodes based on the new similarity, includes: Extract the gene coding sequence of the updated node, and split the gene coding sequence of the updated node into multiple coding segments, each coding segment corresponding to a running parameter feature or an environmental feature. Extract the gene coding sequences of other nodes in the fault-related network topology that have link connections with the updated node, and also split the gene coding sequences of other nodes in the fault-related network topology that have link connections with the updated node into the same number and type of coding segments. Calculate the similarity between the updated node and the corresponding coded segments of other nodes to obtain multiple segment similarity values. Each segment similarity value is determined based on the proportion of identical bases in the corresponding coded segment. Based on the feature importance corresponding to different coded segments, a weight coefficient is assigned to each segment similarity value. The higher the feature importance, the larger the weight coefficient. After standardizing the similarity values of each coding segment, a weighted sum is calculated based on the weight coefficients to obtain the overall gene coding sequence similarity. The similarity difference is calculated by comparing the overall gene coding sequence similarity with the gene coding sequence similarity corresponding to the original association strength. If the similarity difference is positive, the association strength attribute of the link is increased according to the proportion of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength. The increase ratio is positively correlated with the proportion of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength. If the similarity difference is negative, the association strength attribute of the link is reduced according to the ratio of the absolute value of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength. The reduction ratio is positively correlated with the ratio of the absolute value of the similarity difference to the similarity of the gene coding sequence corresponding to the original association strength. If the updated node had no connection links with other nodes before, and the similarity of the comprehensive gene coding sequence reached a preset association threshold, then a new link was established between the updated node and other nodes, and the value corresponding to the similarity of the comprehensive gene coding sequence was determined as the initial association strength attribute of the new link.
6. The cable fault location method based on big data according to claim 1, characterized in that, The construction of a spatiotemporally coupled positioning network by combining cable laying path data with the topology evolution trajectory includes: Acquire cable laying path data, which includes the location coordinates of the physical laying sections of the cable, the length of each section, and the connection relationships between the sections. Extract all evolution state snapshots from the topological evolution trajectory. Each evolution state snapshot contains the updated fault feature gene unit node and the cable segment identifier corresponding to the updated fault feature gene unit node. Match the cable segment identifier in each evolution state snapshot with the physical laying segment in the cable laying path data to determine the location coordinates and segment length of the physical laying segment corresponding to each fault feature gene unit node; A two-dimensional coordinate system is established with time as the vertical axis and the location coordinates of the cable laying path as the horizontal axis. In the two-dimensional coordinate system, the time point corresponding to each evolution state snapshot and the location coordinates of the fault characteristic gene unit node at that time point are marked. Connecting the position coordinates of the same fault feature gene unit node at different time points forms the spatiotemporal trajectory line of the fault feature gene unit node. The thickness of the spatiotemporal trajectory line is determined based on the association strength attribute of the fault feature gene unit node. The higher the association strength, the thicker the trajectory line. Mark the connection relationship of each physical laying section in a two-dimensional coordinate system, and connect the position coordinates of adjacent physical laying sections with line segments to form the spatial baseline of the cable laying path; By overlaying the spatiotemporal trajectory lines with the spatial baseline, the fault type of each fault feature gene unit node corresponding to the spatiotemporal trajectory line is labeled, forming a spatiotemporal coupled positioning network that includes time dimension, spatial dimension and fault information.
7. The cable fault location method based on big data according to claim 6, characterized in that, The step of matching the cable segment identifier in each evolution state snapshot with the physical laying segment in the cable laying path data to determine the location coordinates and segment length of the physical laying segment corresponding to each fault feature gene unit node includes: Extract cable segment identifiers from all fault characteristic gene unit nodes from the evolutionary state snapshot to form a segment identifier list; Extract the identifiers and corresponding location coordinates and section lengths of the physical laying sections from the cable laying path data, and construct a physical section information table. The fields of the physical section information table include the physical section identifier, the starting location coordinates, the ending location coordinates, and the section length. Iterate through each cable segment identifier in the segment identifier list and search for the physical segment identifier that matches the cable segment identifier in the physical segment information table; After finding the physical segment identifier that matches the cable segment identifier, extract the start position coordinates, end position coordinates, and segment length corresponding to the physical segment identifier; Calculate the coordinates of the midpoint of the physical laying section, wherein the coordinates of the midpoint are the average of the coordinates of the starting position and the coordinates of the ending position of the physical laying section. The midpoint coordinates, the starting coordinates, the ending coordinates, and the length of the physical laying section are bound to the corresponding fault feature gene unit nodes to form a target mapping relationship. The process involves repeatedly extracting cable segment identifiers from all fault feature gene unit nodes in the evolution state snapshot to form a segment identifier list; extracting the identifiers of physical laying segments and their corresponding location coordinates and segment lengths from the cable laying path data to construct a physical segment information table; traversing each cable segment identifier in the segment identifier list to find a matching physical segment identifier in the physical segment information table; extracting the corresponding starting location coordinates and segment length after finding a matching physical segment identifier; calculating the midpoint location coordinates of the physical laying segment; and binding the midpoint location coordinates and segment length with the corresponding fault feature gene unit node to form a target mapping relationship, so that each fault feature gene unit node corresponds to a unique physical laying segment's location coordinates and segment length.
8. The cable fault location method based on big data according to claim 6, characterized in that, A two-dimensional coordinate system is established with time as the vertical axis and the location coordinates of the cable laying path as the horizontal axis. In this system, the time point corresponding to each evolutionary state snapshot and the location coordinates of the fault characteristic gene unit node at that time point are marked, including: Determine the range of the vertical axis and the horizontal axis of the two-dimensional coordinate system. The vertical axis range covers the time from the start point of the topological evolution trajectory to the time from the end point of the topological evolution trajectory, and the horizontal axis range covers the minimum position coordinate of the cable laying path to the maximum position coordinate of the cable laying path. Time scales are marked on the vertical axis of a two-dimensional coordinate system. The interval of the time scale is determined based on the acquisition interval of the evolution state snapshot. The shorter the acquisition interval, the smaller the time scale interval. The position coordinate scale is marked on the horizontal axis of the two-dimensional coordinate system. The interval of the position coordinate scale is determined based on the total length of the cable laying path. The longer the total length, the larger the interval of the position coordinate scale. Extract the acquisition time point of each evolutionary state snapshot from the topological evolution trajectory, and map the acquisition time point to the corresponding scale position of the vertical axis of the two-dimensional coordinate system; Extract the midpoint coordinates of all fault feature gene unit nodes in each evolution state snapshot, and map these midpoint coordinates to the corresponding scale position of the horizontal axis of the two-dimensional coordinate system. In a two-dimensional coordinate system, each fault feature gene unit node corresponds to a coordinate point. The vertical axis of this coordinate point is the acquisition time point of the evolution state snapshot, and the horizontal axis of this coordinate point is the midpoint position coordinate of the fault feature gene unit node. Each coordinate point is marked with a pre-defined symbol. The color of the symbol is determined based on the fault type of the fault feature gene unit node. Different fault types correspond to different colors. The size of the symbol is determined based on the association strength of the fault feature gene unit node. The higher the association strength, the larger the symbol.
9. The cable fault location method based on big data according to claim 1, characterized in that, The step of determining the physical laying section where the cable fault is located based on the spatiotemporal coupling positioning network and generating a cable fault location instruction containing section identifiers and fault characteristic gene unit association information includes: Analyze the spatiotemporal trajectory lines in the spatiotemporal coupled positioning network, and identify clustering regions where spatiotemporal trajectory lines converge. The clustering region is the overlapping area of multiple spatiotemporal trajectory lines within the same horizontal axis coordinate range and similar vertical axis coordinate range. Extract all fault feature gene unit nodes corresponding to spatiotemporal trajectory lines within the aggregation area, and collect the fault type, correlation strength, and corresponding physical laying section identifier of the fault feature gene unit nodes; The frequency of occurrence of each fault type within the clustered area is statistically analyzed, and the fault type with the highest frequency of occurrence is identified as the main fault type. Calculate the ratio of the number of nodes in each physical laying section within the cluster area to the total number of nodes in the cluster area, and determine the physical laying section with the highest ratio as the candidate fault section. Extract the starting position coordinates, ending position coordinates, and length of the fault candidate segment; simultaneously extract the gene coding sequences of all fault feature gene unit nodes within the fault candidate segment. The main fault types, the location information of the fault candidate segments, the segment identifiers of the fault candidate segments, and the gene coding sequences of the corresponding fault feature gene unit nodes are integrated to form core fault location information. The core fault location information is processed into a structured form, and the structured fault location information is encapsulated into a cable fault location instruction.
10. A cable fault location system based on big data, characterized in that, The device includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the cable fault location method based on big data as described in any one of claims 1-9.