Cable fault locating method and system based on big data

CN122525288APending Publication Date: 2026-08-07SHANDONG QUANXING YINQIAO OPTICAL & ELECTRIC CABLE SCI & TECH DEV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG QUANXING YINQIAO OPTICAL & ELECTRIC CABLE SCI & TECH DEV
Filing Date
2026-05-08
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

一方面,这些方法往往仅依赖于单一的故障检测手段,例如仅通过测量电缆的电阻变化或检测局部放电信号来定位故障

Benefits of technology

[0006] Based on the above, this embodiment of the invention acquires a multi-source dynamic cable data set, covering real-time cable operation status data, historical fault evolution data, and dynamic data of the cable's surrounding environment. It utilizes big data streaming technology to construct a cable fault association knowledge graph, enabling real-time establishment of associations between multi-source data and cable fault types. Node weights are dynamically updated as real-time data flows in. The multi-source data is classified to obtain cable data classification results with spatiotemporal labels. Combined with the fault feature evolution trajectory, a cable fault location reasoning basis containing spatiotemporal association logic of multi-source data and a fault location derivation chain is generated. Accurate fault location reasoning is performed from both spatiotemporal dimensions. Finally, based on the reasoning basis, dynamic fault location deduction is executed to determine the real-time location information of the cable fault and generate a cable fault location command containing dynamic coordinates of the fault location, which is sent to the operation and maintenance terminal. This achieves real-time, accurate, and comprehensive fault location, greatly improving the efficiency and reliability of cable fault location and facilitating rapid power supply restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

The application provides a cable fault positioning method and system based on big data, and relates to the technical field of cable operation and maintenance. First, a multi-source cable dynamic data set containing real-time cable operation state data, historical fault evolution data and cable surrounding environment dynamic data is obtained. Then, a cable fault correlation knowledge graph is constructed based on big data stream processing technology, realizing real-time correlation of multi-source data and fault types and dynamic updating of node weights. Next, the multi-source data is classified to obtain classification results with space-time labels. Combined with the classification results and fault feature evolution trajectories, positioning reasoning basis is generated. Finally, according to the reasoning basis, fault position dynamic reasoning is performed to determine real-time position information and generate positioning instructions containing fault position dynamic coordinates and send them to the operation and maintenance terminal, improving the efficiency and reliability of cable fault positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cable maintenance technology, and more specifically, to a cable fault location method and system based on big data. Background Technology

[0002] In the field of cable system operation and maintenance, accurate and timely location of cable faults is crucial for ensuring stable power supply and reducing power outage losses. Traditional cable fault location methods have many limitations. On the one hand, these methods often rely on only a single fault detection method, such as locating the fault solely by measuring changes in cable resistance or detecting partial discharge signals. However, the occurrence of cable faults is a complex process, influenced by a combination of factors, and a single data source cannot comprehensively and accurately reflect the characteristics and location of the fault.

[0003] On the other hand, existing methods lack effective utilization of the dynamism and correlation of data when processing it. There are close connections between cable operating status data, historical fault evolution data, and dynamic data of the surrounding environment, but traditional methods fail to fully explore the potential correlations between these data, making it impossible to construct a comprehensive and dynamic fault knowledge system. Furthermore, in the fault location process, traditional methods typically employ static location models, which are difficult to adapt to real-time changes in cable operating status and the dynamic evolution of fault characteristics, resulting in inaccurate and untimely location results that cannot meet the demands of modern power systems for efficient and accurate fault location. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a cable fault location method based on big data, the method comprising: Acquire a multi-source cable dynamic data set, which includes real-time cable operating status data, historical fault evolution data, and dynamic data of the cable's surrounding environment; A cable fault association knowledge graph is constructed based on big data streaming technology. The cable fault association knowledge graph is used to establish a real-time association between a multi-source dynamic cable data set and cable fault types. The node weights of the cable fault association knowledge graph are dynamically updated as real-time data flows in. The multi-source cable dynamic data set is classified according to the spatiotemporal coupling rules of the cable fault association knowledge graph to obtain cable data classification results with spatiotemporal labels; The cable fault location reasoning basis is generated by combining the cable data classification results with the fault feature evolution trajectory. The cable fault location reasoning basis includes the spatiotemporal correlation logic of multi-source data and the fault location derivation chain. Based on the cable fault location reasoning, the fault location is dynamically deduced to determine the real-time location information of the cable fault, a cable fault location instruction containing the dynamic coordinates of the fault location is generated, and the cable fault location instruction containing the dynamic coordinates of the fault location is sent to the cable maintenance terminal.

[0005] In another aspect, embodiments of the present invention also provide a cable fault location system based on big data, including a processor and a machine-readable storage medium connected to the processor. The machine-readable storage medium is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the above-described method.

[0006] Based on the above, this embodiment of the invention acquires a multi-source dynamic cable data set, covering real-time cable operation status data, historical fault evolution data, and dynamic data of the cable's surrounding environment. It utilizes big data streaming technology to construct a cable fault association knowledge graph, enabling real-time establishment of associations between multi-source data and cable fault types. Node weights are dynamically updated as real-time data flows in. The multi-source data is classified to obtain cable data classification results with spatiotemporal labels. Combined with the fault feature evolution trajectory, a cable fault location reasoning basis containing spatiotemporal association logic of multi-source data and a fault location derivation chain is generated. Accurate fault location reasoning is performed from both spatiotemporal dimensions. Finally, based on the reasoning basis, dynamic fault location deduction is executed to determine the real-time location information of the cable fault and generate a cable fault location command containing dynamic coordinates of the fault location, which is sent to the operation and maintenance terminal. This achieves real-time, accurate, and comprehensive fault location, greatly improving the efficiency and reliability of cable fault location and facilitating rapid power supply restoration. Attached Figure Description

[0007] Figure 1 This is a schematic diagram of the execution flow of the cable fault location method based on big data provided in the embodiments of the present invention.

[0008] Figure 2 This is a schematic diagram of exemplary hardware and software components of the cable fault location system based on big data provided in an embodiment of the present invention. Detailed Implementation

[0009] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a cable fault location method based on big data according to an embodiment of the present invention. The following is a detailed description of the cable fault location method based on big data.

[0010] Step S110: Obtain a multi-source cable dynamic data set, which includes real-time cable operating status data, historical fault evolution data, and dynamic data of the cable's surrounding environment.

[0011] In this embodiment, an urban underground cable network is used as the application scenario. Real-time cable operation status data is collected through various sensors deployed on the cable lines. These sensors are distributed in different sections of the cable and can monitor various operating parameters of the cable in real time. Historical fault evolution data comes from the historical database of the cable operation and maintenance management system, containing data records of the entire process from the occurrence to the completion of various past faults. Dynamic data of the cable's surrounding environment is collected through environmental monitoring equipment installed in cable tunnels, cable wells, and surrounding areas, covering data on various environmental factors affecting cable operation.

[0012] Specifically, real-time cable operating status data includes current, voltage, and temperature data during cable operation. This data is continuously collected by sensors at set time intervals and transmitted to the data processing center. For example, current data reflects the charge flow within the cable, voltage data reflects the potential difference between the cable's ends, and temperature data shows the cable's heat generation during operation. Historical fault evolution data includes information such as the time of each fault occurrence, fault type, cable operating parameters at the time of the fault, parameter changes during the fault development process, fault handling measures, and results. Dynamic data of the cable's surrounding environment includes temperature, humidity, vibration, and pollutant concentration data. Changes in ambient temperature may affect the cable's heat dissipation efficiency, humidity levels may cause changes in cable insulation performance, vibration may originate from nearby construction or traffic factors, and pollutant concentration relates to cable corrosion.

[0013] Step S120: Construct a cable fault association knowledge graph based on big data streaming technology. The cable fault association knowledge graph is used to establish a real-time association between multi-source dynamic cable data sets and cable fault types. The node weights of the cable fault association knowledge graph are dynamically updated as real-time data flows in.

[0014] Step S121: Analyze the data stream characteristics of the multi-source cable dynamic data set, and distinguish the streaming transmission frequency of real-time cable operation status data, the time span attribute of historical fault evolution data, and the collection cycle attribute of dynamic data of the cable's surrounding environment.

[0015] Analyzing the data flow characteristics of multi-source cable dynamic data sets is a fundamental step in constructing a cable fault association knowledge graph. For real-time cable operating status data, the streaming transmission frequency varies depending on the parameter type and monitoring requirements. For example, current and voltage data typically require higher transmission frequencies to capture their instantaneous changes in real time, while temperature data changes relatively slowly, allowing for a lower transmission frequency. By analyzing the transmission frequency of real-time data, we can understand the data update speed and data volume.

[0016] The time span attribute of historical fault evolution data refers to the time range covered by each fault data segment. Different types of faults have different time spans in their evolution. Some faults may develop rapidly in a short period of time, while others may undergo a longer development process, with a larger time span. Analyzing the time span attribute of historical fault evolution data can help us understand the development cycle characteristics of different fault types, which is helpful in establishing time-related relationships in knowledge graphs.

[0017] The acquisition cycle of dynamic environmental data surrounding cables is determined by the settings of the environmental monitoring equipment; different environmental parameters may require different acquisition cycles. For example, a shorter acquisition cycle is used for rapidly changing vibration data, while a longer acquisition cycle can be used for slowly changing pollutant concentration data. Clearly defining the acquisition cycle of dynamic environmental data allows for matching environmental data with real-time cable operating status data over time, enabling the analysis of the timeliness of environmental factors' impact on cable faults.

[0018] When analyzing data stream characteristics, it is necessary to analyze the transmission logs and data tagging information of various types of data. For real-time cable operation status data, the streaming frequency is determined by examining the timestamp intervals of data transmission. For historical fault evolution data, the start and end times of each fault record are extracted, the time span is calculated, and statistical analysis is performed to determine the time span distribution characteristics of different fault types. For dynamic data of the cable's surrounding environment, the configuration parameters of the environmental monitoring equipment are examined to obtain the collection cycle settings for various environmental parameters, and the accuracy of the collection cycle is verified by combining the timestamps of the actual collected data.

[0019] Step S122: Extract core feature factors related to cable fault evolution from each data stream. The core feature factors include real-time fluctuation characteristics of cable operating parameters, environmental disturbance characteristics when historical faults occur, and operating data change characteristics during the fault propagation process.

[0020] Step S1221: Analyze the data transmission protocol and data structure of each data stream, and identify the fields directly related to the cable fault evolution. The fields directly related to the cable fault evolution include the cable real-time current field, the cable real-time voltage field, the cable real-time temperature field, the temperature and humidity change rate field of the cable's surrounding environment, and the fault propagation speed field.

[0021] Different types of data streams employ different transmission protocols and data structures, requiring separate parsing. For real-time cable operation status data, the transmission protocol may use industrial Ethernet or a dedicated sensor transmission protocol, and the data structure includes sensor identifiers, acquisition timestamps, and various operating parameter fields. By parsing these protocols and data structures, the meaning and location of each field in the data can be clarified.

[0022] Identifying fields directly related to cable fault evolution requires knowledge of cable faults and historical experience. Real-time cable current, voltage, and temperature fields directly reflect the cable's operating status; abnormal changes in these fields are often related to faults. The ambient temperature and humidity change rate field reflects the rate of environmental change; rapid temperature and humidity changes can adversely affect the cable. The fault propagation rate field reflects the spread of the fault along the cable line.

[0023] Step S1222: Extract the parameter change curve within the continuous time window from the real-time current field and the real-time voltage field of the cable, calculate the slope change rate and peak occurrence frequency of the parameter change curve, and use the slope change rate and peak occurrence frequency of the parameter change curve as the core components of the real-time fluctuation characteristics of the cable operating parameters.

[0024] Extracting parameter variation curves and calculating relevant features from the real-time current and voltage fields of the cable can effectively capture the real-time fluctuations of cable operating parameters. The selection of the continuous time window needs to be determined based on the variation characteristics of the cable parameters. Choosing an appropriate time window size ensures that it includes sufficient parameter variation information without causing information redundancy due to an excessively large window.

[0025] The slope of the parameter change curve reflects the rate of parameter change. A larger slope indicates more drastic parameter changes per unit time, potentially indicating a risk of failure. The frequency of peak occurrences indicates the number of times the parameter exceeds a set threshold within a time window. A higher frequency of peak occurrences indicates more frequent parameter fluctuations and poorer cable stability.

[0026] In practice, the real-time current and voltage data of the cable are first divided into time windows, forming a continuous sequence within each window. Then, parameter variation curves are plotted based on these data sequences, with time on the horizontal axis and current or voltage value on the vertical axis. Next, the slope of the curve at different time points is calculated, and the rate of change of the slope is determined. The magnitude of the rate of change of the slope measures the severity of parameter changes. Simultaneously, the number of peak values ​​appearing in the parameter variation curve within each time window is counted, i.e., the peak frequency. The rate of change of the slope and the peak frequency within each time window are considered the core components of the real-time fluctuation characteristics of the cable operating parameters within that window.

[0027] Step S1223: Filter the abrupt change records of the dynamic data of the cable surrounding environment before the fault occurred in the historical fault evolution data, and extract the abnormal increments of the cable surrounding environment parameters in the abrupt change records. The abnormal increments of the cable surrounding environment parameters include the short-term sudden rise and fall of temperature and humidity in the cable surrounding environment, the instantaneous increase of external mechanical force around the cable, and the sudden increase of pollutant concentration in the cable surrounding environment. The abnormal increments of the cable surrounding environment parameters are used as the environmental disturbance characteristics at the time of the historical fault.

[0028] By filtering historical fault evolution data, records of abrupt changes in environmental dynamics prior to the fault occurrence can be used to identify environmental disturbances associated with the fault. Before a fault occurs, certain parameters of the surrounding environment may exhibit abnormal abrupt changes, which could be triggers for the fault. Therefore, extracting abrupt increases in environmental parameters from these abrupt change records is of great significance.

[0029] Sudden increases or decreases in temperature and humidity around cables within a short period can affect the cable's insulation performance and heat dissipation. Sudden increases in external mechanical forces can cause physical damage to the cable, while a sudden increase in the concentration of pollutants in the surrounding environment can accelerate the cable's corrosion process. These abnormal increases in environmental parameters collectively constitute the environmental disturbance characteristics of historical fault occurrences.

[0030] During the operation, the occurrence time of each fault is first located from historical fault evolution data. Then, dynamic environmental data of the cable's surrounding environment within a set time period prior to that time is extracted. This environmental data is analyzed to identify abrupt changes, which are data records showing significant changes in environmental parameters within a short period. For temperature and humidity data, the magnitude of sudden increases and decreases is calculated, i.e., the temperature and humidity differences before and after the abrupt change. For external mechanical force data, the instantaneous increase is calculated, i.e., the difference between the force value after and before the abrupt change. For pollutant concentration data, the sudden increase is calculated, i.e., the difference between the concentration value after and before the abrupt change. These calculated abnormal increments are then compiled and used as environmental disturbance characteristics at the time of historical fault occurrences.

[0031] Step S1224: Collect cable operating status data during the fault propagation stage from the historical fault evolution data, extract the parameter attenuation law of the cable operating status data as the fault propagation distance changes, and use the parameter attenuation law as the operating data change feature during the fault propagation process. The parameter attenuation law as the fault propagation distance changes includes the decreasing gradient of cable current with fault propagation distance, the loss rate of cable voltage with fault propagation distance, and the changing trend of cable temperature with fault propagation distance.

[0032] By collecting cable operating status data during the fault propagation stage from historical fault evolution data, it is possible to analyze the changing patterns of cable operating parameters during fault propagation. When a fault propagates along a cable line, cable operating parameters at different distances will exhibit specific attenuation or change characteristics, which can reflect the propagation path and scope of the fault.

[0033] The decreasing gradient of cable current with fault propagation distance reflects the current attenuation under the influence of the fault; a larger gradient indicates faster current attenuation. The loss rate of cable voltage with fault propagation distance reflects the degree of voltage loss during propagation; a higher loss rate indicates a more significant voltage drop. The trend of cable temperature with fault propagation distance shows the temperature distribution during fault propagation, which may show an increasing or decreasing trend, depending on the type of fault and its propagation mode.

[0034] Specifically, cable operating status data during the fault propagation stage is selected from historical fault evolution data, including operating parameter records at different time points and cable locations. Based on the fault's starting location and propagation direction, the corresponding fault propagation distance for each operating parameter record is determined. Then, current data at different propagation distances at the same time point are analyzed, calculating the decrease in current value with increasing distance to obtain the current's descent gradient with fault propagation distance. For voltage data, voltage values ​​at different propagation distances are similarly analyzed, calculating the ratio of voltage loss with increasing distance to distance to obtain the voltage loss rate with fault propagation distance. For temperature data, temperature changes at different propagation distances are observed, summarizing the temperature change trend with fault propagation distance, such as an initial increase followed by a decrease, or a continuous increase.

[0035] Step S1225: Perform spatiotemporal correlation verification on the extracted cable operating parameters, environmental disturbance characteristics at the time of historical fault occurrence, and operating data change characteristics during the fault propagation process, and retain the feature items directly related to the cable fault evolution timeline and the cable fault spatial propagation path.

[0036] The spatiotemporal correlation verification of the extracted features is to ensure that these features are directly related to the evolution of cable faults. The spatiotemporal correlation verification is performed from two dimensions: time and space. In the time dimension, it checks whether the features match the timeline of fault evolution, and in the spatial dimension, it verifies whether the features are related to the spatial propagation path of the fault.

[0037] In terms of time, the occurrence time of various features is compared with key time nodes in the cable fault evolution, such as the warning period before the fault occurs, the moment the fault occurs, and different stages of fault propagation. Features that closely match these key time nodes are retained, while features that are not clearly related to the fault evolution timeline are removed.

[0038] In the spatial dimension, the relationship between the cable location corresponding to various features and the fault spatial propagation path is analyzed. It checks whether the location of the feature is on the fault propagation path, or whether there is a certain spatial correlation with the fault propagation path. Features directly related to the fault spatial propagation path are retained, while features without spatial correlation are excluded.

[0039] For example, if a cable's operating parameter exhibits a real-time fluctuation characteristic shortly before a fault occurs, and the corresponding cable location is within the initial fault location area, then this characteristic will be retained in the spatiotemporal correlation verification. However, if an environmental disturbance characteristic occurs at a time significantly different from the fault occurrence time, and its corresponding location is not on the fault propagation path, then this characteristic may be discarded. Spatiotemporal correlation verification improves the specificity and effectiveness of the features.

[0040] Step S1226: Dynamically assign weights to the feature items after the spatiotemporal correlation verification. The weights are determined based on the frequency of the feature items in the real-time data stream and the advance of cable fault warning, forming an initial set of core feature factors. The higher the frequency and the greater the advance of cable fault warning, the higher the weight.

[0041] Dynamically assigning weights to feature items after spatiotemporal correlation verification can highlight the role of important features in fault correlation analysis. The weights are determined based on the frequency of the feature items in the real-time data stream and the lead time for cable fault warnings; these two factors reflect the prevalence and warning value of the feature items.

[0042] The higher the frequency of a feature in the real-time data stream, the more common that feature is during cable operation and the stronger its correlation with cable faults; therefore, it is given a higher weight. The greater the lead time for cable fault warnings, that is, the earlier the feature appears before the fault occurs, the more obvious the warning effect of that feature is on the fault, and the more time can be gained for fault handling; therefore, it is also given a higher weight.

[0043] In practice, the process begins by counting the number of times each feature item appears per unit time in the real-time data stream to obtain its frequency. Then, the occurrence time of each feature item relative to the fault occurrence time in historical fault cases is analyzed to calculate the early warning lead time. Based on preset weighting rules, the frequency and early warning lead time are converted into corresponding weight values. These two weight values ​​are then combined to obtain the dynamic weight of each feature item. All feature items are sorted from highest to lowest weight value, and a subset of feature items with higher weight values ​​are selected to form the initial set of core feature factors.

[0044] Step S1227: Perform reverse matching between the core feature factors in the initial set of core feature factors and historical fault cases to verify the contribution of each core feature factor to the determination of cable fault type. Retain the core feature factors with the highest contribution to the determination of cable fault type by a predetermined proportion, and finally determine the core feature factors related to cable fault evolution.

[0045] Step S12271: Extract complete case data of multiple types of cable faults from the historical fault case database. Each complete case data includes the record of changes in the core characteristic factors of the entire cable fault occurrence process and the final cable fault type label.

[0046] When extracting complete case data from the historical fault case database, it is necessary to cover various types of cable faults, such as insulation aging faults, mechanical damage faults, and overheating faults, to ensure comprehensive verification. The extraction of each complete case data point must be performed according to a unified standard, ensuring that it includes records of changes in core characteristic factors before, during, and after the fault occurs. For example, for insulation aging fault cases, it is necessary to extract the changes in characteristic factors from when the cable insulation performance begins to decline, to the state of characteristic factors at the time of the fault, and then to the evolution of characteristic factors as the fault further spreads. Simultaneously, each case data point should be clearly labeled with the final fault type tag, such as "insulation aging fault" or "mechanical damage fault."

[0047] During the extraction process, the data in the historical failure case database needs to be filtered and organized to remove cases with incomplete data or missing key information. For cases that meet the requirements, the changes in core feature factors are sorted out in chronological order to form a complete feature factor time series. Simultaneously, failure type labels are standardized to ensure that failures of the same type use consistent label names, avoiding validation errors caused by inconsistent labels. After extraction, these complete case data are categorized and stored according to failure type for subsequent training and validation set partitioning.

[0048] Step S12272: Construct an independent validation model for each core feature factor in the initial set of core feature factors. The input of the independent validation model is the change record of the core feature factor in the complete case data, and the output of the independent validation model is the cable fault type prediction result.

[0049] An independent validation model is built for each core feature factor, enabling individual evaluation of each feature factor's ability to determine the fault type. These independent validation models can employ model structures suitable for processing sequence data, such as recurrent neural network models. Each model focuses on analyzing the relationship between the corresponding core feature factor's change records and the fault type.

[0050] When constructing the model, the input layer structure of each model is first determined. The dimension of the input layer is determined based on the time series length of the core feature factor change records and the feature dimension. For example, if the change records of a certain core feature factor contain sequence data with multiple time points, the input layer needs to be able to accept a sequence input of that length. The hidden layer of the model has multiple neurons to extract key patterns and regularities in the feature factor changes. The number of neurons in the output layer corresponds to the number of cable fault types, and the output of each neuron represents the predicted probability of the corresponding fault type.

[0051] For each core feature factor, its variation record in the complete case data is used as the input sample for the model, and the corresponding fault type label is used as the expected output. In this way, each independent validation model can learn the mapping relationship between a single core feature factor and the fault type. For example, a validation model built for the core feature factor of cable temperature variation takes the temperature variation sequence over time in different fault cases as input and outputs the fault type prediction result corresponding to that sequence.

[0052] Step S12273: Divide the complete case data in the historical fault case database into a training set and a validation set according to a preset ratio. Use the training set to train each independent validation model and adjust the parameters of the independent validation model until the cable fault type prediction accuracy of the training set reaches a preset threshold.

[0053] Dividing the dataset into training and validation sets ensures that the model can be effectively evaluated on unseen data. The preset ratio can be determined based on the total amount of case data; for example, 70% of the complete case data could be allocated to the training set and 30% to the validation set. The partitioning process uses random sampling, while ensuring that the proportions of different fault types in the training and validation sets are consistent with those in the original dataset, avoiding bias in model training towards a particular fault type due to uneven data distribution.

[0054] When training each independent validation model using the training set, an iterative training approach is employed. In each training round, changes in core feature factors from the training set are recorded and input into the model. The model then makes predictions based on the current parameters, yielding a fault type prediction. The prediction result is compared with the actual fault type label to calculate the prediction error. Based on the prediction error, the model parameters, such as weights and biases, are adjusted using a backpropagation algorithm to reduce the prediction error.

[0055] During training, the prediction accuracy of the training set is calculated periodically, and training stops when the accuracy reaches a preset threshold. The preset threshold needs to be set by comprehensively considering the model's performance requirements and the risk of overfitting. If the threshold is set too high, the model may overfit the training data and perform poorly on the validation set; if the threshold is set too low, the model may not have fully learned the relationship between features and fault types. The model's loss function also needs to be monitored during training. When the loss function value stabilizes at a low level and no longer decreases significantly, this can also be used as one of the criteria for stopping training.

[0056] Step S12274: Test the trained independent validation models using the validation set, and record the model performance index of each independent validation model for different cable fault types.

[0057] Testing the trained model using a validation set allows us to evaluate its generalization ability on data not used in the training process. During testing, changes in core feature factors from the validation set are recorded and input into the corresponding trained independent validation model, which then outputs a fault type prediction. The prediction results are compared with the actual fault type labels on the validation set, and performance metrics for each model on the validation set are calculated.

[0058] Model performance metrics include accuracy, precision, recall, and F1 score. Accuracy reflects the proportion of correct predictions made by the model overall; precision, for a specific fault type, reflects the proportion of cases predicted as that type that actually occurred; recall, for a specific fault type, reflects the proportion of cases that actually occurred as that type that were correctly predicted by the model; and the F1 score is a combined metric of precision and recall, used to balance the performance of both.

[0059] For each independent validation model, its performance metrics for each type of fault are calculated and recorded in detail. For example, a validation model may have high precision for insulation aging faults but low recall for mechanical damage faults. Simultaneously, the overall performance metrics of the model on the validation set are recorded as a reference for evaluating the model's overall performance.

[0060] Step S12275: Construct a comprehensive scoring model for the contribution of core feature factors to cable fault type determination. After standardizing the performance indicators of each model of the independent verification model, perform weighted calculation according to preset weights to obtain the comprehensive contribution score of each core feature factor.

[0061] The purpose of constructing a comprehensive contribution scoring model is to integrate multiple performance indicators into a unified score, thereby quantifying the contribution of core feature factors. First, the performance indicators of each model need to be standardized to eliminate dimensional differences between them, enabling effective comparison and combination.

[0062] Standardization can be achieved by converting metric values ​​to a range of 0 to 1. For example, for accuracy, the minimum accuracy of all models is subtracted from the accuracy of each model, and then divided by the difference between the maximum and minimum accuracy values ​​to obtain the standardized accuracy value. The same standardization method is used for metrics such as precision, recall, and F1 score to ensure that all standardized metric values ​​fall within the same numerical range.

[0063] The preset weights are determined based on the importance of each performance indicator in fault type determination. For example, accuracy reflects the overall predictive ability of the model and can be given a higher weight; the F1 score combines precision and recall and is also given a higher weight; precision and recall are assigned corresponding weights according to actual needs. In this embodiment, the weights for accuracy, F1 score, precision, and recall can be set to 0.3, 0.3, 0.2, and 0.2 respectively.

[0064] Each standardized performance indicator is multiplied by its corresponding preset weight, and the products are then summed to obtain the overall contribution score for each core feature factor. A higher overall contribution score indicates a greater contribution of that core feature factor to cable fault type determination. For example, if a core feature factor has a standardized accuracy of 0.8, a standardized F1 score of 0.7, a standardized precision of 0.6, and a standardized recall of 0.7, then its overall contribution score is 0.71.

[0065] Step S12276: Sort the comprehensive contribution scores of all core feature factors in descending order, and select the core feature factors with the highest comprehensive contribution scores by a predetermined proportion as candidate core feature factors.

[0066] Sort the overall contribution scores in descending order to clearly reflect the relative contributions of each core feature factor. The sorting process arranges the core feature factors in descending order of their overall contribution scores, forming an ordered list. The preset proportions can be determined based on the total number of core feature factors and actual needs; for example, selecting the top 60% of core feature factors as candidate core feature factors.

[0067] After sorting, the number of candidate core feature factors is determined according to a preset ratio. For example, if the initial set of core feature factors contains 20 feature factors and the preset ratio is 60%, then the top 12 feature factors are selected as candidate core feature factors. During the selection process, if feature factors with the same overall contribution score are encountered, they can be re-sorted based on their performance on key performance indicators (such as accuracy or F1 score) to determine the final selection result.

[0068] Selecting candidate core feature factors can reduce the number of feature factors, retain features with high contribution, and improve the efficiency and accuracy of subsequent fault location and analysis. Meanwhile, the preset ratio can be adjusted according to the needs of the actual application scenario. If more feature factors are needed to ensure the comprehensiveness of the analysis, the preset ratio can be appropriately increased; if it is necessary to simplify the model and improve processing speed, the preset ratio can be appropriately decreased.

[0069] Step S12277: Substitute the candidate core feature factors into new historical fault cases for cross-validation, and check the overall accuracy of the candidate core feature factor combination in determining the cable fault type. If the overall accuracy of the candidate core feature factor combination in determining the cable fault type meets the preset requirements, it is determined as the final core feature factor related to the cable fault evolution. If the overall accuracy of the candidate core feature factor combination in determining the cable fault type does not meet the preset requirements, the selection ratio of candidate core feature factors is increased, and the cross-validation process is repeated until the overall accuracy of the candidate core feature factor combination in determining the cable fault type meets the standard, and the core feature factor related to the cable fault evolution is finally determined.

[0070] Cross-validation is used to test the overall performance of candidate core feature factor combinations, ensuring that they maintain good accuracy across different failure cases. New historical failure cases refer to case data that were not involved in previous model training and validation; this case data can more objectively evaluate the generalization ability of candidate feature factor combinations.

[0071] Candidate core feature factors are combined and applied to new historical fault cases. Change records of these candidate core feature factors are extracted from the cases and input into the corresponding model or analysis system to obtain fault type determination results. The determination results are compared with the actual fault type labels to calculate the overall accuracy rate, which is the ratio of the number of correctly determined cases to the total number of cases.

[0072] The preset requirement for overall accuracy is determined based on the actual fault location accuracy requirements. For example, the preset requirement for overall accuracy is not less than 85%. If the overall accuracy of the candidate core feature factor combination reaches or exceeds the preset requirement, it indicates that the combination can better reflect the characteristics of cable fault evolution and can be determined as the final core feature factor.

[0073] If the overall accuracy does not meet the preset requirements, the selection ratio of candidate core feature factors needs to be increased, for example, from 60% to 70%, and candidate core feature factors need to be reselected. Then, the cross-validation process is repeated using the new combination of candidate core feature factors, and the overall accuracy is calculated again. This process is repeated until the overall accuracy of the candidate core feature factor combination reaches the preset requirements. During the process of increasing the selection ratio, the percentage increase each time can be determined according to the actual situation, such as increasing by 10% each time. Through cross-validation and ratio adjustment, the final determined core feature factor combination can be ensured to have high reliability and effectiveness.

[0074] Step S123: Construct the initial topology of the cable fault association knowledge graph, using core feature factors as nodes of the cable fault association knowledge graph, and using the association frequency between core feature factors in historical fault cases as the initial edge weights to form a node connection network.

[0075] Constructing the initial topology is the foundational framework for building the cable fault association knowledge graph. After identifying the core feature factors related to cable fault evolution, each core feature factor is treated as an independent node in the knowledge graph. These nodes cover multiple aspects, including real-time fluctuation characteristics of cable operating parameters, environmental disturbance characteristics during historical fault occurrences, and operational data changes during fault propagation.

[0076] The initial edge weights are determined based on the correlation frequency between core feature factors in historical fault cases. Correlation frequency refers to the number of times two core feature factors appear simultaneously or show a significant correlation in historical fault cases. For example, in multiple insulation aging fault cases, abnormal increases in cable temperature and abnormal changes in humidity around the cable often occur simultaneously, resulting in a high correlation frequency between these two core feature factor nodes.

[0077] When calculating association frequencies, all historical failure cases are traversed, and the co-occurrence of each pair of core feature factors is statistically analyzed. For each case, it is checked whether any two core feature factors appear simultaneously or have a significant correlation in that case. If so, the association frequency of this pair of feature factors is incremented by 1. After the statistics are completed, each association frequency is normalized by dividing by the total number of cases to obtain the initial edge weights. The initial edge weights range from 0 to 1, with larger values ​​indicating a stronger association between the two core feature factors.

[0078] Based on the core feature factor nodes and initial edge weights, a node connection network is constructed to form the initial topology of the cable fault association knowledge graph. The connection relationships and weights between nodes reflect the historical associations between the core feature factors.

[0079] Step S124: Deploy the big data streaming engine and connect the real-time cable operation status data to the data stream channel of the big data streaming engine. When new core feature factors appear in the real-time data or the values ​​of existing core feature factors exceed the historical fluctuation range, the node weight update mechanism of the node connection network is triggered.

[0080] Deploying a big data streaming engine enables efficient processing and real-time analysis of real-time cable operation status data. The big data streaming engine features low latency and high throughput, allowing for rapid processing of continuously flowing, massive amounts of real-time data.

[0081] When connecting real-time cable operation status data to the data stream channel, the data access interface needs to be configured to ensure that the data flows into the engine according to the set format and protocol. The data stream channel receives, parses, and preprocesses the data in real time, converting it into a format that the engine can process. Simultaneously, a data buffering mechanism is established to handle fluctuations in data traffic and ensure the stability of engine processing.

[0082] The triggering conditions for the node weight update mechanism are set as follows: When a new core feature factor appears in the real-time data, it indicates that a previously unidentified feature related to fault evolution has been discovered. A new node needs to be added to the node connection network, and connections between this new node and other relevant nodes need to be established. The initial edge weights can be set based on the correlation of similar features. When the value of an existing core feature factor exceeds its historical fluctuation range, it indicates an abnormal change in that feature factor, which may foreshadow the occurrence or development of a fault. In this case, the weight update mechanism needs to be triggered to adjust the edge weights associated with that feature factor.

[0083] The historical fluctuation range is determined by analyzing the distribution of core characteristic factors in historical failure cases and normal operation data. For example, the maximum, minimum, and standard deviation of the characteristic factor in historical data are calculated, and the range exceeding the average plus or minus three standard deviations is set as the historical fluctuation range. When the characteristic factor value in real-time data exceeds this range, the weight update mechanism is triggered.

[0084] Step S125: Based on the similarity matching results between real-time data and historical fault evolution data, adjust the edge weights of corresponding nodes in the node connection network. The higher the similarity, the greater the increase in edge weights; conversely, the edge weight decay is triggered.

[0085] Similarity matching between real-time data and historical fault evolution data is an important basis for adjusting edge weights. Similarity matching determines the degree of similarity between the current cable condition reflected in the data and the historical fault condition by comparing the similarity between the core feature factor combinations in the real-time data and the core feature factor combinations in historical fault cases.

[0086] Similarity matching can be calculated using methods such as cosine similarity or Euclidean distance. For the core feature factor combination vector in real-time data and the core feature factor combination vector in historical fault cases, the similarity value between them is calculated. A higher similarity value indicates a greater similarity between the features of the real-time data and the historical fault case, and a higher probability that a similar fault will occur in the current cable condition.

[0087] The edge weights are adjusted based on the similarity matching results. When the similarity is high, the edge weights between core feature factor nodes related to the historical failure case are increased. The increase is proportional to the similarity value; the higher the similarity, the greater the increase. For example, the increase is greater when the similarity value is 0.9 than when the similarity value is 0.6. Increasing the edge weights strengthens the associations between these feature factors, making the knowledge graph more accurately reflect the current failure risk.

[0088] When the similarity is low, it indicates a significant difference between the features in the real-time data and those in historical failure cases. This triggers the edge weight decay mechanism, reducing the edge weights between related nodes. The decay magnitude is determined by the similarity value; the lower the similarity, the greater the decay. Edge weight decay prevents the knowledge graph from retaining outdated or irrelevant relationships, ensuring its timeliness and accuracy. When adjusting edge weights, upper and lower limits need to be set to prevent distortion of relationships caused by excessively large or small weight values. Simultaneously, the time, reason, and magnitude of each weight adjustment should be recorded.

[0089] Step S126: Integrate the real-time updated node and edge weights to form a cable fault association knowledge graph containing spatiotemporal dimension association rules. The spatiotemporal dimension association rules include the association strength differences of the same core feature factor in different time windows and different cable sections.

[0090] Integrating the real-time updated node and edge weights involves standardizing and structuring the dynamically adjusted node connection network to form a complete cable fault association knowledge graph. During the integration process, newly added nodes, adjusted edge weights, and existing nodes and edge weights need to be stored and managed uniformly to ensure the consistency and integrity of the knowledge graph data.

[0091] The extraction of spatiotemporal correlation rules is completed during the integration process, determined by analyzing the changes in the correlation strength of core feature factors across different time windows and cable sections. The difference in correlation strength of the same core feature factor across different time windows refers to the different values ​​exhibited by the correlation weights of this feature factor with other feature factors at different time stages. For example, within the warning time window before a fault occurs, the correlation strength between cable temperature and current characteristics is high; while within the normal operation time window, the correlation strength is low.

[0092] The difference in correlation strength across different cable sections refers to the varying correlation weights of the same core feature factor with other feature factors in different cable sections. For example, in cable joint sections, the correlation strength between vibration characteristics and mechanical damage fault characteristics is high; while in straight cable sections, this correlation strength is low. Integrating these spatiotemporal correlation rules into the cable fault correlation knowledge graph allows the knowledge graph to not only reflect the correlation relationships between feature factors but also to demonstrate the temporal and spatial variations of these correlation relationships.

[0093] Step S130: Classify the multi-source cable dynamic data set according to the spatiotemporal coupling rules of the cable fault association knowledge graph to obtain cable data classification results with spatiotemporal labels.

[0094] Step S131: Extract the spatiotemporal coupling classification rules from the cable fault association knowledge graph, and analyze the core feature factors in the spatiotemporal coupling classification rules, such as the time correlation threshold, spatial segment matching accuracy, and fault type association confidence requirements.

[0095] Extracting spatiotemporally coupled classification rules from the cable fault association knowledge graph is the first step in the classification process. These rules are summarized based on historical data and fault association relationships, and serve as the basis for data classification. The time correlation threshold of core feature factors specifies the required degree of correlation between different core feature factors over time. Only when the degree of temporal correlation between two feature factors reaches or exceeds this threshold are they considered to be correlated in time.

[0096] Spatial segment matching accuracy determines the fineness of the cable segment division and the matching standard to which the data belongs. Higher matching accuracy means that the cable segment division is more detailed and the matching between data and segments is more accurate. Fault type association confidence requirement sets the minimum level of confidence for data to be associated with a certain fault type. Only when the association confidence between data and fault type reaches this requirement can the data be classified into the category corresponding to that fault type.

[0097] During the extraction and parsing process, it is necessary to obtain the specific content of the spatiotemporal coupling classification rules through the knowledge graph query interface. For the time correlation threshold of core feature factors, the time interval range or time synchronization rate requirement corresponding to different combinations of feature factors in the parsing rules needs to be specified. For spatial segment matching accuracy, the rule's division length, coordinate range, and allowable range of matching errors between data location and segment need to be clearly defined. For fault type association confidence requirements, the minimum confidence value or level standard set for different fault types in the extraction rules needs to be specified.

[0098] Step S132: Label each data point in the multi-source cable dynamic data set with the collection timestamp and cable section identifier to form a data unit with basic spatiotemporal information.

[0099] Labeling each data point in a multi-source cable dynamic dataset with a collection timestamp and cable segment identifier is the process of adding basic spatiotemporal information to the data. The collection timestamp records the exact time the data was collected, accurate to the second or millisecond, which is the basis for time window segmentation and time correlation analysis. The cable segment identifier indicates the cable segment corresponding to the data collection, usually represented by a segment number, with each segment number corresponding to a specific cable line location and range.

[0100] During the annotation process, for real-time cable operation status data and dynamic data of the cable's surrounding environment, the collection timestamp is directly extracted from the data's metadata. If the data does not contain timestamp information, it is calculated and supplemented based on the data reception time and transmission delay. For cable segment identification, the cable segment to which the data acquisition sensor belongs is determined based on the installation location of the sensor. Each sensor is pre-associated with a corresponding segment number, and the corresponding cable segment identification can be found and annotated on the data through the sensor identification. After annotation, each data point becomes a data unit containing basic spatiotemporal information, preparing for subsequent matching and classification.

[0101] Step S133: According to the spatiotemporal coupling classification rule, the core feature factors of each data unit with basic spatiotemporal information are matched with the nodes in the same time window and the same cable segment in the cable fault association knowledge graph. The spatiotemporal correlation degree between the data unit with basic spatiotemporal information and each fault type is calculated. The calculation of the spatiotemporal correlation degree between the data unit with basic spatiotemporal information and each fault type is based on the time synchronization of the core feature factors, the spatial location overlap, and the edge weights of the cable fault association knowledge graph.

[0102] Step S1331: Extract the time window division standard and spatial segment division precision from the spatiotemporal coupling classification rules. Divide the time dimension of the multi-source cable dynamic data set into continuous time windows according to the time window division standard, and divide the spatial dimension of the multi-source cable dynamic data set into cable segments of equal precision according to the spatial segment division precision.

[0103] Extracting the time window division criteria and spatial segmentation precision from the spatiotemporal coupling classification rules serves as the basis for dividing time and space dimensions. The time window division criteria specify the length and division method of the time window. For example, the time dimension can be divided into continuous and non-overlapping time windows according to fixed time intervals (such as every minute or every hour).

[0104] The spatial segmentation accuracy determines the segmentation length and principle of the cable segment. According to this accuracy, the spatial dimension of the cable line is divided into multiple cable segments of equal length or equal accuracy, and each segment has a clear start and end position and range.

[0105] During the segmentation process, based on the time window segmentation criteria and the data collection timestamps, all data in the multi-source cable dynamic data set are allocated to corresponding time windows. For the spatial dimension, based on the actual route and length of the cable lines, combined with the spatial segmentation precision, the cable lines are divided into several continuous cable segments, and each segment is assigned a unique segment number. This temporal and spatial segmentation allows the data to be matched and analyzed within a unified spatiotemporal framework.

[0106] Step S1332: Assign a corresponding time window identifier and spatial segment identifier to each data unit with basic spatiotemporal information.

[0107] After dividing the time windows and spatial segments, it is necessary to assign a corresponding time window identifier and spatial segment identifier to each data unit with basic spatiotemporal information. The time window identifier is a unique number for each time window. Based on the time range of the data unit's acquisition timestamp, the time window to which it belongs is determined, and the identifier of that time window is assigned to the data unit.

[0108] The spatial segment identifier is based on the original cable segment identifier of the data unit, corresponding to the newly divided spatial segment number, ensuring that the spatial location information of the data unit is consistent with the divided spatial segments. By assigning identifiers, each data unit clearly belongs to its specific time window and spatial segment.

[0109] Step S1333: Extract the core feature factor set of each data unit with basic spatiotemporal information, match each core feature factor in the core feature factor set with the corresponding node in the same time window and the same cable segment in the cable fault association knowledge graph, and record the number of successfully matched core feature factors and the current edge weight of each matched node.

[0110] After extracting the core feature factor set of the data unit, each core feature factor is matched with the corresponding node in the same time window and spatial segment of the cable fault association knowledge graph. During the matching process, by comparing information such as the type, attribute, and numerical range of the feature factors, it is determined whether the feature factors of the data unit match the nodes in the knowledge graph.

[0111] For each successfully matched core feature factor, its number and the current edge weight of each matched node in the knowledge graph are recorded. Edge weights reflect the strength of the relationship between that node and other nodes, and are an important parameter for calculating spatiotemporal correlation. Through matching and recording, the association information between data units and knowledge graph nodes can be obtained.

[0112] Step S1334: Calculate the time synchronization of core feature factors. Compare the difference between the acquisition time of the core feature factors of the data unit with basic spatiotemporal information and the latest update time of the knowledge graph node of cable fault association. The smaller the difference, the higher the time synchronization of core feature factors. The value of the time synchronization of core feature factors is calculated by inverting the ratio of the difference to the length of the time window.

[0113] When calculating the time synchronization of core feature factors, we first obtain the collection time of the core feature factors of the data unit and the latest update time of the corresponding node in the knowledge graph. Then, we calculate the difference between these two times. The smaller the difference, the more synchronized the feature factors of the data unit are with the feature factors of the knowledge graph node in time, and the higher the time synchronization.

[0114] The specific numerical value of the core characteristic factor, time synchronization, is calculated by inverting the ratio of the time difference to the time window length. For example, if the time difference is a certain value and the time window length is a fixed value, the smaller the ratio, the larger the inverted time synchronization value, indicating higher time synchronization. This calculation method quantifies time synchronization into a specific numerical value, facilitating subsequent comprehensive calculations.

[0115] Step S1335: Calculate the spatial location overlap. Based on the spatial segment identifier of the data unit with basic spatiotemporal information and the cable segment range associated with the cable fault association knowledge graph node, determine the overlap ratio between the location of the data unit with basic spatiotemporal information and the associated segment of the cable fault association knowledge graph node. The overlap ratio is the spatial location overlap.

[0116] When calculating spatial overlap, the cable segment range where the data unit is located is determined based on the spatial segment identifier, and the cable segment range associated with the knowledge graph node is also obtained. By comparing these two segment ranges, the ratio of their overlapping length to the length of the segment where the data unit is located is calculated; this ratio is the spatial overlap.

[0117] A higher overlap ratio indicates a greater degree of spatial overlap between the location of the data unit and the associated segment of the knowledge graph node, and a higher degree of spatial overlap. The numerical range of spatial overlap is between 0 and 1.

[0118] Step S1336: Construct a spatiotemporal correlation calculation model between data units with basic spatiotemporal information and each fault type. After normalizing the sum of edge weights of successfully matched core feature factors, the time synchronization value of core feature factors, and the spatial location overlap, perform weighted fusion according to preset weight coefficients. The weighted fusion result is the spatiotemporal correlation between data units with basic spatiotemporal information and corresponding fault types.

[0119] When constructing a spatiotemporal correlation calculation model, it is first necessary to normalize three parameters: the sum of edge weights of successfully matched core feature factors, the temporal synchronization value of core feature factors, and the spatial overlap degree. The purpose of normalization is to transform parameters with different dimensions and numerical ranges into a unified numerical interval (usually 0 to 1) for comprehensive calculation.

[0120] After normalization, the three parameters are weighted and fused according to preset weight coefficients. The preset weight coefficients are determined based on the degree of influence of each parameter on the spatiotemporal correlation. For example, the sum of edge weights may have a greater impact on the correlation, so a higher weight coefficient is assigned. The result of the weighted fusion is the spatiotemporal correlation between the data unit and the corresponding fault type, which comprehensively reflects the overall correlation between the data unit and the fault type in terms of time, space, and feature association.

[0121] Step S1337: Repeat the above matching and calculation process for each data unit with basic spatiotemporal information to obtain the spatiotemporal correlation degree value between each data unit with basic spatiotemporal information and each fault type, and establish a correspondence table between the data unit and each fault type.

[0122] For each data unit with basic spatiotemporal information in the multi-source cable dynamic data set, the above-mentioned matching, parameter calculation, and spatiotemporal correlation calculation processes are repeatedly performed. Through processing one by one, the spatiotemporal correlation values ​​between each data unit and various fault types are obtained.

[0123] Then, these values ​​are organized into a spatiotemporal correlation table between data units and various fault types. The table clearly shows the degree of correlation between each data unit and different fault types. This correlation table is an important basis for subsequent data filtering and classification. By querying the correlation values ​​in the table, the fault type category to which the data unit belongs can be quickly determined.

[0124] Step S134: Select data units with basic spatiotemporal information whose spatiotemporal correlation with each fault type meets the fault type correlation confidence requirement, group them according to fault type to form an initial classification dataset, and each initial classification dataset carries a corresponding time window label and cable section label.

[0125] Based on the confidence level requirements for fault type association, data units that meet the association requirements are selected from the spatiotemporal correlation table between data units and each fault type. For each fault type, all data units whose association with that fault type reaches or exceeds the confidence level requirements are selected and grouped according to fault type to form an initial classification dataset.

[0126] Each initial classification dataset carries a corresponding time window label and cable section label. These labels indicate the temporal and spatial distribution characteristics of the data units within the dataset. For example, for an initial classification dataset targeting overheating faults, the time window labels might be concentrated during periods of high-temperature weather, and the cable section labels might be concentrated in sections with high cable loads. Through grouping and labeling, the data is initially organized according to fault type and spatiotemporal attributes.

[0127] Step S135: Call the historical core feature factor evolution sequence of the same fault type from the cable fault association knowledge graph, align the data units with basic spatiotemporal information in the initial classification dataset with the historical core feature factor evolution sequence in terms of time dimension, and retain the data units with basic spatiotemporal information that conform to the historical core feature factor evolution trend.

[0128] The evolution sequence of historical core feature factors, which belongs to the same fault type as each initial classification dataset, is retrieved from the cable fault association knowledge graph. This sequence records the changes and trends of core feature factors over time during the occurrence of similar faults in the past.

[0129] The data units in the initial classification dataset are arranged in chronological order and aligned with the historical evolution sequence of core feature factors in the time dimension. By comparing the changes in the core feature factors of the data units with the historical evolution trend, those data units that conform to the historical evolution trend are retained, while those that do not conform to the trend are removed. This improves the accuracy of data classification and eliminates data that may be inconsistent with the fault evolution trend due to interference or abnormal collection.

[0130] Step S136: Based on the cable section identifier, perform spatial clustering on the retained data units with basic spatiotemporal information to form spatial classification subsets with cable sections as units. Each spatial classification subset contains multiple time window data of the same fault type within the cable section.

[0131] For the data units retained after time-dimensional alignment, spatial clustering is performed based on their cable segment identifiers. All data units belonging to the same cable segment are grouped together to form spatial classification subsets based on cable segment. Each spatial classification subset contains data from multiple time windows within that cable segment that belong to the same fault type.

[0132] Spatial clustering enables data to be further organized according to spatial location, facilitating the analysis of differences in data characteristics across different cable sections under the same fault type. For example, a spatial subset of a cable joint section may contain data related to joint overheating faults across multiple time windows. Analyzing this data can reveal the fault occurrence patterns and characteristics of that joint section.

[0133] Step S137: Integrate each spatial classification subset and its corresponding spatiotemporal label to generate cable data classification results with spatiotemporal labels. The spatiotemporal labels include the data acquisition time window, cable section number, and fault type association level.

[0134] All spatial classification subsets are integrated to form a complete cable data classification result. During the integration process, the spatiotemporal labels corresponding to each data unit are retained. These labels include the time window of data acquisition, the cable section number to which it belongs, and the association level with the fault type.

[0135] The fault type association level is determined based on the spatiotemporal correlation between the data unit and the fault type; the higher the correlation, the higher the association level. The cable data classification results with spatiotemporal tags clearly demonstrate the distribution of data in time, space, and fault type. Technicians can quickly obtain data information related to specific fault types within a specific time window and a specific cable section by querying these classification results.

[0136] Step S140: Combine the cable data classification results with the fault feature evolution trajectory to generate a cable fault location reasoning basis, which includes multi-source data spatiotemporal correlation logic and fault location derivation chain.

[0137] Step S141: Extract the spatiotemporal feature dataset corresponding to each fault type from the cable data classification results with spatiotemporal labels. The spatiotemporal feature dataset includes cable operation parameter fluctuation records, dynamic data change records of the cable's surrounding environment, and cable section correlation data for different time windows under the fault type.

[0138] Extracting spatiotemporal feature datasets from cable data classification results with spatiotemporal labels is a fundamental step in generating inferences. For each fault type, all associated data units are selected, carrying information about different time windows and cable sections. These data units are then organized according to chronological order and spatial location to form the spatiotemporal feature dataset corresponding to that fault type.

[0139] The cable operating parameter fluctuation records include changes in parameters such as current, voltage, and temperature within different time windows, reflecting changes in the cable's operating state before and after a fault. The dynamic data change records of the cable's surrounding environment cover changes in environmental factors such as temperature, humidity, vibration, and pollutant concentration, helping to analyze the impact of the environment on the fault. Cable section association data records information about the cable section to which the data unit belongs. For example, for cable insulation aging faults, the spatiotemporal feature dataset may contain records of gradually increasing temperature, humidity changes, and corresponding current and voltage fluctuations in the fault-associated section across multiple time windows.

[0140] Step S142: Perform feature evolution analysis on the spatiotemporal feature dataset to construct a fault feature evolution trajectory. The fault feature evolution trajectory includes the changing trend of core feature factors in the time dimension, the diffusion path in the spatial dimension, and the evolution law of the correlation strength between core feature factors.

[0141] Feature evolution analysis of spatiotemporal feature datasets aims to uncover the changing patterns of fault features over time and space, thereby constructing fault feature evolution trajectories. In the time dimension, the analysis examines the numerical changes of core feature factors (such as temperature and current fluctuation amplitude) across different time windows to determine whether they exhibit trends of increase, decrease, or periodic change. These trends reflect the speed and stage of fault development.

[0142] Spatially, by tracking cable segments where fault-related characteristics appear in different time windows, the path of fault propagation from its initial location to the surrounding areas is depicted, revealing the direction and extent of fault propagation. Simultaneously, analyzing the changes in the correlation strength between core characteristic factors—for example, whether the correlation between current fluctuations and temperature increases or decreases during fault development—helps to understand the internal mechanisms of the fault. By synthesizing the results of these three analyses, a complete fault characteristic evolution trajectory is constructed, showcasing the dynamic development process of the fault.

[0143] Step S143: Retrieve the feature evolution trajectory library of historical fault cases from the cable fault association knowledge graph, dynamically match the current fault feature evolution trajectory with the trajectory in the feature evolution trajectory library of historical fault cases, calculate the similarity between the current fault feature evolution trajectory and the trajectory in the feature evolution trajectory library of historical fault cases, and select the historical fault case with the highest similarity as a reference.

[0144] The feature evolution trajectory database of historical fault cases is retrieved from the cable fault association knowledge graph to provide historical experience reference for current fault analysis. The historical trajectory database stores a large number of feature evolution trajectories of various types of faults that have occurred in the past. These trajectories contain typical patterns of different fault types in terms of time, space and feature association.

[0145] The current fault feature evolution trajectory is dynamically matched with trajectories in the historical trajectory database. The similarity score is calculated by comparing the similarity in core feature factor change trends, spatial diffusion paths, and the evolutionary patterns of inter-feature correlation strength. Higher similarity indicates a closer resemblance between the current and historical faults in their development patterns. Historical fault cases with the highest similarity are selected as references, allowing for the application of fault location experience and handling methods from these historical cases.

[0146] Step S144: Based on the differences between the current spatiotemporal feature dataset and the reference, construct a multi-source data spatiotemporal correlation logic. The multi-source data spatiotemporal correlation logic includes correlation rules between the current spatiotemporal feature dataset and the reference in terms of time window offset, spatial segment differences, and changes in the intensity of core feature factors.

[0147] Step S1441: Extract the spatiotemporal feature dataset of historical failure cases from the reference data, and align the spatiotemporal feature dataset of historical failure cases with the current spatiotemporal feature dataset in terms of dimensions.

[0148] After extracting the spatiotemporal feature dataset of historical failure cases from the reference, it is necessary to align it with the current spatiotemporal feature dataset in terms of dimensions. Dimension alignment is to ensure that the two datasets are consistent in terms of data structure, feature types, and temporal and spatial partitioning, so as to enable effective comparison and difference analysis.

[0149] Specifically, the data formats and feature metrics of both datasets are standardized to ensure that features such as current, voltage, and temperature in the historical dataset are defined and calculated in the same way as their corresponding features in the current dataset. Simultaneously, the criteria for dividing the time window and the precision of cable segment division are adjusted to ensure that the time window length and segment division length are consistent between the historical and current datasets. Dimensional alignment is used to eliminate comparison errors caused by differences in data structures.

[0150] Step S1442: Calculate the time window offset, compare the occurrence time of fault features in the current spatiotemporal feature dataset with that in the historical fault case spatiotemporal feature dataset, determine the offset difference between the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset in the time dimension, and construct time association rules based on the offset difference. The time association rules include the correspondence between the offset difference and the fault propagation speed.

[0151] When calculating the time window offset, the occurrence time of each fault feature in the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset are compared one by one. For each core feature factor, the time window in which it first appears in the current dataset and the time window in which it first appears in the historical dataset are found, and the interval between the two time windows is calculated, which is the time window offset difference of the feature factor.

[0152] By combining the offset differences of all core feature factors, the overall time window offset is determined. Based on the analysis of a large amount of historical data, the correlation between the offset difference and the fault propagation speed is summarized. For example, the larger the offset difference, the faster or slower the fault propagation speed may be. Based on this, time association rules are constructed to describe the association patterns in the time dimension.

[0153] Step S1443: Calculate the spatial segment difference degree, analyze the cable segment differences in fault feature distribution between the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset under the same time window, count the number of cable segments with different fault feature distributions and the feature intensity of each cable segment with different fault feature distributions, and construct spatial association rules based on the spatial segment difference degree. The spatial association rules include the association between the number of cable segments with different fault feature distributions and the fault location deviation.

[0154] When calculating the spatial segment dissimilarity, within the same time window, cable segments with fault feature distributions in the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset are compared. Cable segments that exhibit fault feature distributions in the current dataset but not in the historical dataset, or vice versa, are identified as those with dissimilarity in fault feature distributions.

[0155] The number of these discrepancy segments is counted, and the intensity of fault characteristics within each segment (such as the magnitude of temperature rise, current fluctuation, etc.) is measured. Based on historical experience and data analysis, spatial association rules are established. These rules describe the relationship between the number of discrepancy segments and the deviation between the actual fault location and the fault location in the reference case. For example, the more discrepancy segments there are, the greater the deviation in fault location may be.

[0156] Step S1444: Calculate the intensity change rate of core feature factors, compare the numerical differences of the corresponding core feature factors between the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset, calculate the percentage difference of the numerical differences for each core feature factor, and perform normalization processing to obtain comparable intensity change rates of core feature factors. Based on the normalized intensity change rates of core feature factors, construct feature association rules, which include the mapping relationship between the intensity change rate of core feature factors and the severity of the fault.

[0157] When calculating the rate of change of the intensity of core feature factors, for each core feature factor, within the same time window and cable segment, the values ​​in the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset are compared. The numerical difference between the two is calculated, and the difference is compared with the values ​​in the historical dataset to obtain the percentage difference.

[0158] The percentage differences of all core feature factors are normalized to a common numerical range, yielding comparable rates of change in core feature factor intensity. By analyzing the relationship between the rate of change in intensity and the actual severity of the fault (such as the scope of impact and the difficulty of repair), a feature association rule is constructed. This rule reflects the indicative role of feature intensity changes in fault severity.

[0159] Step S1445: Integrate the temporal association rules, spatial association rules, and feature association rules to construct a multi-source data spatiotemporal association logic framework. Each association rule in the multi-source data spatiotemporal association logic framework corresponds to a reasoning condition. If the reasoning condition is met, the corresponding fault location derivation direction is triggered.

[0160] By integrating temporal association rules, spatial association rules, and feature association rules, a spatiotemporal association logic framework for multi-source data is constructed. In this framework, each association rule is treated as an independent inference condition; when data meets the condition of a certain rule, the corresponding fault location derivation direction is triggered.

[0161] For example, when the time window offset meets a certain condition in the time association rule, the derivation direction may be biased towards areas with faster fault propagation; when the spatial segment difference meets a certain condition in the spatial association rule, the derivation direction may be adjusted to concentrate towards the difference segments. Through the above fusion method, association rules of different dimensions are integrated into a unified logical framework, enabling the reasoning process to comprehensively consider multiple factors such as time, space, and features.

[0162] Step S1446: Prioritize the association rules in the spatiotemporal association logic framework of multi-source data. Determine the priority based on the impact of the association rules on the fault location accuracy. The priority of temporal association rules is higher than that of spatial association rules, and the priority of spatial association rules is higher than that of feature association rules.

[0163] Prioritizing association rules within a multi-source data spatiotemporal correlation logic framework is crucial for determining the final derivation direction based on priority when conflicting derivation directions arise among different rules during the inference process. Priority is determined by the degree to which association rules influence fault location accuracy; rules with greater influence have higher priority.

[0164] In practical applications, time factors often have a significant impact on the timeliness and accuracy of fault location; therefore, time association rules have a higher priority than spatial association rules. The accuracy of spatial location directly affects the precision of fault location, so spatial association rules have a higher priority than feature association rules. This priority ranking ensures that rules with a greater impact on location accuracy are considered first during the inference process, thereby improving the accuracy of fault location.

[0165] Step S1447: Add a dynamic adjustment mechanism to the multi-source data spatiotemporal correlation logic framework. When real-time data updates cause changes in the applicable conditions of any correlation rule, adjust the inference parameters of that correlation rule to form a complete multi-source data spatiotemporal correlation logic.

[0166] A dynamic adjustment mechanism is added to the spatiotemporal correlation logic framework of multi-source data to adapt to changes brought about by real-time data updates. When new real-time data is received, the applicability conditions of each correlation rule are re-examined. If the applicability conditions of any rule change due to data updates, the inference parameters of that rule are adjusted accordingly.

[0167] For example, as real-time data is updated, the time window offset may change, causing the applicable conditions of the time association rules to change. In this case, it is necessary to recalculate the inference parameters and adjust the weights of the inference direction. Through the above dynamic adjustment mechanism, the spatiotemporal association logic of multi-source data can respond to data changes in real time, maintain the timeliness and accuracy of the inference logic, and ultimately form a complete spatiotemporal association logic of multi-source data.

[0168] Step S145: Construct a fault location derivation chain. The starting point of the fault location derivation chain is the current spatiotemporal feature data input, the intermediate links are the hierarchical application of spatiotemporal correlation logic of multi-source data, each intermediate link outputs the fault location candidate range, and the subsequent intermediate link narrows the fault location candidate range based on the result of the previous intermediate link.

[0169] Constructing a fault location derivation chain is a process of progressively narrowing down the fault location range. The starting point of the derivation chain is the current spatiotemporal characteristic data, which includes fault-related temporal, spatial, and characteristic information. The intermediate stages are hierarchical applications of spatiotemporal correlation logic for multi-source data. Each intermediate stage is based on the results of the previous stage and combines different correlation rules for inference and analysis.

[0170] The first intermediate step may initially determine a broad range of candidate fault locations based on temporal association rules and the temporal characteristics of the current data. The second intermediate step, building on the candidate range from the first step, applies spatial association rules to further narrow down the range. Subsequent steps continue to apply other association rules or more refined analyses, continuously narrowing down the candidate range until a relatively precise fault location range is obtained. Each intermediate step outputs a corresponding candidate fault location range, ensuring the traceability of the derivation process.

[0171] Step S146: Integrate the fault feature evolution trajectory, the spatiotemporal correlation logic of multi-source data, and the fault location derivation chain to form the basis for cable fault location reasoning.

[0172] By integrating the fault feature evolution trajectory, the spatiotemporal correlation logic of multi-source data, and the fault location deduction chain, a basis for cable fault location reasoning is formed. The fault feature evolution trajectory provides the overall context of fault development, the spatiotemporal correlation logic of multi-source data provides the rules and basis for reasoning, and the fault location deduction chain provides specific reasoning steps and the process of narrowing down the scope.

[0173] Step S147: Dynamically label the cable fault location reasoning basis, and label the corresponding spatiotemporal feature data source, the associated node of the cable fault association knowledge graph and the historical fault case reference identifier for each logical deduction link.

[0174] Dynamically labeling the reasoning basis for cable fault location aims to enhance the transparency and traceability of the reasoning process. At each logical deduction stage, the specific source of the spatiotemporal characteristic data used in that stage is clearly labeled, such as the sensor number used for data acquisition and the timestamp of data transmission, facilitating subsequent data verification and traceability.

[0175] Simultaneously, the associated nodes of the cable fault association knowledge graph upon which this step is based are labeled, explaining the specific nodes and edge weights in the knowledge graph referenced for reasoning. Furthermore, the identifiers of historical fault cases referenced are included, such as case number and occurrence time, ensuring that historical references during the reasoning process are traceable. Through dynamic labeling, technicians can clearly understand the basis and source of each reasoning step, increasing their confidence in the reasoning results and facilitating troubleshooting and correction when deviations occur in the reasoning.

[0176] Step S150: Perform dynamic simulation of the fault location based on the cable fault location reasoning, determine the real-time location information of the cable fault, generate a cable fault location instruction containing the dynamic coordinates of the fault location, and send the cable fault location instruction containing the dynamic coordinates of the fault location to the cable maintenance terminal.

[0177] Step S151: Analyze the fault location derivation chain in the cable fault location reasoning basis, and extract the spatiotemporal constraints and fault feature matching requirements of each link in the fault location derivation chain.

[0178] Analyzing the fault location derivation chain in the reasoning process for cable fault location is fundamental to performing dynamic fault location simulation. The fault location derivation chain comprises multiple reasoning steps, each with its specific spatiotemporal constraints and fault feature matching requirements. The spatiotemporal constraints specify the applicable time window and cable section range for that step's reasoning, while the fault feature matching requirements clarify the numerical range and trend of the core characteristic factors that this step must satisfy.

[0179] By analyzing the derivation chain, the conditions and requirements of each step are extracted one by one, forming a clear list of derivation rules. For example, the spatiotemporal constraints of a certain step may be a specific time window and several consecutive cable sections, and the fault characteristic matching requirement may be that the temperature exceeds a certain threshold and the current fluctuation amplitude is within a specific range.

[0180] Step S152: Extract the core fault feature data with spatiotemporal tags from the cable fault location reasoning basis. The core fault feature data with spatiotemporal tags includes the peak fluctuation data of cable operating parameters in the current time window, the environmental disturbance data of the corresponding cable section, and the fault propagation trend data.

[0181] Extracting core fault feature data with spatiotemporal labels from the cable fault location inference basis is crucial input for dynamic fault location simulation. Peak fluctuation data of cable operating parameters within the current time window reflects extreme changes in cable operating parameters at that moment, indicating the level of fault activity.

[0182] The environmental disturbance data for the corresponding cable section includes changes in environmental factors in the fault-related section, such as sudden changes in temperature and humidity, and increased vibration, which helps to analyze the impact of the environment on the fault location. The fault propagation trend data, based on the fault characteristic evolution trajectory, predicts the direction and speed of fault development in time and space.

[0183] Step S153: Input the fault core feature data with spatiotemporal tags into the starting link of the fault location derivation chain, and select the initial fault location candidate range that meets the requirements according to the spatiotemporal constraints in the starting link of the fault location derivation chain. The initial fault location candidate range covers multiple continuous cable sections.

[0184] After the core fault feature data with spatiotemporal labels is input into the initial stage of the fault location derivation chain, the initial stage will filter the data according to preset spatiotemporal constraints. The spatiotemporal constraints typically include the time window range and cable section range corresponding to the data; only data that meets these ranges will be included in the initial screening.

[0185] For example, the time constraints for the initial stage might be limited to the current time window and the previous few consecutive time windows, while the spatial constraints might be limited to the cable segment associated with the core fault feature data and several adjacent segments. By comparing the core fault feature data with these constraints, all cable segments that meet the conditions are selected, and these segments together constitute the initial candidate range for fault location. The initial candidate range typically covers multiple consecutive cable segments.

[0186] Step S154: Proceed to the next stage of the fault location derivation chain. Based on the fault feature evolution trajectory analysis, analyze the change pattern of the core fault feature data within the initial fault location candidate range, eliminate cable sections whose core fault feature data does not conform to the fault feature evolution trajectory trend, and narrow down the initial fault location candidate range to a preset number of cable sections.

[0187] After moving to the next stage of the fault location derivation chain, the focus is on analyzing the cable segments within the initial fault location candidate range based on the fault feature evolution trajectory. The fault feature evolution trajectory includes the changing trends of core feature factors in the time dimension and the diffusion path in the spatial dimension. By comparing the changing patterns of the core fault feature data of each cable segment within the initial candidate range with the trajectory trend, segments that do not conform to the trajectory trend can be identified.

[0188] For example, if the fault feature evolution trajectory shows that the temperature characteristic should show a continuous upward trend over time, but the temperature data of a certain cable segment within the initial candidate range shows a downward trend, then that segment does not conform to the evolution trajectory trend and can be eliminated. By gradually eliminating segments that do not meet the conditions in the above manner, the initial fault location candidate range is narrowed down to a preset number of cable segments, making the candidate range more focused.

[0189] Step S155: Based on the historical fault cases in the fault location derivation chain, compare the core fault feature data of the current narrowed fault location candidate range with the fault location features of similar historical fault cases, calculate the similarity between the core fault feature data and the fault location features of similar historical fault cases, and retain the first preset number of cable sections with the highest similarity as the intermediate fault location candidate range.

[0190] Based on historical fault case studies, the narrowed candidate fault locations are further refined. A detailed comparison is made between the core fault feature data of each cable segment within the current candidate range and the feature data of fault locations in similar historical fault cases. This comparison includes aspects such as the magnitude, rate of change, and correlation strength of core feature factors. The similarity between the two is calculated to measure the degree of matching between the current segment's features and historical fault location features. The top 10 cable segments with the highest similarity are retained as intermediate fault location candidate ranges. This step leverages historical experience to further improve the accuracy of the candidate range, making fault location inferences more reliable.

[0191] Step S156: Based on the latest changes in the real-time cable operation status data, dynamically verify the candidate range of intermediate fault locations. If any cable segment shows new fault characteristics in the real-time cable operation status data, increase the priority of that cable segment in the candidate range of intermediate fault locations; otherwise, decrease the priority of that cable segment in the candidate range of intermediate fault locations.

[0192] Based on the latest updates to real-time cable operating status data, the candidate range for intermediate fault locations is dynamically verified. The latest changes in real-time data may contain new fault characteristic information, reflecting the latest development of the fault. For each cable segment within the intermediate candidate range, its latest real-time operating status data is checked. If new fault-related characteristics appear (such as sudden current drops, abnormal voltage fluctuations, etc.), the priority of that segment is increased, indicating that it is more likely to be a fault location. Conversely, if no new fault characteristics appear in the real-time data of a segment, and the changes in existing characteristics tend to level off, the priority of that segment is decreased. Through dynamic verification, the priority ranking of the candidate range for intermediate fault locations can reflect the latest dynamics of the fault in real time.

[0193] Step S157: Based on the priority ranking results of the intermediate fault location candidate range, select the cable segment with the highest priority as the core area of ​​the fault location, and determine the precise location of the fault in the core area of ​​the fault location by combining the distribution density of the core fault feature data within the core area of ​​the fault location.

[0194] Based on the priority ranking of the candidate fault locations, the cable segment with the highest priority is selected as the core fault location area. Within this core area, the distribution density of the core fault feature data is further analyzed; areas with higher distribution density are more likely to experience a fault.

[0195] For example, by analyzing the distribution of core fault features such as temperature and vibration data at different locations within the core area, the specific location where the feature data is most concentrated and changes most significantly can be identified, and this location can be determined as the precise location of the fault.

[0196] Step S158: Integrate the core area identifier of the fault location with the precise coordinates of the fault location to generate real-time location information of the cable fault including a timestamp.

[0197] The system integrates the identification information of the core area of ​​the fault location (such as the section number) and the precise coordinates of the fault location (such as specific three-dimensional coordinate values), and adds a corresponding timestamp to generate complete real-time cable fault location information. The timestamp records the specific time when the location information was determined, ensuring the timeliness of the information.

[0198] Step S159: Generate a cable fault location instruction containing dynamic coordinates of the fault location based on the real-time location information of the cable fault. The cable fault location instruction also includes fault type, fault severity and fault development trend prediction information.

[0199] Based on real-time cable fault location information, a cable fault location command is generated, including dynamic coordinates of the fault location. In addition to dynamic coordinates, the command also integrates fault type information (such as insulation aging, mechanical damage, etc.), fault severity assessment results (such as mild, moderate, severe), and fault development trend information based on fault characteristic evolution trajectory and real-time data prediction (such as spread rate, likelihood of increased impact area, etc.). This information provides maintenance personnel with a comprehensive fault overview, helping them develop reasonable repair plans and improve the efficiency and targeted nature of fault handling.

[0200] Step S1510: Send the cable fault location command containing the dynamic coordinates of the fault location to the cable maintenance terminal through an encrypted transmission channel to ensure the security and integrity of the command during transmission.

[0201] An encrypted transmission channel is selected to send cable fault location commands to the cable maintenance terminal. Encryption prevents commands from being illegally intercepted or tampered with during transmission, ensuring the security and integrity of the commands. A verification mechanism is employed during transmission to check the integrity of the command data, ensuring the accuracy of the commands received by the maintenance terminal. Upon receiving the command, the cable maintenance terminal displays information such as the dynamic coordinates, type, and severity of the fault, providing precise guidance for maintenance personnel's on-site repair work and enabling rapid fault response and handling.

[0202] Figure 2 The illustration shows exemplary hardware and software components of a big data-based cable fault location system 100 that can implement the ideas of this application, according to some embodiments of this application. For example, a processor 120 can be used in the big data-based cable fault location system 100 and to perform the functions in this application.

[0203] For example, the big data-based cable fault location system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the big data-based cable fault location system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to the program instructions. The big data-based cable fault location system 100 also includes an I / O interface 150 between the computer and other input / output devices.

[0204] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned cable fault location method based on big data is implemented.

[0205] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A cable fault location method based on big data, characterized in that, The method includes: Acquire a multi-source cable dynamic data set, which includes real-time cable operating status data, historical fault evolution data, and dynamic data of the cable's surrounding environment; A cable fault association knowledge graph is constructed based on big data streaming technology. The cable fault association knowledge graph is used to establish a real-time association between a multi-source dynamic cable data set and cable fault types. The node weights of the cable fault association knowledge graph are dynamically updated as real-time data flows in. The multi-source cable dynamic data set is classified according to the spatiotemporal coupling rules of the cable fault association knowledge graph to obtain cable data classification results with spatiotemporal labels; The cable fault location reasoning basis is generated by combining the cable data classification results with the fault feature evolution trajectory. The cable fault location reasoning basis includes the spatiotemporal correlation logic of multi-source data and the fault location derivation chain. Based on the cable fault location reasoning, the fault location is dynamically deduced to determine the real-time location information of the cable fault, a cable fault location instruction containing the dynamic coordinates of the fault location is generated, and the cable fault location instruction containing the dynamic coordinates of the fault location is sent to the cable maintenance terminal.

2. The cable fault location method based on big data according to claim 1, characterized in that, The cable fault association knowledge graph constructed based on big data streaming technology includes: The data flow characteristics of the multi-source cable dynamic data set are analyzed to distinguish the streaming transmission frequency of real-time cable operation status data, the time span attribute of historical fault evolution data, and the collection cycle attribute of dynamic data of the cable's surrounding environment. The core feature factors related to cable fault evolution are extracted from each data stream. These core feature factors include real-time fluctuation characteristics of cable operating parameters, environmental disturbance characteristics at the time of historical fault occurrence, and changes in operating data during the fault propagation process. The initial topology of the cable fault association knowledge graph is constructed, with core feature factors as nodes of the cable fault association knowledge graph and the association frequency between core feature factors in historical fault cases as the initial edge weights, forming a node connection network. Deploy a big data streaming engine and connect real-time cable operation status data to the data stream channel of the big data streaming engine. When new core feature factors appear in the real-time data or the values ​​of existing core feature factors exceed the historical fluctuation range, trigger the node weight update mechanism of the node connection network. Based on the similarity matching results between real-time data and historical fault evolution data, the edge weights of corresponding nodes in the node connection network are adjusted. The higher the similarity, the greater the increase in edge weights, and vice versa, edge weight decay is triggered. By integrating the real-time updated node and edge weights, a cable fault association knowledge graph containing spatiotemporal dimension association rules is formed. The spatiotemporal dimension association rules include the difference in association strength of the same core feature factor in different time windows and different cable sections.

3. The cable fault location method based on big data according to claim 2, characterized in that, The extraction of core feature factors related to cable fault evolution from each data stream includes: The data transmission protocol and data structure of each data stream are analyzed to identify the fields directly related to the cable fault evolution. These fields include the real-time cable current field, the real-time cable voltage field, the real-time cable temperature field, the temperature and humidity change rate of the cable's surrounding environment field, and the fault propagation speed field. Extract the parameter change curves within a continuous time window from the real-time current field and real-time voltage field of the cable, calculate the slope change rate and peak frequency of the parameter change curves, and use the slope change rate and peak frequency of the parameter change curves as the core components of the real-time fluctuation characteristics of cable operating parameters. The mutation records of the dynamic data of the cable surrounding environment before the fault occurred in the historical fault evolution data are screened, and the abnormal increments of the cable surrounding environment parameters in the mutation records are extracted. The abnormal increments of the cable surrounding environment parameters include the short-term sudden rise and fall of temperature and humidity in the cable surrounding environment, the instantaneous increase of external mechanical force in the cable surrounding environment, and the sudden increase of pollutant concentration in the cable surrounding environment. The abnormal increments of the cable surrounding environment parameters are used as the environmental disturbance characteristics at the time of the historical fault. The cable operating status data during the fault propagation stage in the historical fault evolution data is collected. The parameter attenuation law of the cable operating status data with the fault propagation distance is extracted. The parameter attenuation law with the fault propagation distance is used as the operating data change feature during the fault propagation process. The parameter attenuation law with the fault propagation distance includes the decreasing gradient of cable current with the fault propagation distance, the loss rate of cable voltage with the fault propagation distance, and the changing trend of cable temperature with the fault propagation distance. Spatiotemporal correlation verification was performed on the extracted cable operating parameters, environmental disturbance characteristics at the time of historical fault occurrence, and operating data change characteristics during the fault propagation process. Feature items directly related to the cable fault evolution timeline and cable fault spatial propagation path were retained. Dynamic weights are assigned to the feature items after spatiotemporal correlation verification. The weights are determined based on the frequency of the feature items in the real-time data stream and the lead time of cable fault warning, forming an initial set of core feature factors. The higher the frequency and the greater the lead time of cable fault warning, the higher the weight. The core feature factors in the initial set of core feature factors are back-matched with historical fault cases to verify the contribution of each core feature factor to the determination of cable fault type. The core feature factors with the highest contribution to the determination of cable fault type are retained according to a preset proportion, and finally the core feature factors related to the evolution of cable faults are determined.

4. The cable fault location method based on big data according to claim 3, characterized in that, The process involves reverse matching the core feature factors in the initial set of core feature factors with historical fault cases to verify the contribution of each core feature factor to the determination of cable fault type. The core feature factors with the highest contribution to cable fault type determination, representing a predetermined proportion, are retained. Finally, the core feature factors related to cable fault evolution are determined, including: Extract complete case data of multiple types of cable faults from the historical fault case database. Each complete case data includes the record of changes in the core characteristic factors of the entire cable fault occurrence process and the final cable fault type label. For each core feature factor in the initial set of core feature factors, an independent validation model is constructed. The input of the independent validation model is the change record of the core feature factor in the complete case data, and the output of the independent validation model is the cable fault type prediction result. The complete case data in the historical fault case database is divided into a training set and a validation set according to a preset ratio. The training set is used to train each independent validation model. The parameters of the independent validation model are adjusted until the cable fault type prediction accuracy of the training set reaches a preset threshold. The trained independent validation models were tested using a validation set, and the model performance metrics for each independent validation model for different cable fault types were recorded. A comprehensive scoring model for the contribution of core feature factors to cable fault type determination is constructed. After standardizing the performance indicators of each model of the independent verification model, the model is weighted according to the preset weights to obtain the comprehensive contribution score of each core feature factor. The comprehensive contribution scores of all core feature factors are sorted in descending order, and the core feature factors with the highest comprehensive contribution scores are selected as candidate core feature factors. Candidate core feature factors are cross-validated by substituting them into new historical fault cases. The overall accuracy of the combination of candidate core feature factors in determining cable fault type is checked. If the overall accuracy of the combination of candidate core feature factors in determining cable fault type meets the preset requirements, it is determined as the final core feature factor related to cable fault evolution. If the overall accuracy of the combination of candidate core feature factors in determining cable fault type does not meet the preset requirements, the proportion of candidate core feature factors is increased, and the cross-validation process is repeated until the overall accuracy of the combination of candidate core feature factors in determining cable fault type meets the standard. Finally, the core feature factor related to cable fault evolution is determined.

5. The cable fault location method based on big data according to claim 1, characterized in that, The classification of the multi-source cable dynamic data set according to the spatiotemporal coupling rules of the cable fault association knowledge graph yields cable data classification results with spatiotemporal labels, including: The spatiotemporal coupling classification rules are extracted from the cable fault association knowledge graph, and the core feature factors in the spatiotemporal coupling classification rules, such as the time correlation threshold, spatial segment matching accuracy, and fault type association confidence requirements, are analyzed. Each data point in the multi-source cable dynamic data set is labeled with a collection timestamp and a cable section identifier to form a data unit with basic spatiotemporal information. According to the spatiotemporal coupling classification rules, the core feature factors of each data unit with basic spatiotemporal information are matched with the nodes in the same time window and the same cable segment in the cable fault association knowledge graph. The spatiotemporal correlation degree between the data unit with basic spatiotemporal information and each fault type is calculated. The calculation of the spatiotemporal correlation degree between the data unit with basic spatiotemporal information and each fault type is based on the time synchronization of the core feature factors, the spatial location overlap, and the edge weights of the cable fault association knowledge graph. Data units with basic spatiotemporal information are selected, and the spatiotemporal correlation between them and each fault type meets the correlation confidence requirement of the fault type. These data units are then grouped according to fault type to form an initial classification dataset, and each initial classification dataset carries a corresponding time window label and cable section label. The evolution sequence of historical core feature factors of the same fault type is called from the cable fault association knowledge graph. The data units with basic spatiotemporal information in the initial classification dataset are aligned with the evolution sequence of historical core feature factors in the time dimension, and the data units with basic spatiotemporal information that conform to the evolution trend of historical core feature factors are retained. Based on the cable section identifier, the data units with basic spatiotemporal information are spatially clustered to form spatial classification subsets with cable sections as the unit. Each spatial classification subset contains multiple time window data of the same fault type within the cable section. By integrating the spatial classification subsets and their corresponding spatiotemporal labels, cable data classification results with spatiotemporal labels are generated. The spatiotemporal labels include the data acquisition time window, cable section number, and fault type association level.

6. The cable fault location method based on big data according to claim 5, characterized in that, According to the spatiotemporal coupling classification rules, the core feature factors of each data unit with basic spatiotemporal information are matched with nodes in the same time window and the same cable segment in the cable fault association knowledge graph, and the spatiotemporal correlation degree between the data unit with basic spatiotemporal information and each fault type is calculated, including: The time window division criteria and spatial segment division precision are extracted from the spatiotemporal coupling classification rules. The time dimension of the multi-source cable dynamic data set is divided into continuous time windows according to the time window division criteria, and the spatial dimension of the multi-source cable dynamic data set is divided into cable segments of equal precision according to the spatial segment division precision. Assign a corresponding time window identifier and spatial segment identifier to each data unit with basic spatiotemporal information; Extract the core feature factor set of each data unit with basic spatiotemporal information, match each core feature factor in the core feature factor set with the corresponding node in the same time window and the same cable segment in the cable fault association knowledge graph, and record the number of successfully matched core feature factors and the current edge weight of each matched node. To calculate the time synchronization of core feature factors, compare the difference between the acquisition time of the core feature factors of the data unit with basic spatiotemporal information and the latest update time of the knowledge graph node related to cable faults. The smaller the difference, the higher the time synchronization of the core feature factors. The value of the time synchronization of core feature factors is calculated by inverting the ratio of the difference to the length of the time window. The spatial location overlap is calculated by determining the overlap ratio between the location of the data unit with basic spatiotemporal information and the cable segment range associated with the cable fault association knowledge graph node, based on the spatial segment identifier of the data unit with basic spatiotemporal information and the cable segment associated with the cable fault association knowledge graph node. The overlap ratio is the spatial location overlap. A spatiotemporal correlation calculation model is constructed between data units with basic spatiotemporal information and various fault types. After normalizing the sum of edge weights of successfully matched core feature factors, the time synchronization value of core feature factors, and the spatial location overlap, the data units are weighted and fused according to preset weight coefficients. The weighted fusion result is the spatiotemporal correlation between data units with basic spatiotemporal information and corresponding fault types. Repeat the matching and calculation process for each data unit with basic spatiotemporal information to obtain the spatiotemporal correlation degree values ​​between each data unit with basic spatiotemporal information and each fault type, and establish a correspondence table between the data unit and each fault type.

7. The cable fault location method based on big data according to claim 1, characterized in that, The method of generating cable fault location inference based on the cable data classification results and fault feature evolution trajectory includes: The spatiotemporal feature datasets corresponding to each fault type are extracted from the cable data classification results with spatiotemporal labels. The spatiotemporal feature datasets include cable operating parameter fluctuation records, dynamic data change records of the cable's surrounding environment, and cable section correlation data for different time windows under the fault type. The spatiotemporal feature dataset is subjected to feature evolution analysis to construct a fault feature evolution trajectory. The fault feature evolution trajectory includes the changing trend of core feature factors in the time dimension, the diffusion path in the spatial dimension, and the evolution law of the correlation strength between core feature factors. The feature evolution trajectory library of historical fault cases is retrieved from the cable fault association knowledge graph. The current fault feature evolution trajectory is dynamically matched with the trajectory in the feature evolution trajectory library of historical fault cases. The similarity between the current fault feature evolution trajectory and the trajectory in the feature evolution trajectory library of historical fault cases is calculated, and the historical fault case with the highest similarity is selected as a reference. Based on the differences between the current spatiotemporal feature dataset and the reference, a multi-source data spatiotemporal correlation logic is constructed. The multi-source data spatiotemporal correlation logic includes the correlation rules between the current spatiotemporal feature dataset and the reference in terms of time window offset, spatial segment differences, and changes in the intensity of core feature factors. A fault location derivation chain is constructed. The starting point of the fault location derivation chain is the current spatiotemporal feature data input, the intermediate links are the hierarchical application of spatiotemporal correlation logic of multi-source data, each intermediate link outputs the fault location candidate range, and the subsequent intermediate link narrows the fault location candidate range based on the result of the previous intermediate link. The evolution trajectory of fault characteristics, the spatiotemporal correlation logic of multi-source data, and the fault location derivation chain are integrated to form the basis for cable fault location reasoning; The reasoning basis for cable fault location is dynamically labeled, and each logical deduction step is labeled with the corresponding spatiotemporal feature data source, the associated node of the cable fault association knowledge graph, and the reference identifier of historical fault cases.

8. The cable fault location method based on big data according to claim 7, characterized in that, The construction of multi-source data spatiotemporal correlation logic based on the differences between the current spatiotemporal feature dataset and the reference data includes: Extract the spatiotemporal feature dataset of historical failure cases from the reference data, and align the spatiotemporal feature dataset of historical failure cases with the current spatiotemporal feature dataset in terms of dimensions. Calculate the time window offset, compare the occurrence time of fault features in the current spatiotemporal feature dataset with that in the historical fault case spatiotemporal feature dataset, determine the offset difference between the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset in the time dimension, and construct time association rules based on the offset difference. The time association rules include the correspondence between the offset difference and the fault propagation speed. Calculate the spatial segment difference degree, analyze the cable segment differences in fault feature distribution between the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset under the same time window, count the number of cable segments with different fault feature distributions and the feature intensity of each cable segment with different fault feature distributions, and construct spatial association rules based on the spatial segment difference degree. The spatial association rules include the correlation between the number of cable segments with different fault feature distributions and the fault location deviation. Calculate the rate of change of core feature factors intensity, compare the numerical differences of the corresponding core feature factors between the current spatiotemporal feature dataset and the historical fault case spatiotemporal feature dataset, calculate the percentage difference of the numerical differences for each core feature factor, and perform normalization to obtain comparable rates of change of core feature factors intensity. Based on the normalized rates of change of core feature factors intensity, construct feature association rules, which include the mapping relationship between the rate of change of core feature factors intensity and the severity of the fault. By integrating time association rules, spatial association rules, and feature association rules, a multi-source data spatiotemporal association logic framework is constructed. Each association rule in the multi-source data spatiotemporal association logic framework corresponds to a reasoning condition. If the reasoning condition is met, the corresponding fault location derivation direction is triggered. The association rules in the spatiotemporal association logic framework of multi-source data are prioritized and the priority is determined based on the degree of influence of the association rules on the fault location accuracy. The priority of temporal association rules is higher than that of spatial association rules, and the priority of spatial association rules is higher than that of feature association rules. A dynamic adjustment mechanism is added to the spatiotemporal correlation logic framework of multi-source data. When real-time data updates cause changes in the applicable conditions of any correlation rule, the inference parameters of that correlation rule are adjusted to form a complete spatiotemporal correlation logic of multi-source data.

9. The cable fault location method based on big data according to claim 1, characterized in that, The step of performing dynamic fault location deduction based on the cable fault location reasoning to determine the real-time location information of the cable fault includes: The fault location derivation chain in the cable fault location reasoning basis is analyzed, and the spatiotemporal constraints and fault feature matching requirements of each link in the fault location derivation chain are extracted. Extract the core fault feature data with spatiotemporal tags from the cable fault location reasoning basis. The core fault feature data with spatiotemporal tags includes the peak fluctuation data of cable operating parameters in the current time window, the environmental disturbance data of the corresponding cable section, and the fault propagation trend data. The fault core feature data with spatiotemporal tags is input into the starting link of the fault location derivation chain. The initial fault location candidate range that meets the requirements is selected according to the spatiotemporal constraints in the starting link of the fault location derivation chain. The initial fault location candidate range covers multiple consecutive cable sections. Moving to the next stage of the fault location derivation chain, based on the fault feature evolution trajectory analysis, the change pattern of the core fault feature data within the initial fault location candidate range is analyzed, and cable sections whose core fault feature data do not conform to the fault feature evolution trajectory trend are excluded, thus narrowing the initial fault location candidate range to a preset number of cable sections. Based on the historical fault cases in the fault location derivation chain, the core fault feature data of the current narrowed fault location candidate range is compared with the fault location features of similar historical fault cases. The similarity between the core fault feature data and the fault location features of similar historical fault cases is calculated, and the first preset number of cable sections with the highest similarity are retained as the intermediate fault location candidate range. Based on the latest changes in real-time cable operation status data, the candidate range of intermediate fault locations is dynamically verified. If any cable segment shows new fault characteristics in its real-time cable operation status data, the priority of that cable segment in the candidate range of intermediate fault locations is increased; otherwise, the priority of that cable segment in the candidate range of intermediate fault locations is decreased. Based on the priority ranking results of the intermediate fault location candidate range, the cable section with the highest priority is selected as the core area of ​​the fault location. Combining the distribution density of the core fault feature data within the core area of ​​the fault location, the precise location of the fault within the core area of ​​the fault location is determined. By integrating the core area identifier of the fault location with the precise coordinates of the fault location, real-time location information of the cable fault, including timestamps, is generated.

10. A cable fault location system based on big data, characterized in that, The big data-based cable fault location system includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the big data-based cable fault location method according to any one of claims 1-9.