A log data processing method based on cable phase measurement

By constructing a multi-dimensional feature vector set and a hierarchical repair strategy, the abnormal problem of cable phase data caused by interference in complex power systems is solved, efficient and accurate data repair and real-time monitoring are achieved, and the data quality and intelligence level of the power system are improved.

CN120179518BActive Publication Date: 2025-10-14FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510084398.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-10-14
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Cable phase data is susceptible to interference during long-term operation, resulting in data anomalies, missing data, noise, and drift. Traditional methods cannot meet the real-time and intelligent requirements of complex power systems, especially under fault conditions, which makes phase data analysis more difficult.

Method used

Construct a multidimensional feature vector set, establish a classification benchmark for short-term and long-term feature anomaly patterns, identify anomalies in real time by setting a log monitoring sliding window, use interpolation algorithm and polynomial fitting and segmented regression to perform layered repair, generate a repair value sequence to reconstruct the data, and form an adaptive closed-loop repair mechanism.

Benefits of technology

It achieves accurate identification and hierarchical repair of short-term fluctuations and long-term drift anomalies, improves the real-time and accuracy of data processing, ensures the continuous reliability of data quality, and significantly improves the data support for power system status monitoring and dispatch optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179518B_ABST
    Figure CN120179518B_ABST
Patent Text Reader

Abstract

The application discloses a kind of log data processing methods based on cable phase measurement, specifically related to log data processing field, including: constructing the feature vector set containing multidimensional description, the classification benchmark of short-term, long-term feature anomaly mode of each cable number is established;Real-time monitoring of the latest log record of each cable, generate respectively corresponding short-term fluctuation, long-term drift anomaly recognition result;Based on short-term fluctuation anomaly characteristics, the first level repair strategy is established, and the repair value sequence is generated using interpolation algorithm to reconstruct the short-term fluctuation anomaly data of phase log;Based on long-term offset anomaly characteristics, the second level repair strategy is established, and the long-term offset anomaly data of phase log is reconstructed based on polynomial fitting and piecewise regression;After repair is executed, the data is re-identified anomaly, judge whether phase log data is repaired, if yes, end log data repair, if not, issue phase anomaly alarm signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of log data processing, and more particularly, to a log data processing method based on cable phase measurement. Background Art

[0002] Cable phase data logs are an important data source for monitoring the operating status of power systems, and their accuracy directly affects the system's operational evaluation, fault diagnosis, and optimized control. However, during long-term operation, cable phase data is susceptible to a variety of interference sources, including equipment failures, electromagnetic noise, sensor drift, and transient interference from the external environment. These problems may cause data quality issues such as anomalies, omissions, noise, and drift in the log data, compromising the accuracy and integrity of the data. Traditional phase data quality detection and repair methods rely heavily on static threshold determination and manual intervention, which cannot meet the real-time and intelligent requirements of complex power systems. Especially under fault conditions, inaccurate data anomalies will spread rapidly, making phase data analysis more difficult.

[0003] Therefore, how to propose a log data processing method based on cable phase measurement to maintain high data quality and stability under different interference conditions, provide reliable data support for real-time monitoring of power systems, and significantly reduce the need for manual intervention and improve the intelligence level of data quality management is an urgent problem to be solved.

[0004] In order to solve the above problems, a technical solution is now provided. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a log data processing method based on cable phase measurement to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] S1: Integrate the multidimensional features of the cable historical phase log data, construct a feature vector set containing multidimensional descriptions, and establish a classification benchmark for the short-term and long-term characteristic anomaly patterns of each cable number;

[0008] S2: Set a log monitoring sliding window to monitor the latest log records of each cable in real time and generate anomaly identification results corresponding to short-term fluctuations and long-term drifts respectively;

[0009] S3: Based on the short-term fluctuation anomaly characteristics, a first-level repair strategy is established, and an interpolation algorithm is used to generate a repair value sequence to reconstruct the short-term fluctuation anomaly data of the phase log;

[0010] S4: A second-level repair strategy is established based on the long-term offset anomaly characteristics. The trend repair data is generated based on polynomial fitting and piecewise regression to reconstruct the long-term offset anomaly data of the phase log.

[0011] S5: Re-identify the abnormalities of short-term fluctuations and long-term drifts on the data after the repair is completed, and determine whether the phase log data has been repaired. If so, end the log data repair; if not, issue a phase abnormality alarm signal.

[0012] In a preferred embodiment, in S1, the multi-dimensional features of the cable historical phase log data are integrated to construct a feature vector set containing multi-dimensional descriptions, and a classification benchmark for the short-term and long-term characteristic abnormal patterns of each cable number is established, specifically including:

[0013] Screen the records in the cable historical phase log data, delete the logs with missing items, incomplete formats and duplicate content, and obtain the original record set with complete data;

[0014] Read the original record set one by one, and create time series phase history log data groups corresponding to different cable numbers based on the data upload timestamp and cable number;

[0015] The time series phase historical log data corresponding to each cable number is grouped and analyzed independently. The static features of the data are extracted from the historical logs, including the mean, extreme value, standard deviation and fluctuation amplitude of the phase data, and a static feature description set is constructed.

[0016] Through time series analysis, the dynamic feature values ​​of historical logs are extracted, including the rate of change of phase data and the sliding window mean of a set window size, to construct a dynamic feature description set.

[0017] Through frequency domain analysis, the frequency domain feature values ​​of historical logs are extracted, including high-frequency amplitude and high-frequency energy ratio, and a frequency domain feature description set is constructed.

[0018] Integrate static features, dynamic features and frequency domain features to establish a feature vector set containing multi-dimensional descriptions;

[0019] The extracted log feature vector set corresponding to the normal operation period of the cable history is divided into short-term and long-term feature categories through feature clustering, and each category corresponds to a comprehensive feature group;

[0020] Based on the comprehensive feature grouping of different feature categories corresponding to the phase history log of each cable number, a classification benchmark for the short-term and long-term feature anomaly patterns of the phase history log of each cable number is established. The matrix expression corresponding to the classification benchmark is:

[0021]

[0022] Among them, D is the matrix corresponding to the classification benchmark, DS and DL are the classification benchmarks of short-term and long-term feature abnormal patterns respectively, and X max 、X min 、X rate 、X avg are the maximum, minimum, rate of change and mean of the phase data respectively, σ is the standard deviation of the phase data, F high 、E high They are the high-frequency amplitude and high-frequency energy ratio of the phase data frequency domain, W avg is the sliding window mean.

[0023] In a preferred embodiment, in S2, a log monitoring sliding window is set to monitor the latest log records of each cable in real time, and the abnormality identification results corresponding to short-term fluctuations and long-term drifts are generated, specifically including:

[0024] Use the cable number as the search identifier to set the cable log monitoring sliding window to obtain the latest log records of each cable in real time;

[0025] For the log data in each monitoring sliding window, a feature vector matrix based on the feature anomaly classification benchmark is constructed. The data in the sliding window is mapped to the feature space one by one, and the key features representing different anomaly patterns are extracted.

[0026] The multi-center iterative distribution method is used to cluster the key features representing different anomaly patterns. The Euclidean distance to the classification benchmark is calculated within the feature vector matrix. A distance offset anomaly threshold is set. By comparing the thresholds, identification results corresponding to short-term fluctuations and long-term drift anomalies are generated. A repair task is established for each log data entry identified as an anomaly.

[0027] According to the characteristic deviation of the anomaly identification results and the order of the anomaly impact, the weight distribution algorithm is used to prioritize the selection of phase log data repair paths and generate a priority sequence for the execution of the repair paths, specifically:

[0028] Obtain the log data anomaly recognition results corresponding to all cable numbers, and uniformly convert each type of anomaly recognition result into a feature offset ratio expression;

[0029] Preset the impact weight coefficients of short-term fluctuations and long-term drift anomalies, perform weighted summation on the abnormal characteristic offset ratios of short-term fluctuations and long-term drifts, calculate the comprehensive anomaly index, and adjust the task sequence number of the cable phase log data repair task according to the reverse order of the comprehensive anomaly index;

[0030] Calculate the single weighted offset of the short-term fluctuation and long-term drift anomaly characteristics of the phase log data corresponding to each task sequence number, perform secondary sorting in reverse order of the weighted offset size, and add the corresponding repair tasks to the log data repair queue according to the final sorting results. The distributed lock is used to control the synchronous execution of the repair tasks.

[0031] In a preferred embodiment, in S3, a first-level repair strategy is established based on the short-term fluctuation anomaly characteristics, and an interpolation algorithm is used to generate a repair value sequence to reconstruct the short-term fluctuation anomaly data of the phase log, specifically including:

[0032] When the executor of the first-level repair strategy obtains the distributed lock, it obtains the sliding log data segment in the log monitoring sliding window corresponding to the first-priority repair task in the log data repair queue;

[0033] The phase fluctuation amplitude in the static features of the classification benchmark is extracted as the upper and lower limits of the abnormal fluctuation range. The data exceeding the upper and lower limits in the sliding log data segment are defined as short-term fluctuation abnormal data points, and their positions are marked.

[0034] An interpolation algorithm is selected as the repair basis. Based on the time index of the short-term fluctuation abnormal data point and its adjacent non-abnormal data points, a local interpolation matching the input dimension with the abnormal fluctuation range is constructed to generate the corresponding repair value sequence.

[0035] The generated repair values ​​are used to replace the short-term fluctuation abnormal data points in sequence, and the reconstructed data are inserted continuously step by step in chronological order;

[0036] At the edge of the repaired abnormal data point, the transition interval value between the repaired value and the adjacent non-abnormal data point is recalculated through weighted moving average to generate a smoothed local data sequence update log and release the distributed lock.

[0037] In a preferred embodiment, in S4, a second-level repair strategy is established based on the long-term offset anomaly characteristics, and the long-term offset anomaly data of the phase log is reconstructed by generating trend repair data based on polynomial fitting and piecewise regression. Specifically, the strategy includes:

[0038] When the executor of the second-level repair strategy obtains the distributed lock, it obtains the sliding log data segment in the log monitoring sliding window corresponding to the first-priority repair task in the log data repair queue;

[0039] Scan the time series data of the sliding log data segment one by one, and identify abnormal intervals of continuous deviation by calculating the degree of match between the dynamic deviation trend of the data point and the sliding mean change trend in the benchmark dynamic feature, and generate corresponding interval identifiers;

[0040] Based on polynomial fitting and segmented regression methods, a mathematical model is constructed to describe the phase change trend of the abnormal interval, and the trend is described by the changes in the key characteristics of the abnormal pattern of the abnormal interval;

[0041] An iterative optimization algorithm is used to gradually adjust the parameter set of the fitting model, with the goal of minimizing the relative residual of the fitting relative to the classification benchmark features of the long-term offset anomaly, and constructing a consistent trend expression between the benchmark features and the key features of the anomaly pattern;

[0042] Based on the tuned trend fitting model, regression calculation is performed on each time point in the abnormal interval to generate a repair value sequence covering all time points in the abnormal interval. The abnormal data is replaced with the repair value sequence data in sequence, and the distributed lock is released after the replacement is completed.

[0043] In a preferred embodiment, in S5, the data after the repair is completed is re-identified for short-term fluctuations and long-term drift anomalies to determine whether the phase log data is repaired. If so, the log data repair is terminated. If not, a phase anomaly alarm signal is issued, specifically including:

[0044] When all tasks in the log data repair queue are completed, the repaired data is compared with the classification benchmark again to generate the corresponding short-term fluctuation and long-term drift anomaly identification results;

[0045] When the short-term fluctuation and long-term drift anomaly identification results are detected as non-abnormal, the data repair is terminated. When the short-term fluctuation or long-term drift anomaly identification results are detected as abnormal, a phase abnormality data alarm signal is sent to the maintenance personnel terminal. The alarm signal contains the specific time period and cable number corresponding to the phase log data where the abnormality occurred.

[0046] The technical effects and advantages of the log data processing method based on cable phase measurement of the present invention are as follows:

[0047] By constructing a multidimensional feature vector set and establishing a classification benchmark, this method achieves precise identification and layered repair of short-term fluctuations and long-term drift anomalies. This method uses a sliding window for log monitoring to achieve real-time monitoring and anomaly identification, improving the real-time and accuracy of data processing. For short-term fluctuation anomalies, an interpolation algorithm generates a repair value sequence to achieve local data reconstruction. For long-term drift anomalies, polynomial fitting and piecewise regression are used to generate trend repair data, accurately restoring the overall trend of the log data. This layered repair strategy enables this method to flexibly respond to different anomaly types, ensuring efficient and accurate repair.

[0048] Furthermore, the present invention employs an adaptive closed-loop repair mechanism. By re-identifying repaired data and dynamically evaluating the repair effect, this creates a repair-recognition closed loop. This mechanism issues an alarm signal when a complete repair is not possible, ensuring the continued reliability of data quality. This method effectively eliminates short-term fluctuations and long-term drift in cable phase log data, significantly improving data integrity and accuracy. This provides high-quality data support for power system condition monitoring and dispatch optimization, while also achieving a high level of automation and intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a schematic diagram of a log data processing method based on cable phase measurement according to the present invention. DETAILED DESCRIPTION

[0050] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0051] Example 1

[0052] Figure 1 The present invention provides a log data processing method based on cable phase measurement, which includes the following steps:

[0053] S1: Integrate the multidimensional features of the cable historical phase log data, construct a feature vector set containing multidimensional descriptions, and establish a classification benchmark for the short-term and long-term characteristic anomaly patterns of each cable number;

[0054] S2: Set a log monitoring sliding window to monitor the latest log records of each cable in real time and generate anomaly identification results corresponding to short-term fluctuations and long-term drifts respectively;

[0055] S3: Based on the short-term fluctuation anomaly characteristics, a first-level repair strategy is established, and an interpolation algorithm is used to generate a repair value sequence to reconstruct the short-term fluctuation anomaly data of the phase log;

[0056] S4: A second-level repair strategy is established based on the long-term offset anomaly characteristics. The trend repair data is generated based on polynomial fitting and piecewise regression to reconstruct the long-term offset anomaly data of the phase log.

[0057] S5: Re-identify the abnormalities of short-term fluctuations and long-term drifts on the data after the repair is completed, and determine whether the phase log data has been repaired. If so, end the log data repair; if not, issue a phase abnormality alarm signal.

[0058] In S1, the multidimensional features of the cable historical phase log data are integrated to construct a feature vector set containing multidimensional descriptions. The classification benchmarks for the short-term and long-term characteristic anomaly patterns of each cable number are established. Specifically, the following are included:

[0059] From the cable phase history log, read the original record set, check the log content one by one, identify and delete records containing missing items or incomplete formats (such as missing timestamps, missing cable numbers or phase data), and eliminate duplicate records (such as records with the same timestamp) to ensure data integrity and consistency.

[0060] Read the original record set one by one, and create time series phase history log data groups corresponding to different cable numbers based on the data upload timestamp and cable number;

[0061] The time-series phase historical log data corresponding to each cable number is grouped and analyzed independently. The static features of the data are extracted from the historical logs, including the mean, extreme value, standard deviation and fluctuation amplitude of the phase data. A static feature description set is constructed. These static features reflect the overall distribution characteristics of the data and help analyze the phase status of the cable.

[0062] For each cable ID's time series data, a sliding window approach is used to extract dynamic features. These features primarily include the rate of change of the phase data and the sliding average within a set window size (based on the amount of time series data and the time interval span). By analyzing the temporal trends of the data, dynamic features describing the data's changes are obtained.

[0063] Frequency domain analysis extracts frequency domain feature values ​​from historical logs, including high-frequency amplitude and high-frequency energy percentage, and constructs a set of frequency domain feature descriptions. High-frequency energy primarily reflects the noise component in the data, and frequency domain features help identify abnormal patterns. The frequency standard for high frequencies can be freely set based on recognition accuracy. Fast Fourier transform (FFT) is used to convert time domain data into the frequency domain and calculate the amplitude of high-frequency components in the frequency domain. The energy percentages of low and high frequencies are compared, and the high-frequency energy percentage is calculated using the following formula:

[0064]

[0065] Where RP is the proportion of high-frequency energy, X(f) is the amplitude of the frequency domain data, and f h It is the frequency standard of the set high frequency. Static features, dynamic features and frequency domain features are integrated to establish a feature vector set containing multi-dimensional description.

[0066] The extracted log feature vector set corresponding to the normal operation period of the cable history is divided into short-term and long-term feature categories through feature clustering, and each category corresponds to a comprehensive feature group;

[0067] Based on the comprehensive feature grouping of different feature categories corresponding to the phase history log of each cable number, a classification benchmark for the short-term and long-term feature anomaly patterns of the phase history log of each cable number is established. The matrix expression corresponding to the classification benchmark is:

[0068]

[0069] Among them, D is the matrix corresponding to the classification benchmark, DS and DL are the classification benchmarks of short-term and long-term feature abnormal patterns respectively, and X max 、X min 、X rate 、X avg are the maximum, minimum, rate of change and mean of the phase data respectively, σ is the standard deviation of the phase data, F high 、E high They are the high-frequency amplitude and high-frequency energy ratio of the phase data frequency domain, W avg is the sliding window mean.

[0070] In S2, a log monitoring sliding window is set to monitor the latest log records of each cable in real time, and generate abnormal identification results corresponding to short-term fluctuations and long-term drifts, including:

[0071] Using the cable number as a search identifier, a corresponding sliding window for log monitoring is set for each cable number to obtain the latest log records for each cable in real time. The monitoring window size can be set to a fixed time period (for example, 1 hour) or dynamically increased based on event triggers. Real-time data is obtained through APIs or stream processing frameworks such as Kafka.

[0072] The log data for each cable ID is mapped into a feature vector using static, dynamic, and frequency domain feature extraction algorithms. Each feature vector contains multiple dimensions, including the mean, extreme values, standard deviation, fluctuation amplitude, and sliding average of the phase data. The feature vector matrix provides a high-dimensional feature space representation of all monitoring data for each cable, ensuring that abnormal patterns in the logs can be captured.

[0073] The multi-center iterative distribution method (or K-means algorithm) is used to cluster the key features in the feature vector matrix. By calculating the Euclidean distance between each feature vector and the classification benchmark, a threshold is set to identify abnormal patterns of short-term fluctuations and long-term drift, and generate corresponding anomaly identification results.

[0074] Set the distance offset anomaly threshold, generate the corresponding short-term fluctuation and long-term drift anomaly identification results through threshold comparison, and establish a repair task for each log data with an abnormal identification result.

[0075] According to the characteristic deviation of the anomaly identification results and the order of the anomaly impact, the weight distribution algorithm is used to prioritize the selection of phase log data repair paths and generate a priority sequence for the execution of the repair paths, specifically:

[0076] Obtain the log data anomaly recognition results corresponding to all cable numbers, and uniformly convert each type of anomaly recognition result into a feature offset ratio expression.

[0077] The impact weight coefficients of short-term fluctuations and long-term drift anomalies are preset (for example, the impact weight of short-term fluctuations is 0.7, and that of long-term drifts is 0.3). The abnormal characteristic offset ratios of short-term fluctuations and long-term drifts are weighted and summed, and the comprehensive anomaly index is calculated. The task sequence number of the cable phase log data repair task is adjusted according to the reverse order of the comprehensive anomaly index.

[0078] Calculate the single weighted offset of the short-term fluctuation and long-term drift anomaly characteristics of the phase log data corresponding to each task sequence number, perform secondary sorting in reverse order of the weighted offset size, and add the corresponding repair tasks to the log data repair queue in sequence according to the final sorting results.

[0079] Use distributed locks (such as those based on Zookeeper or Redis) to synchronize the execution of repair tasks, ensuring that only one repair task is executed at a time. After each task is completed, the distributed lock is released, allowing the next task to begin, preventing data repair thread confusion.

[0080] In S3, a first-level repair strategy is established based on the short-term fluctuation anomaly characteristics. The interpolation algorithm is used to generate a repair value sequence to reconstruct the short-term fluctuation anomaly data of the phase log. Specifically, the strategy includes:

[0081] In a distributed environment, the executor first controls access to the repair tasks in the log data repair queue by acquiring a distributed lock. Tasks in the repair queue are sorted by priority, and the highest-priority repair task is extracted. Based on the cable number and timestamp corresponding to the task, the executor queries the data storage system for the relevant data segments within the log monitoring sliding window. These data segments typically consist of historical phase data, including corresponding time series data records.

[0082] When the executor of the first-level repair strategy obtains the distributed lock, the sliding log data segment in the log monitoring sliding window corresponding to the first-order repair task in the log data repair queue is obtained.

[0083] The phase fluctuation amplitudes from the static features of the classification benchmark are extracted and used as the upper and lower limits of short-term fluctuation anomalies. By comparing the sliding log data segments item by item, data points that exceed the upper and lower limits are identified, marked as short-term fluctuation anomaly data points, and their locations are recorded.

[0084] The interpolation algorithm is selected as the repair basis. According to the time index of the short-term fluctuation abnormal data point and its adjacent non-abnormal data points before and after, a local interpolation with matching input dimension and abnormal fluctuation range interval is constructed to generate a corresponding repair value sequence. An interpolation formula is constructed based on the adjacent data points before and after the abnormal data point:

[0085]

[0086] In the formula, RE is the repair value, z0 and z1 are the values of the adjacent non-abnormal data points before and after, t0 and t1 are the time indexes corresponding to the adjacent non-abnormal data points before and after, t is the time index of the current abnormal data point, and ∈ is a repair tolerance parameter used to prevent abnormal loss caused by excessive repair.

[0087] The generated repair value is used to replace the short-term fluctuation abnormal data point in sequence, and the reconstructed data is continuously inserted by using time sequence.

[0088] At the edge of the repaired abnormal data point, the transition interval value between the repaired value and the adjacent non-abnormal data point is recalculated by weighted moving average to generate a smoothed local data sequence update log, and the distributed lock is released.

[0089] In S4, a second-level repair strategy is established based on the characteristics of long-term offset anomaly, and a trend repair data is generated based on polynomial fitting and segmented regression to reconstruct the long-term offset anomaly data of the phase log. Specifically, it includes:

[0090] When the execution program of the second-level repair strategy obtains the distributed lock, the sliding log data segment in the log monitoring sliding window corresponding to the first-order repair task in the log data repair queue is obtained.

[0091] For each period of data obtained, the execution program analyzes the log data by calculating the matching degree of the dynamic deviation trend of the data point and the sliding mean trend of the reference dynamic feature. The dynamic deviation trend refers to the trend of the data point relative to its previous and subsequent time points within a certain period of time. The reference dynamic feature is generated from the dynamic feature of the historical phase log established previously, representing the trend change of the data under normal circumstances. When the data point deviates greatly from the reference dynamic feature and continues, it is identified as an abnormal data point.

[0092] By calculating the deviation degree of each data point, the system can effectively identify the abnormal interval with continuous deviation. For example, if multiple consecutive data points deviate from the reference feature beyond the set threshold range, the interval is marked as an abnormal interval, and the corresponding interval identifier is generated. The interval identifier is recorded and passed to the subsequent processing module to ensure that the subsequent repair algorithm can identify and focus on the abnormal interval.

[0093] The trend is described by the changes in the key features of the abnormal pattern in the abnormal interval, specifically:

[0094] For each data segment in the abnormal interval, we use polynomial fitting to capture the nonlinear trend of the data in the interval. Assume that the observation value of the data interval is {(x i ,y i )}, where x i is the timestamp, y i is the corresponding phase data, and the polynomial function k(x) is expressed as:

[0095] k(x)=a n x n +a n-1 x n-1 +…+a1x 1 +a0;

[0096] Where n is the order of the polynomial, a0, a1, ..., a n-1 、a n is the coefficient of fitting, x represents {(x i ,y i )} in the timestamp variable.

[0097] An iterative optimization algorithm is used to gradually adjust the parameter set of the fitting model, with the goal of minimizing the relative residual of the fitting relative to the classification benchmark features of the long-term offset anomaly (specifically set based on the fitting accuracy requirements, the default setting is 15% to prevent the problem of anomaly loss caused by overfitting). The consistency trend expression of the benchmark features and the key features of the anomaly pattern is constructed. The deviation between the data point and the fitting curve is calculated as follows:

[0098]

[0099] Where m is the number of data points, P is the sum of squared errors, and the best fitting coefficient is obtained by minimizing P within the allowable range of relative residuals.

[0100] Based on the tuned trend fitting model, regression calculation is performed on each time point in the abnormal interval to generate a repair value sequence covering all time points in the abnormal interval. The abnormal data is replaced with the repair value sequence data in sequence, and the distributed lock is released after the replacement is completed.

[0101] In S5, the data after the repair is completed is re-identified for short-term fluctuations and long-term drift anomalies to determine whether the phase log data has been repaired. If so, the log data repair is terminated. If not, a phase anomaly alarm signal is issued. Specifically, the following steps are performed:

[0102] When all tasks in the log data repair queue are completed, the system marks the task status as "Completed". The repaired data is compared again with the classification benchmark to generate the corresponding short-term fluctuation and long-term drift anomaly identification results.

[0103] If both short-term fluctuations and long-term drift anomalies are identified as "non-abnormal", data repair is terminated. If short-term fluctuations or long-term drift anomalies are identified as "abnormal", an alarm signal is generated. The alarm signal includes information such as the abnormal time period, fluctuation amplitude, drift direction and cable number, and is sent to the maintenance personnel's terminal in real time. The system sends the alarm signal to the maintenance personnel's terminal device through a specified communication protocol (such as MQTT, HTTP, etc.). The signal not only includes the time period and cable number of the abnormal data, but also provides the specific abnormality type (short-term fluctuations or long-term drift) and the degree of abnormal deviation for maintenance personnel to analyze and use.

[0104] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.

[0105] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0106] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0107] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0108] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0109] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0110] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0111] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0112] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0113] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A log data processing method based on cable phase measurement, characterized in that: The steps include: S1: Integrate the multidimensional features of the cable historical phase log data to construct a feature vector set containing multidimensional descriptions, and establish a classification benchmark for the short-term and long-term characteristic anomaly patterns of each cable number; S2: Set a log monitoring sliding window to monitor the latest log records of each cable in real time and generate anomaly identification results corresponding to short-term fluctuations and long-term drifts respectively; S3: Based on the short-term fluctuation anomaly characteristics, a first-level repair strategy is established, and an interpolation algorithm is used to generate a repair value sequence to reconstruct the short-term fluctuation anomaly data of the phase log; S4: A second-level repair strategy is established based on the long-term offset anomaly characteristics. The trend repair data is generated based on polynomial fitting and piecewise regression to reconstruct the long-term offset anomaly data of the phase log. S5: Re-identify the abnormalities of short-term fluctuations and long-term drifts on the data after the repair is completed, and determine whether the phase log data has been repaired. If so, end the log data repair; if not, issue a phase abnormality alarm signal.

2. The method for processing log data based on cable phase measurement according to claim 1, characterized in that: In S1, the multidimensional features of the cable historical phase log data are integrated to construct a feature vector set containing multidimensional descriptions. The classification benchmarks for the short-term and long-term characteristic anomaly patterns of each cable number are established. Specifically, the following are included: Screen the records in the cable historical phase log data, delete the logs with missing items, incomplete formats and duplicate content, and obtain the original record set with complete data; Read the original record set one by one, and create time series phase history log data groups corresponding to different cable numbers based on the data upload timestamp and cable number; The time series phase historical log data corresponding to each cable number is grouped and analyzed independently. The static features of the data are extracted from the historical logs, including the mean, extreme value, standard deviation and fluctuation amplitude of the phase data, and a static feature description set is constructed. Through time series analysis, the dynamic feature values ​​of historical logs are extracted, including the rate of change of phase data and the sliding window mean of a set window size, to construct a dynamic feature description set. Through frequency domain analysis, the frequency domain feature values ​​of historical logs are extracted, including high-frequency amplitude and high-frequency energy ratio, and a frequency domain feature description set is constructed. Integrate static features, dynamic features and frequency domain features to establish a feature vector set containing multi-dimensional descriptions; The extracted log feature vector set corresponding to the normal operation period of the cable history is divided into short-term and long-term feature categories through feature clustering, and each category corresponds to a comprehensive feature group; Based on the comprehensive feature grouping of different feature categories corresponding to the phase history log of each cable number, a classification benchmark for the short-term and long-term feature anomaly patterns of the phase history log of each cable number is established. The matrix expression corresponding to the classification benchmark is: Among them, D is the matrix corresponding to the classification benchmark, DS and DL are the classification benchmarks of short-term and long-term feature abnormal patterns respectively, and X max 、X min 、X rate 、X avg are the maximum, minimum, rate of change and mean of the phase data respectively, σ is the standard deviation of the phase data, F high 、E high They are the high-frequency amplitude and high-frequency energy ratio of the phase data frequency domain, W avg is the sliding window mean.

3. The method for processing log data based on cable phase measurement according to claim 2, characterized in that: In S2, a log monitoring sliding window is set to monitor the latest log records of each cable in real time, and generate abnormal identification results corresponding to short-term fluctuations and long-term drifts, including: Use the cable number as the search identifier to set the cable log monitoring sliding window to obtain the latest log records of each cable in real time; For the log data in each monitoring sliding window, a feature vector matrix based on the feature anomaly classification benchmark is constructed. The data in the sliding window is mapped to the feature space one by one, and the key features representing different anomaly patterns are extracted. The multi-center iterative distribution method is used to cluster the key features representing different anomaly patterns. The Euclidean distance to the classification benchmark is calculated within the feature vector matrix. A distance offset anomaly threshold is set. By comparing the thresholds, identification results corresponding to short-term fluctuations and long-term drift anomalies are generated. A repair task is established for each log data entry identified as an anomaly. According to the characteristic deviation of the anomaly identification results and the order of the anomaly impact, the weight distribution algorithm is used to prioritize the selection of phase log data repair paths and generate a priority sequence for the execution of the repair paths, specifically: Obtain the log data anomaly recognition results corresponding to all cable numbers, and uniformly convert each type of anomaly recognition result into a feature offset ratio expression; Preset the impact weight coefficients of short-term fluctuations and long-term drift anomalies, perform weighted summation on the abnormal characteristic offset ratios of short-term fluctuations and long-term drifts, calculate the comprehensive anomaly index, and adjust the task sequence number of the cable phase log data repair task according to the reverse order of the comprehensive anomaly index; Calculate the single weighted offset of the short-term fluctuation and long-term drift anomaly characteristics of the phase log data corresponding to each task sequence number, perform secondary sorting in reverse order of the weighted offset size, and add the corresponding repair tasks to the log data repair queue according to the final sorting results. The distributed lock is used to control the synchronous execution of the repair tasks.

4. The method for processing log data based on cable phase measurement according to claim 3, characterized in that: In S3, a first-level repair strategy is established based on the short-term fluctuation anomaly characteristics. The interpolation algorithm is used to generate a repair value sequence to reconstruct the short-term fluctuation anomaly data of the phase log. Specifically, the strategy includes: When the executor of the first-level repair strategy obtains the distributed lock, it obtains the sliding log data segment in the log monitoring sliding window corresponding to the first-priority repair task in the log data repair queue; The phase fluctuation amplitude in the static features of the classification benchmark is extracted as the upper and lower limits of the abnormal fluctuation range. The data exceeding the upper and lower limits in the sliding log data segment are defined as short-term fluctuation abnormal data points, and their positions are marked. An interpolation algorithm is selected as the repair basis. Based on the time index of the short-term fluctuation abnormal data point and its adjacent non-abnormal data points, a local interpolation matching the input dimension with the abnormal fluctuation range is constructed to generate the corresponding repair value sequence. The generated repair values ​​are used to replace the short-term fluctuation abnormal data points in sequence, and the reconstructed data are inserted continuously step by step in chronological order; At the edge of the repaired abnormal data point, the transition interval value between the repaired value and the adjacent non-abnormal data point is recalculated through weighted moving average to generate a smoothed local data sequence update log and release the distributed lock.

5. The method for processing log data based on cable phase measurement according to claim 4, characterized in that: In S4, a second-level repair strategy is established based on the long-term offset anomaly characteristics. The trend repair data is generated based on polynomial fitting and piecewise regression to reconstruct the long-term offset anomaly data of the phase log. Specifically, the strategy includes: When the executor of the second-level repair strategy obtains the distributed lock, it obtains the sliding log data segment in the log monitoring sliding window corresponding to the first-priority repair task in the log data repair queue; Scan the time series data of the sliding log data segment one by one, and identify abnormal intervals of continuous deviation by calculating the degree of match between the dynamic deviation trend of the data point and the sliding mean change trend in the benchmark dynamic feature, and generate corresponding interval identifiers; Based on polynomial fitting and segmented regression methods, a mathematical model is constructed to describe the phase change trend of the abnormal interval, and the trend is described by the changes in the key characteristics of the abnormal pattern of the abnormal interval; An iterative optimization algorithm is used to gradually adjust the parameter set of the fitting model, with the goal of minimizing the relative residual of the fitting relative to the classification benchmark characteristics of the long-term offset anomaly, and constructing a consistent trend expression between the benchmark characteristics and the key characteristics of the anomaly pattern; Based on the tuned trend fitting model, regression calculation is performed on each time point in the abnormal interval to generate a repair value sequence covering all time points in the abnormal interval. The abnormal data is replaced with the repair value sequence data in sequence, and the distributed lock is released after the replacement is completed.

6. The method for processing log data based on cable phase measurement according to claim 5, characterized in that: In S5, the data after the repair is completed is re-identified for short-term fluctuations and long-term drift anomalies to determine whether the phase log data has been repaired. If so, the log data repair is terminated. If not, a phase anomaly alarm signal is issued. Specifically, the following steps are performed: When all tasks in the log data repair queue are completed, the repaired data is compared with the classification benchmark again to generate the corresponding short-term fluctuation and long-term drift anomaly identification results; When the short-term fluctuation and long-term drift anomaly identification results are detected as non-abnormal, the data repair is terminated. When the short-term fluctuation or long-term drift anomaly identification results are detected as abnormal, a phase abnormality data alarm signal is sent to the maintenance personnel terminal. The alarm signal contains the specific time period and cable number corresponding to the phase log data where the abnormality occurred.

Citation Information

Patent Citations

  • Log collection management method and system

    CN118152355A

  • Heating abnormity reason analysis method and heating system

    CN118391728A