A Method for Checking Abnormal Charging and Discharging Data Using Data Lineage Topological Relationships
By constructing a blood-related topological relationship and anomaly detection model for charging and swapping data, the real-time and accuracy of abnormal detection in the charging and swapping network are solved, efficient and accurate abnormal data processing is achieved, and system stability and operational efficiency are improved.
Patent Information
- Application Number
- CN202510406410.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art has insufficient real-time and accuracy of abnormal detection in the charging and swapping network, and cannot fully cover the business logic and data flow rules of charging and swapping data, resulting in insufficient sensitivity and timeliness of abnormal detection.
Combining the blood relationship and topological structure of the data, a charging and swapping abnormal data inspection method is constructed. Through data collection, blood relationship topological map construction, abnormal detection model group construction, abnormal positioning and traceability and processing, it integrates depth-first search, random forest algorithm, support vector machine and time series analysis to realize multi-dimensional abnormal detection of charging and swapping data.
It improves the efficiency and accuracy of abnormal detection of charging and swapping data, enhances the stability and reliability of the system, and can promptly detect and process abnormal data, reduce business interruptions, and improves operational efficiency and user experience.
Smart Images

Figure CN119917984B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for checking abnormal charging and swapping data by using data lineage topological relationships. Background Art
[0002] With the rapid development of electric vehicles and charging infrastructure, the charging and swapping field has faced comprehensive challenges in the big data era. The amount of data processed by enterprises and charging station operators has increased sharply, including not only basic information such as user charging behavior, power consumption, and equipment status, but also complex associated data such as geographical location, time pattern, and equipment interoperability. Facing this massive and highly interconnected data set, how to efficiently manage, effectively utilize, and timely identify potential abnormal situations, such as abnormal charging behavior and equipment fault warning, has become the focus of the industry.
[0003] Regarding the characteristics of charging and swapping data, currently, a data anomaly checking strategy based on data lineage topological relationships can be adopted to optimize the data processing and monitoring process. Specifically, this strategy includes:
[0004] Data analysis method based on the charging and swapping data lineage relationship: By pre - constructing a lineage relationship graph between charging and swapping data, that is, clarifying the source, flow, and mutual dependence relationship of each piece of data such as charging records and equipment status changes. When data analysis or anomaly detection is required, the system can quickly locate and extract the data set closely associated with the target data field from the data warehouse or real - time data stream. Based on these comprehensive and accurate data associations, in - depth analysis can be carried out, which can significantly improve the accuracy and response speed of anomaly detection. For example, quickly identifying a high failure rate or abnormal charging mode of charging stations in a certain area during a certain period.
[0005] Abnormal charging and swapping detection method integrating topological structure and attribute information: Given that the charging and swapping network has natural spatial distribution characteristics and complex interaction relationships between devices, a method of network topological analysis can be used, combined with the attribute information of charging and swapping facilities (such as power, voltage, temperature, etc.), to construct a charging and swapping attribute network. Using algorithms such as graph theory and machine learning, abnormal nodes or abnormal patterns can be detected in the network, such as abnormal charging time distribution, abnormally high power loss, or abnormal frequent switching of equipment. This method can not only effectively identify problems of a single device, but also reveal potential risks in the entire charging and swapping network, providing strong support for operation optimization and fault prevention.
[0006] In the field of charging and swapping data, although the data analysis method based on data lineage has shown significant advantages in improving data processing and analysis efficiency, when directly applied to data anomaly detection, especially when dealing with the large and complex datasets generated by the charging and swapping network, some challenges may be faced. These challenges are mainly reflected in the real-time and accuracy of anomaly detection. Since data lineage mainly focuses on the source and flow of data, and is not necessarily optimized directly for the identification of anomaly patterns. This means that in quickly identifying sudden failures, abnormal charging behaviors, or equipment performance degradation during the charging and swapping process, additional mechanisms may be required to enhance the sensitivity and timeliness of anomaly detection.
[0007] On the other hand, the anomaly detection method based on the deep combination of topological structure and attribute information does have advantages in capturing the complexity of the charging and swapping network and the characteristics of equipment attributes. However, this method may not fully consider the uniqueness of the data lineage in charging and swapping data, such as the transfer rules of data among charging stations, users, and vehicles, as well as the business logic and rules behind these data transfers. Therefore, directly applying this method to the anomaly detection of charging and swapping data warehouses or business data may overlook some anomaly patterns caused by specific business scenarios or data lineage, thus affecting the comprehensiveness and accuracy of anomaly detection.
[0008] To overcome these challenges, a more ideal solution is to combine the data lineage analysis with the anomaly detection method based on topological structure and attribute information to form a comprehensive charging and swapping data anomaly detection framework. This framework should be able to make full use of data lineage to trace the source and propagation path of abnormal data, and at the same time, identify and verify potential anomaly patterns through in-depth analysis of topological structure and attribute information. This can not only ensure the real-time and accuracy of anomaly detection, but also comprehensively cover various anomalies in the charging and swapping network, providing more accurate and efficient decision-making support for operators. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a method for checking abnormal charging and swapping data using data lineage topology. By integrating the visual presentation of the data lineage graph, the construction of fine data topology relationships, and the efficient data anomaly detection and instant feedback mechanism, it effectively solves the problems of data development and management anomalies encountered in the production process of charging and swapping big data.
[0010] To solve the above technical problem, the technical solution adopted by the present invention is: A method for checking abnormal charging and swapping data using data lineage topology, comprising the following steps:
[0011] S1. Data collection: Collect relevant data from various data sources in the charging and swapping system, organize them into a data set, and preprocess the data set to obtain a preprocessed multi-dimensional charging and swapping data set.
[0012] S2. Construction of the charging and swapping data lineage topology graph: Establish the data lineage topology graph among data items in the charging and swapping system, and classify and label the nodes in the data lineage topology graph; then, use the depth-first search algorithm to obtain key nodes and analyze the data lineage relationship; finally, after the analysis is completed, generate the charging and swapping data lineage topology graph.
[0013] S3. Construction of the anomaly detection model group.
[0014] S4. Anomaly location and tracing: Use the constructed data lineage topology graph to conduct tracing analysis on the detected anomalies, and clarify the origin of the abnormal data and its propagation path in the entire charging and swapping network.
[0015] S5. Handling of charging and swapping data anomalies: Develop a handling strategy according to the type, severity, and impact range of the anomalies.
[0016] Preferably, the data sources include the charging station control system, the user payment platform, and the vehicle information system.
[0017] Preferably, the step S1 specifically includes the following steps:
[0018] S11. Collect charging and swapping data from different time periods, different locations, and different platform sources, and organize them into a data set.
[0019] S12. Manually discriminate the dimensions in the data set, and perform dimensionality reduction, removal of duplicate values, and removal of outliers to obtain the charging and swapping data table DA, the recent fault location table DB, the area table DC, the protection unit table DD, the first olt data summary table 1DE, the second olt data summary table 2DF, the third olt data summary table 3DG, the fourth olt data summary table 4DH, and the data application table DI.
[0020] Preferably, the step S2 includes the following steps:
[0021] S21. Use natural language processing and regular expressions to accurately identify the charging and swapping data lineage node set VDA, the recent fault location lineage node set VDB, the regional lineage node set VDC, the protection unit lineage node set VDD, the first olt data summary lineage node set VDE, the second olt data summary lineage node set VDF, the third olt data summary lineage node set VDG, the fourth olt data summary lineage node set VDH, and the data application lineage node set VDI from SQL query statements, data flow diagrams, and ETL scripts. At the same time, identify the dependency relationships and transfer paths E between data items to provide basic information for subsequent construction of data lineage relationships;
[0022] S22. Take the charging and swapping data lineage node set VDA, the recent fault location lineage node set VDB, the regional lineage node set VDC, and the protection unit lineage node set VDD as the starting nodes of the data lineage topology graph, and the data application lineage node set VDI as the ending node. Take the sets VDE, VDF, VDG, and VDH between the starting node and the ending node as the backbone nodes of the data lineage topology graph to construct the skeleton of the data lineage topology graph;
[0023] S23. According to the topology graph skeleton, combined with the lineage relationships between the sets, automatically generate the topology graph G.
[0024] Preferably, the step S3 includes the following steps:
[0025] S31. Construct an anomaly classification model based on the data lineage topology graph, collect multi-dimensional data in the charging and swapping system, and mark key nodes through the construction of the data lineage relationship topology graph to form a data lineage network;
[0026] S32. Construct a user behavior anomaly detection model based on user behavior analysis, establish real-time and batch anomaly detection models. The batch detection model extracts features from the topology graph and determines node anomalies through the random forest algorithm. User behavior analysis uses the support vector machine SVM to identify anomalies in charging behavior;
[0027] S33. Construct a periodic anomaly detection model, and identify the abnormal trend of seasonal load changes through the periodic anomaly detection model STL decomposition and Prophet model;
[0028] S34. Construct a real-time anomaly detection model, and the real-time detection model analyzes the charging data stream in real time based on a sliding window to locate sudden anomalies in equipment downtime.
[0029] Preferably, the step S31 specifically includes: First, extract each node and path from the constructed data lineage topology graph, that is, the key features of device description, select, combine, and generate derivative features for the features, add specific tags, and construct a charging and swapping dataset containing normal samples and abnormal samples. Then, divide the cleaned charging and swapping dataset into a training set and a test set according to a certain ratio; import the training set into a random forest model for learning; before the random forest model performs anomaly detection on the features of each node and path in the topology structure, set the dimension number of the sample feature subset of each decision tree in the random forest model to , where p is the total number of features; during the model training process, through the cross-validation method, set the candidate parameter set param_grid, and use grid search to determine the optimal number of decision trees M; for each decision tree, calculate the importance of node features respectively, and calculate the anomaly probability according to the status of different paths; the output result of each decision tree is 0 or 1, representing normal and abnormal respectively; through the majority voting method, synthesize the output of each decision tree to obtain the final determination result of the node;
[0030] In terms of anomaly scoring and calculation, synthesize the prediction results of multiple decision trees, and define the topology anomaly score S topology :
[0031] ;
[0032] where M represents the number of trees in the random forest model, and T m (x) is the classification result of the m-th tree for the sample x;
[0033] Finally, set the anomaly threshold T topology ; when the topology anomaly score S topology >T topology , the random forest model will mark the corresponding node as abnormal and give the corresponding anomaly label.
[0034] Preferably, the step S32 specifically includes: First, collect user behavior data, including charging time, frequency, charging pile status, and user historical behavior. After cleaning and preprocessing the original user behavior data, extract a feature vector composed of multiple key features from the user behavior data to construct a key feature matrix , where n is the number of samples and d is the feature dimension;
[0035] Then, input the matrix features into a nonlinear support vector machine SVM for training; by introducing a Gaussian kernel:
[0036] ;
[0037] where, K() represents the Gaussian kernel function, where γ is a parameter controlling the kernel width. An optimal hyperplane is found to effectively distinguish normal charging behaviors from abnormal charging behaviors, and a classifier that can classify the input behavior data into "normal charging behavior" or "abnormal charging behavior" is obtained;
[0038] After training is completed, the trained SVM model is used to classify new user behaviors; for each input sample x, the model will output a label , where -1 represents abnormal charging behavior and +1 represents normal charging behavior; by analyzing the classification results, abnormal charging behaviors are identified;
[0039] When an abnormal charging behavior is detected, all the information of the current relevant nodes and edges of the abnormal charging behavior is captured through the data lineage topology graph, and the data is input into the data lineage topology graph to identify the specific cause of the abnormality.
[0040] Preferably, the step S33 specifically includes: First, preprocess the charging and swapping data by handling missing values, removing obvious outliers, standardizing, and removing long-term trends to ensure data stability and reduce the noise level; then use STL decomposition to decompose the time series, and the decomposition formula is:
[0041] ;
[0042] Among them, Y t represents the time series, represents the trend component, reflecting the long-term change trend of charging and swapping demand over time; represents the seasonal component, reflecting the fixed periodic changes in charging and swapping demand, is the residual part, representing other short-term fluctuations or noises that cannot be captured by the trend and seasonality;
[0043] After STL decomposition is completed, perform periodic modeling on the seasonal component; according to the seasonal characteristics of charging and swapping demand, the Prophet model first separates the trend term from the seasonal term and first fits the long-term trend change of the data; subsequently, the Prophet model generates sine and cosine basis functions through Fourier series to represent periodic changes, and uses the user-specified period to automatically adapt to the seasonal components of daily fluctuations, differences between weekends and weekdays, and generates a repeatable periodic function; for the special demand fluctuations brought by holidays, the Prophet model adds holidays as features through the holiday effect module and learns the unique impact of holidays on charging demand to adjust the expected value and capture the changes in holiday demand;
[0044] After the Prophet model is trained, it will generate future prediction data as a reference for periodic normal values; then calculate the standardized residuals:
[0045] ;
[0046] where Residual t is the difference between the actual observed value and the predicted value of the Prophet model, μ is the mean of the residuals, and σ is the standard deviation of the residuals; set a threshold , for each time point t, when , the Prophet model enters the anomaly detection stage; thus identifying the moments that exhibit anomalies based on the periodic trend and seasonal variations.
[0047] Preferably, the step S34 is specifically as follows: Use a real-time anomaly detection system to continuously monitor the charging and swapping data stream, and import the acquired data into an anomaly detection model based on the data lineage topology structure to obtain the current state of the device, ensuring rapid response, including sudden device shutdown and sudden anomalies in abnormal charging behavior.
[0048] Preferably, it further includes step S6 anomaly feedback, collecting feedback information during the anomaly detection process and transmitting it back.
[0049] The present invention provides a method for checking abnormal charging and swapping data using data lineage topology relationships, which has the following beneficial effects:
[0050] 1. Improve the efficiency of abnormal charging and swapping data detection:
[0051] The technical solution of the present invention realizes the accurate tracking of data flow during the charging and swapping process by constructing the lineage topology relationship of charging and swapping data. During anomaly detection, using this topology relationship can quickly locate potential problem areas, significantly reducing the time and cost of checking a large amount of charging and swapping data one by one. This not only speeds up the anomaly response speed but also improves the overall operation efficiency.
[0052] 2. Improve the accuracy of abnormal charging and swapping data detection:
[0053] The technical solution of the present invention innovatively combines the topology structure and specific attribute information of charging and swapping data (such as charging volume, charging duration, device status, etc.) for anomaly detection. This multi-dimensional and in-depth analysis method can more accurately identify abnormal values and abnormal patterns in the data. Whether it is the abnormal behavior of a single charging station or potential problems in the entire charging and swapping network, they can be promptly captured and accurately analyzed.
[0054] 3. Enhance the stability and reliability of the charging and swapping system:
[0055] The technical solution of the present invention enables abnormal data in the charging and swapping system to be detected in time and effectively processed. This avoids the spread and accumulation of abnormal data within the system, thereby reducing business interruptions or decision-making errors caused by data errors. At the same time, the rapid response and processing of abnormal data also improve the self-repair ability of the system, enhancing the stability and reliability of the system. In the context of the rapid development of the charging and swapping business, this advantage is of great significance for ensuring user experience and maintaining brand image. Brief Description of the Drawings
[0056] The present invention will be further described below in conjunction with the drawings and embodiments:
[0057] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0058] As Figure 1 shown, a method for checking abnormal charging and swapping data using data lineage topological relationships includes the following steps:
[0059] S1. Data collection: Collect relevant data from various data sources in the charging and swapping system, organize it into a data set, and preprocess the data set to obtain a preprocessed multi-dimensional charging and swapping data set;
[0060] S2. Construction of the charging and swapping lineage topological graph: Establish a lineage topological graph of the relationships between data items in the charging and swapping system, and classify and label the nodes in the data lineage topological graph; then, use the depth-first search algorithm to obtain key nodes and analyze the data lineage relationships; finally, after the analysis is completed, generate the charging and swapping lineage topological graph;
[0061] S3. Construction of an abnormal detection model group;
[0062] S4. Abnormal location and tracing: Use the constructed data lineage topological graph to perform tracing analysis on the detected abnormalities, and clarify the source of the abnormal data and its propagation path in the entire charging and swapping network;
[0063] S5. Processing of abnormal charging and swapping data: According to the type, severity, and scope of influence of the abnormality, formulate a processing strategy.
[0064] Preferably, the data sources include a charging station control system, a user payment platform, and a vehicle information system.
[0065] Preferably, step S1 specifically includes the following steps:
[0066] S11. Collect charging and swapping data from different time periods, different locations, and different platform sources, and organize it into a data set;
[0067] S12. Manually discriminate the dimensions in the dataset, and perform dimensionality reduction, duplicate value removal, and outlier removal operations to obtain the charging and swapping data table DA, the recent fault location table DB, the area table DC, the protection unit table DD, the first olt data summary table 1DE, the second olt data summary table 2DF, the third olt data summary table 3DG, the fourth olt data summary table 4DH, and the data application table DI.
[0068] Preferably, step S2 includes the following steps:
[0069] S21. Use natural language processing and regular expressions to accurately identify the charging and swapping data lineage node sets VDA = (VA1, VA2,..., VAa) from SQL query statements, data flow diagrams, and ETL scripts, where a represents the number of lineage nodes in this set; the recent fault location lineage node set VDB = (VB1, VB2,..., VBb), where b represents the number of lineage nodes in this set; the area lineage node set VDC = (VC1, VC2,..., VCc), where c represents the number of lineage nodes in this set; the protection unit lineage node set VDD = (VD1, VD2,..., VDd), where d represents the number of lineage nodes in this set; the first olt data summary lineage node set VDE = (VE1, VE2,..., VEe), where e represents the number of lineage nodes in this set; the second olt data summary lineage node set VDF = (VF1, VF2,..., VFf), where f represents the number of lineage nodes in this set; the third olt data summary lineage node set VDG = (VG1, VG2,..., VGg), where g represents the number of lineage nodes in this set; the fourth olt data summary lineage node set VDH = (VH1, VH2,..., VHh), where h represents the number of lineage nodes in this set; the data application lineage node set VDI = (VI1, VI2,..., VIi), where i represents the number of lineage nodes in this set; at the same time, identify the dependency relationships and transmission paths E = (E1, E2,..., Ej) between data items, where j represents the number of dependency relationships and transmission paths, providing basic information for subsequent data lineage relationship construction;
[0070] S22. Use the charging and swapping data lineage node set VDA, the recent fault location lineage node set VDB, the area lineage node set VDC, and the protection unit lineage node set VDD as the starting nodes of the data lineage topology graph, the data application lineage node set VDI as the ending node, and the sets VDE, VDF, VDG, and VDH between the starting node and the ending node as the backbone nodes of the data lineage topology graph to construct the data lineage topology graph skeleton;
[0071] S23. Automatically generate a topology graph G based on the topology graph skeleton and the blood relationship between sets.
[0072] Preferably, step S3 includes the following steps:
[0073] S31. Construct an anomaly classification model based on the data blood relationship topology graph, collect multi-dimensional data in the charging and swapping system, and through the construction of the data blood relationship topology graph, mark key nodes to form a data blood relationship network;
[0074] S32. Construct a user behavior anomaly detection model based on user behavior analysis, establish real-time and batch anomaly detection models. The batch detection model extracts features from the topology graph and determines node anomalies through the random forest algorithm. User behavior analysis uses the support vector machine SVM to identify abnormal charging behaviors;
[0075] S33. Construct a periodic anomaly detection model, and through the STL decomposition and Prophet model of the periodic anomaly detection model, identify the abnormal trend of seasonal load changes;
[0076] S34. Construct a real-time anomaly detection model, and the real-time detection model analyzes the charging data stream in real time based on a sliding window to locate sudden anomalies in equipment downtime.
[0077] STL decomposition (Seasonal-Trend decomposition using LOESS): STL decomposition is a time series decomposition technique that uses the locally weighted regression (LOESS) method to divide a time series into a trend component, a seasonal component, and a residual component. STL decomposition is particularly suitable for time series data with complex seasonal and trend changes.
[0078] Prophet model (Prophet Model): It is a statistical model for time series prediction, aiming to capture trends, seasonality, and holiday effects in time series. The model uses Fourier series to generate periodic changes and automatically adjusts the seasonal component according to the periods specified by users (such as years, months, weeks), and is suitable for handling holiday effects and regular seasonal fluctuations.
[0079] Anomaly detection model: It is a model used to identify and classify data that does not conform to the expected pattern. In the charging and swapping system, the anomaly detection model can identify possible equipment failures or abnormal charging behaviors by analyzing the characteristics and behaviors of data.
[0080] Seasonal Anomaly Detection: A method based on time series analysis for identifying abnormal data that deviates from seasonal trends and periodic changes. It models the seasonal components and trends of the data to identify abnormal patterns on top of normal periodic variations, and is particularly suitable for identifying issues such as abnormal loads caused by seasonal demand changes.
[0081] Preferably, step S31 specifically includes: First, extract each node and path from the constructed data lineage topology graph, that is, key features of device description, such as node attributes (abnormal fluctuations in charging amount, sharp increase in device failure rate, charging duration distribution, power consumption trend, etc.), edge attributes (data transfer time, data transfer success rate, etc.). Select, combine and generate derivative features for the features, add specific labels, and construct a charging and swapping dataset containing normal samples and abnormal samples. Then; divide the cleaned charging and swapping dataset into a training set and a test set according to a certain ratio; import the training set into the random forest model for learning; before the random forest model performs anomaly detection on the features of each node and path in the topology structure, set the number of dimensions of the sample feature subset of each decision tree in the random forest model to , where p is the total number of features; during model training, through cross-validation, set the candidate parameter set param_grid = {'n_estimators':[10, 50, 100, 200, 500]}, and use grid search to determine the optimal number of decision trees M = grid_search.best_params_['n_estimators']; for each decision tree, calculate the importance of node features respectively, and calculate the anomaly probability according to the status of different paths; the output result of each decision tree is 0 or 1, representing normal and abnormal respectively; through majority voting, synthesize the output of each decision tree to obtain the final determination result of the node;
[0082] In terms of anomaly scoring and calculation, comprehensively consider the prediction results of multiple decision trees and define the topology anomaly score S topology :
[0083] ;
[0084] where M represents the number of trees in the random forest model, and T m (x) is the classification result of the m-th tree for sample x;
[0085] Finally, set the anomaly threshold T topology =Percentile95(S topology , normal); when the topology anomaly score S topology >T topologyWhen it is the case, the random forest model marks the corresponding node as an anomaly and gives the corresponding anomaly label.
[0086] Preferably, the step S32 specifically includes: First, collect user behavior data, including charging time, frequency, charging pile status, and user historical behavior. After cleaning and preprocessing the original user behavior data, extract a feature vector composed of multiple key features from the user behavior data, and construct a key feature matrix , where n is the number of samples and d is the feature dimension;
[0087] Then, input the matrix features into a non-linear support vector machine SVM for training; by introducing a Gaussian kernel:
[0088] ;
[0089] Among them, K () represents the Gaussian kernel function, γ is a parameter that controls the kernel width, find an optimal hyperplane so that normal charging behaviors and abnormal charging behaviors can be effectively distinguished, and obtain a classifier that can classify the input behavior data into "normal charging behavior" or "abnormal charging behavior";
[0090] After the training is completed, use the trained SVM model to classify new user behaviors; for each input sample x, the model will output a label , where -1 represents abnormal charging behavior and +1 represents normal charging behavior; by analyzing the classification results, identify abnormal charging behaviors;
[0091] When an abnormal charging behavior is detected, capture the information of all the current relevant nodes and edges of the abnormal charging behavior through the data lineage topology graph, input the data into the data lineage topology graph, and identify the specific reasons for the anomaly.
[0092] Preferably, the step S33 specifically includes: First, preprocess the charging and swapping data by methods such as handling missing values, removing obvious outliers, standardizing, and removing long-term trends to ensure the data is stable and reduce the noise level; then use STL decomposition to decompose the time series, and the decomposition formula is:
[0093] ;
[0094] Among them, Y t represents the time series, represents the trend component, reflecting the long-term change trend of the charging and swapping demand over time; represents the seasonal component, reflecting the fixed periodic changes in the charging and swapping demand, This is the residual part, representing short-term fluctuations or noises that cannot be captured by trends and seasonality;
[0095] After the STL decomposition is completed, periodic modeling is performed on the seasonal components; according to the seasonal characteristics of the charging and discharging demand, the Prophet model first separates the trend term and the seasonal term, and first fits the long-term trend change of the data; subsequently, the Prophet model generates sine and cosine basis functions through Fourier series to represent periodic changes, and automatically adapts to the seasonal components of daily fluctuations, weekends, and weekdays using the user-specified period (such as year, month, week), generating a repeatable periodic function; for the special demand fluctuations brought by holidays, the Prophet model adds holidays as features through the Holiday Effect module and learns the unique impact of holidays on the charging demand to adjust the expected value and capture the changes in holiday demand;
[0096] After the Prophet model is trained, the Prophet model will generate future prediction data (synthetic values including trends and seasonal changes), which is used as a reference for periodic normal values; then calculate the standardized residuals:
[0097] ;
[0098] where Residual t is the difference between the actual observed value and the predicted value of the Prophet model, μ is the mean of the residuals, and σ is the standard deviation of the residuals; set the threshold , for each time point t, when the Prophet model enters the anomaly detection stage; thus identifying the moments that are abnormal based on the periodic trend and seasonal changes.
[0099] Preferably, the step S34 is specifically as follows: adopt a real-time anomaly detection system to continuously monitor the charging and discharging data stream, and import the obtained data into an anomaly detection model based on the data lineage topology structure to obtain the current state of the device, ensuring rapid response, including sudden device shutdown and sudden anomalies in abnormal charging behaviors. Timely feedback the results of anomaly detection and processing to the data management department and business operation team of the charging and discharging system to help optimize the data processing process, improve data quality, and service level. Continuously optimize the anomaly detection algorithm and strategy to adapt to the continuous changes in the development of the charging and discharging business.
[0100] Preferably, it further includes step S6 anomaly feedback, collecting and transmitting back the feedback information during the anomaly detection process.
[0101] The present invention provides a method for checking abnormal charging and swapping data by using the data lineage topology relationship, which relates to the technical field of data processing. By collecting multi-dimensional data from the charging and swapping system, constructing a data lineage topology graph, marking key nodes, and forming a data lineage network. Constructing a batch anomaly detection model based on a random forest, a user behavior anomaly detection model based on a support vector machine, and a periodic anomaly detection model based on STL decomposition and Prophet model to identify abnormal behaviors, equipment failures, and seasonal load changes in charging and swapping. At the same time, designing a real-time anomaly detection system to analyze the charging data stream in real time based on a sliding window and quickly respond to sudden anomalies such as equipment downtime. After detecting an anomaly, use the topology graph to trace the source, analyze the propagation path and impact of the abnormal data, and implement processing for different anomaly types, feedback and optimize the system in a timely manner to improve the stability and data processing efficiency of the charging and swapping system.
[0102] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.
Claims
1. A method for checking abnormal charging and discharging data using the data lineage topological relationship, characterized in that, It includes the following steps: S1. Data collection: Collect relevant data from various data sources in the charging and swapping system, organize it into a data set, and preprocess the data set to obtain a preprocessed multi-dimensional charging and swapping data set; S2. Construction of the charging and swapping data lineage topology graph: Establish the data lineage topology graph among various data items in the charging and swapping system, classify and label the nodes in the data lineage topology graph; then, use the depth-first search algorithm to obtain the key nodes and analyze the data lineage relationship; finally, after the analysis is completed, generate the charging and swapping data lineage topology graph; S3. Construction of the anomaly detection model group; S4. Anomaly location and tracing: Use the constructed data lineage topology graph to conduct a tracing analysis on the detected anomalies, and clarify the origin of the abnormal data and its propagation path in the entire charging and swapping network; S5. Handling of charging and swapping data anomalies: Develop a handling strategy according to the type, severity, and impact range of the anomalies; The step S3 includes the following steps: S31. Construct an anomaly classification model based on the data lineage topology graph, collect multi-dimensional data in the charging and swapping system, and through the construction of the data lineage topology graph, label the key nodes to form a data lineage network; S32. Construct a user behavior anomaly detection model based on user behavior analysis, establish real-time and batch anomaly detection models. The batch detection model extracts features from the topology graph and determines node anomalies through the random forest algorithm. User behavior analysis uses the support vector machine SVM to identify abnormal charging behaviors; S33. Construct a periodic anomaly detection model, and identify the abnormal trend of seasonal load changes through the periodic anomaly detection model STL decomposition and Prophet model; S34. Construct a real-time anomaly detection model, and the real-time detection model analyzes the charging data stream in real time based on a sliding window to locate the sudden anomalies of equipment downtime.
2. The method for checking abnormal charging and discharging data using the data lineage topology relationship according to claim 1, wherein The data sources include the charging station control system, the user payment platform, and the vehicle information system.
3. The method for checking abnormal charging and discharging data using the data lineage topology relationship according to claim 1 or 2, characterized in that, The step S1 specifically includes the following steps: S11. Collect charging and swapping data from different time periods, different locations, and different platform sources, and organize it into a data set; S12. Manually discriminate the dimensions in the data set, and perform processing such as dimensionality reduction, removal of duplicate values, and removal of outliers to obtain the charging and swapping data table DA, the recent fault location table DB, the area table DC, the protection unit table DD, the first olt data summary table 1DE, the second olt data summary table 2DF, the third olt data summary table 3DG, the fourth olt data summary table 4DH, and the data application table DI.
4. The method for checking abnormal charging and discharging data using the data lineage topology relationship according to claim 3, wherein The step S2 includes the following steps: S21. Using natural language processing and regular expressions, accurately identify the collection of charging and swapping data lineage nodes VDA, the collection of recent fault location lineage nodes VDB, the collection of regional lineage nodes VDC, the collection of protection unit lineage nodes VDD, the first olt data summary lineage node collection VDE, the second olt data summary lineage node collection VDF, the third olt data summary lineage node collection VDG, the fourth olt data summary lineage node collection VDH, and the data application lineage node collection VDI from SQL query statements, data flow diagrams, and ETL scripts. At the same time, identify the dependency relationships and transmission paths E between data items to provide basic information for subsequent data lineage relationship construction; S22. Take the collection of charging and swapping data lineage nodes VDA, the collection of recent fault location lineage nodes VDB, the collection of regional lineage nodes VDC, and the collection of protection unit lineage nodes VDD as the starting nodes of the data lineage topology graph, and the data application lineage node collection VDI as the ending node. Take the collections VDE, VDF, VDG, and VDH between the starting node and the ending node as the backbone nodes of the data lineage topology graph to construct the skeleton of the data lineage topology graph; S23. According to the topology graph skeleton, combined with the lineage relationships between the collections, automatically generate the topology graph G.
5. The method for checking abnormal charging and discharging data using the data lineage topological relationship according to claim 1, wherein The specific steps of step S31 are as follows: First, extract each node and path from the constructed data lineage topology graph, that is, the key features of device description, select, combine and generate derivative features for the features, add specific tags, and construct a charging and swapping dataset containing normal samples and abnormal samples. Then, divide the cleaned charging and swapping dataset into a training set and a test set according to a certain ratio; import the training set into the random forest model for learning; before the random forest model performs anomaly detection on the features of each node and path in the topology structure, set the dimension number of the sample feature subset of each decision tree in the random forest model to be , where p is the total number of features; during the model training process, set the candidate parameter set param_grid through the cross-validation method, and use grid search to determine the optimal number of decision trees M; for each decision tree, calculate the importance of node features respectively, and calculate the anomaly probability according to the states of different paths; the output result of each decision tree is 0 or 1, representing normal and abnormal respectively; through the majority voting method, comprehensively combine the outputs of each decision tree to obtain the final determination result of the node; In terms of anomaly scoring and calculation, the prediction results of multiple decision trees are integrated to define the topological anomaly score S topology : ; Among them, M represents the number of trees in the random forest model, and T m (x) is the classification result of the m-th tree for the sample x; Finally, set the anomaly threshold T topology ; when the topology anomaly score S topology >T topology , the random forest model will mark the corresponding node as an anomaly and give the corresponding anomaly label.
6. The method for checking abnormal charging and discharging data using the data lineage topology relationship according to claim 1, characterized in that, The specific steps of S32 include: First, collect user behavior data, including charging time, frequency, charging pile status, and user historical behavior. After cleaning and preprocessing the original user behavior data, extract a feature vector composed of multiple key features from the user behavior data, and construct a key feature matrix , where n is the number of samples and d is the feature dimension; Then, input the matrix features into a non-linear support vector machine SVM for training; by introducing a Gaussian kernel: ; Among them, K () represents the Gaussian kernel function, and γ is a parameter controlling the kernel width. Find an optimal hyperplane so that normal charging behaviors and abnormal charging behaviors can be effectively distinguished, and obtain a classifier that can classify the input behavior data into "normal charging behavior" or "abnormal charging behavior". After training is completed, the trained SVM model is used to classify new user behaviors; for each input sample x, the model will output a label , where -1 indicates abnormal charging behavior and +1 indicates normal charging behavior; by analyzing the classification results, abnormal charging behaviors are identified; When an abnormal charging behavior is detected, capture the information of all nodes and edges currently related to the abnormal charging behavior through the data lineage topology graph, input the data into the data lineage topology graph, and identify the specific cause of the abnormality.
7. The method for checking abnormal charging and discharging data using the data lineage topological relationship according to claim 1, characterized in that, The specific steps of S33 include: First, preprocess the charging and swapping data by handling missing values, removing obvious outliers, standardizing, and removing long-term trends to ensure data stability and reduce the noise level; then use STL decomposition to decompose the time series, and the decomposition formula is: ; Among them, Y t represents a time series, represents the trend component, reflecting the long-term change trend of charging and swapping demand over time; represents the seasonal component, reflecting the fixed periodic changes in charging and swapping demand, is the residual part, representing other short-term fluctuations or noises that cannot be captured by the trend and seasonality; After STL decomposition, perform periodic modeling on the seasonal components; according to the seasonal characteristics of the charging and swapping demand, the Prophet model first separates the trend term and the seasonal term, and first fits the long-term trend change of the data; subsequently, the Prophet model generates sine and cosine basis functions through Fourier series to represent periodic changes, and uses the user-specified period to automatically adapt to the seasonal components of daily fluctuations, differences between weekends and weekdays, and generates a repeatable periodic function; for the special demand fluctuations brought by holidays, the Prophet model adds holidays as features through the holiday effect module and learns the unique impact of holidays on charging demand to adjust the expected value and capture the changes in holiday demand; After the Prophet model training is completed, the Prophet model will generate future prediction data, which is used as a reference for periodic normal values; then calculate the standardized residuals: ; Among them, Residual t is the difference between the actual observed value and the predicted value of the Prophet model, μ is the mean of the residuals, and σ is the standard deviation of the residuals; a threshold is set. For each time point t, when occurs, the Prophet model enters the anomaly detection stage; thus, the moments that exhibit anomalies based on the periodic trend and seasonal changes are identified.
8. The method for checking abnormal charging and discharging data using the data lineage topology relationship according to claim 1, characterized in that The specific steps of step S34 are as follows: Use a real-time anomaly detection system to continuously monitor the charging and swapping data stream, and import the obtained data into an anomaly detection model based on the data lineage topology structure to obtain the current status of the device, ensuring rapid response, including sudden device downtime and sudden anomalies in abnormal charging behavior.
9. The method for checking abnormal charging and discharging data using the data lineage topological relationship according to claim 1, wherein It also includes step S6 anomaly feedback, collecting feedback information during the anomaly detection process and transmitting it back.
Citation Information
Patent Citations
Link reliability tracking and optimizing method combined with real-time monitoring
CN119325104A
Power enterprise data full-link fault positioning method based on data consanguinity
CN119449587A