Substation multi-source data processing system based on artificial intelligence

By designing a multi-source data processing system and using machine learning models to dynamically adjust the constraint solution, the problem of traditional systems being difficult to process multi-source data is solved, efficient data integration and processing is achieved, and the operation and management level of the substation is improved.

CN120045551APending Publication Date: 2025-05-27国网重庆市电力公司市南供电分公司 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411871217.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional substation data processing systems are difficult to cope with the needs of multi-source data processing, and there are problems such as data silos, insufficient processing capabilities, poor real-time performance, data accuracy and security.

Method used

Design a multi-source data processing system, through the collaborative work of multiple data source providers, data storage systems and data processing systems, and adopts machine learning models to dynamically adjust and optimize constraint solutions to achieve standardized data processing, quality detection and security guarantee.

Benefits of technology

It realizes effective integration and processing of multi-source data, improves the flexibility and accuracy of data processing, ensures data consistency and comparability, and improves the operation and management level of the substation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045551A_ABST
    Figure CN120045551A_ABST
Patent Text Reader

Abstract

The invention relates to a transformer substation multi-source data processing system based on artificial intelligence. The transformer substation multi-source data processing system comprises a plurality of data source providers, a data storage system and a data processing system. A data source provider stores collected data in a data storage system, and a data processing system obtains data from the data storage system and generates a main data set. For each data source provider, the data processing system formulates and dynamically optimizes a constraint scheme, which is applied to the primary data set to generate a screened data set. The screening data set comprises a plurality of screened node setting files. A data processing system identifies a key data set stored by a key data source provider and generates a screened data set for data of other data source providers. The method is suitable for the fields of transformer substations, intelligent power grids, industrial automation, intelligent transportation and the like, and the flexibility, accuracy and decision support capability of data processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-source data processing, and in particular to a substation multi-source data processing system based on artificial intelligence. Background Art

[0002] With the development of power systems and the growth of power demand, the role of substations in power transmission and distribution has become increasingly important. As a key node in the power system, the operation status of substations directly affects the stability and reliability of the entire power system. In order to ensure the safe and efficient operation of substations, it is necessary to monitor, collect and process various operating data of substations in real time. These data include but are not limited to voltage, current, power, frequency, temperature, equipment status, etc. Traditional substation data processing systems mainly rely on a single data source and use relatively simple processing methods, which is difficult to cope with the needs of multi-source data processing in modern power systems.

[0003] Traditional substation data processing systems usually rely only on data provided by equipment in the substation, such as transformers, circuit breakers, relay protection devices, etc. The data of these devices are usually transmitted to the control center through fieldbus or industrial communication network for centralized processing. This single data source processing method has the following limitations: Due to the single data source, it is impossible to fully reflect the operating status of the substation. For example, it is difficult to fully understand the operating status of other equipment by relying only on transformer data. The data of each device is independent of each other, forming a data island, which makes it difficult to achieve comprehensive integration and comprehensive analysis of data. Traditional data processing systems usually use simple algorithms for data processing, which is difficult to cope with the processing needs of large-scale, multi-source data. With the increase in the number and type of substation equipment, the amount of data is growing exponentially, and the processing capacity of traditional systems is gradually unable to meet actual needs. In the process of data transmission and processing, traditional systems have certain delays, making it difficult to achieve real-time data processing. At the same time, due to the lack of effective data verification and quality detection mechanisms, data loss and false alarms are prone to occur, affecting the accuracy and reliability of data. Traditional systems mainly rely on manual analysis and processing, and lack intelligent data analysis and processing capabilities. It is impossible to use modern artificial intelligence and machine learning technologies to conduct in-depth data mining and analysis, making it difficult to extract valuable information and patterns from massive amounts of data.

[0004] In order to overcome the limitations of traditional systems and achieve comprehensive monitoring and efficient processing of substation operation data, it is necessary to introduce multi-source data integration technology. Multi-source data integration refers to the integration of multiple types of data from different sources to achieve unified management, processing and analysis of data, thereby improving the efficiency and effectiveness of data processing. Multi-source data integration technology has the following advantages in substation data processing: By integrating multiple data sources from inside and outside the substation, the operating status of the substation can be fully reflected. For example, the data of equipment such as transformers, circuit breakers, relay protection devices, and environmental sensors can be integrated to fully understand the operation of the substation.

[0005] Although multi-source data integration technology has broad application prospects in substation data processing, it still faces many challenges and problems in practical applications: multi-source data comes from different sources, and the data formats and standards are also different. How to achieve standardized processing and compatibility of data from different sources is one of the main challenges facing multi-source data integration. In the process of multi-source data integration, data quality and credibility are important factors that must be considered. How to ensure the accuracy, integrity and consistency of data and prevent data loss, false alarms and tampering are problems that multi-source data integration technology needs to solve. In the process of multi-source data integration, a large amount of sensitive data and privacy information is involved. How to ensure the privacy and security of data and prevent data leakage and illegal access are important challenges facing multi-source data integration technology. Summary of the invention

[0006] In order to solve the above problems in the prior art, the present invention proposes a multi-source data processing system, including: multiple data source providers 1, a data storage system 2 and a data processing system 3; multiple data source providers 1 store the collected data in the data storage system 2; the data processing system 3 obtains relevant data from the data storage system 2 and generates a main data set 4; the data processing system 3 formulates a corresponding constraint scheme 5 for each data source provider, and the data processing system 3 dynamically adjusts and optimizes the constraint scheme 5 through a machine learning model 6; the data processing system 3 applies the constraint scheme 5 to the main data set 4 to generate a filtered data set 8; the filtered data set 8 includes a data set of multiple filtered node setting archives; the data processing system 3 identifies the key data set 9 in the main data set 4 stored in the data storage system 2 by the key data source provider based on the filtered data set 8; for multiple other data source providers, the data processing system 3 applies the constraint scheme 5 to other data sets in the main data set 4 to generate a filtered data set 10.

[0007] The constraint quality index of constraint scheme 5 is:

[0008] Q i=f(D i ,F i ,T i )

[0009] Among them, Q i represents the constraint quality index, D i Indicates data format requirements, F i represents the data quality standard, T i Indicates the data update frequency.

[0010] The data processing system 3 further comprises a data quality detection module 7, which is used to detect and mark abnormal data in the data storage system 2, and to regularly verify the data stored in the data storage system 2 by setting data quality standards.

[0011] The data format requirements of the constraint scheme 5 include data type definition, data unit standardization and timestamp format; the data quality standards include data integrity, data accuracy and data consistency; the data update frequency is set according to the real-time requirements of the data type, including high-frequency data and low-frequency data.

[0012] The data processing system 3 applies the constraint scheme 5 to the main data set 4 to generate a screening data set 8, including:

[0013] a) performing format verification on the data obtained from the data storage system, wherein the format verification includes data type verification, unit conversion and timestamp format verification;

[0014] b) Apply data quality standards, check the completeness, accuracy and consistency of data, and mark and remove data that does not meet the quality standards;

[0015] c) Adjust the data update frequency according to the real-time requirements of the data type, where high-frequency data is updated by seconds and low-frequency data is updated by hours or days;

[0016] d) The screening data set 8 includes a data set of multiple screened node setting files, each node setting file contains a set of data records of key parameters.

[0017] The data processing system 3 dynamically adjusts and optimizes the constraint scheme 5 through the machine learning model 6, including the following steps:

[0018] a) Collect historical data and real-time data for preprocessing, including data cleaning, normalization and feature extraction;

[0019] b) using supervised learning and unsupervised learning methods to train the machine learning model 6, wherein the supervised learning is used to predict the quality and trend of the data, and the unsupervised learning is used to detect anomalies in the data;

[0020] c) Validate and evaluate the model;

[0021] d) Based on the trained model, optimize and update the current constraint scheme 5, including data format optimization, data quality standard adjustment and data update frequency adjustment;

[0022] e) When data quality fluctuations are detected, the constraint scheme of the data source is automatically adjusted.

[0023] Identifying the key data set 9 according to the screening data set includes the following steps:

[0024] a) extracting data features from the screening data set, wherein the data features include statistical features and time features;

[0025] b) Use machine learning algorithms to evaluate the impact of each data feature on substation operation;

[0026] c) identifying key data based on the importance scores of the data features;

[0027] d) Verify and confirm key data through data accuracy verification and consistency checks.

[0028] The data processing system 3 applies the constraint scheme 5 to other data sets in the main data set 4 to generate a filtered data set 10, including the following steps:

[0029] a) Identify and categorize other data sources, including environmental monitoring sensors, equipment maintenance record systems, and external electricity market data;

[0030] b) Apply the data format requirements, data quality standards and data update frequency in the constraint scheme to other data sets to perform format standardization, quality inspection and frequency control;

[0031] c) According to the preset screening rules, filter out the data that does not meet the quality standards and format requirements, and generate a screened data set.

[0032] The steps of data detection by the data quality detection module 7 are as follows:

[0033] a) Set data quality standards, including data completeness, accuracy, consistency and timeliness;

[0034] b) Regularly perform quality checks on data stored in the data storage system, including checking for null or missing values, comparing current values ​​with historical data and expected values, and evaluating data consistency and update frequency;

[0035] c) When abnormal data is detected, an alarm is automatically generated and corresponding processing measures are taken, and the corresponding processing measures include requesting re-collection of data, adjusting the constraint scheme of the data source, and recording the abnormal data.

[0036] The multi-source data processing system is applied to data processing of substations or smart grids, industrial automation, and smart transportation.

[0037] The multi-source data processing system of the present invention dynamically adjusts and optimizes the constraint scheme through a machine learning model, thereby improving the flexibility and accuracy of data processing. The system can effectively integrate multiple data sources, generate a screened data set, and thus identify and extract key data. Through unified constraint scheme processing, the consistency and comparability of data are guaranteed, and it is suitable for fields such as substations, smart grids, industrial automation, and smart transportation, and significantly improves data processing efficiency and decision support capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present application, but do not constitute an improper limitation of the present invention. In the drawings:

[0039] Figure 1 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0040] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments, wherein the illustrative embodiments and descriptions are only used to explain the present invention but are not intended to limit the present invention.

[0041] The present invention relates to a multi-source data processing system for substation operation. The specific implementation method includes the collaborative work of multiple data source providers, data storage systems and data processing systems. The operation management of the substation is optimized by uniformly managing, processing and analyzing data from different sources. The specific implementation method of the present invention is described in detail below in conjunction with the accompanying drawings.

[0042] Example 1: System structure

[0043] like Figure 1 As shown, a substation operation multi-source data processing system includes multiple data source providers 1, a data storage system 2 and a data processing system 3. The multiple data source providers 1 store the collected data in the data storage system 2. The data processing system 3 obtains relevant data from the data storage system 2 and generates a main data set 4.

[0044] Data source providers 1 include various devices and sensors inside and outside the substation, which can collect substation operation data such as voltage, current, power, frequency, temperature, humidity, equipment, maintenance, etc. Each data source provider 1 has data collection and transmission functions, and transmits data to the data storage system 2 through the field bus, industrial communication network or wireless communication network.

[0045] The data storage system 2 is used to store data from multiple data source providers 1. The data storage system 2 adopts a distributed storage architecture to ensure the reliability and scalability of data. Specifically, the data storage system 2 includes multiple storage nodes, each of which is responsible for storing data of a specific type or source. The data storage system 2 adopts redundant storage and backup mechanisms to ensure the security and integrity of data.

[0046] The data processing system 3 is the core component of the present invention, and is responsible for processing and analyzing the data stored in the data storage system 2. The specific implementation content of the data processing system 3 is as follows:

[0047] The data processing system 3 obtains relevant data from the data storage system 2 through a preset interface. The data acquisition process includes data reading, preprocessing and cleaning to ensure the validity and accuracy of the data.

[0048] The data processing system 3 generates a main data set 4 from the acquired data according to preset rules and algorithms. The main data set 4 includes substation operation data such as voltage, current, power, frequency, temperature, humidity, equipment, and maintenance.

[0049] The data processing system 3 formulates a corresponding constraint scheme 5 for each data source provider. The constraint scheme 5 includes data format requirements, data quality standards, data update frequency, etc. Specifically, the constraint scheme can be expressed by the following formula:

[0050] Q i =f(D i ,F i ,T i )

[0051] Among them, Q i represents the constraint quality index, D i Indicates data format requirements, F i represents the data quality standard, T i Indicates the data update frequency; the specific form of f() may be a weighted sum function, a linear relationship, a nonlinear relationship, or other complex combinations, depending on the actual needs of the system and the model design.

[0052] The data processing system 3 applies the constraint scheme 5 to the main data set 4 to generate a filtered data set 8 .

[0053] The screening data set 8 includes a plurality of screened node setting archive data sets. The data processing system 3 applies the constraint scheme 5 to the main data set 4 to filter out data that does not meet the requirements and generate a data set that meets the standards.

[0054] The data processing system 3 identifies the key data set 9 in the main data set 4 stored in the data storage system 2 by the key data source provider based on the screened data set 8. The key data set 9 refers to data that has a significant impact on the operation of the substation, and these data are strictly screened and verified to ensure their accuracy and reliability.

[0055] For multiple other data source providers, the data processing system 3 applies the constraint scheme 5 to other data sets of the main data set 4 to generate a filtered data set 10. These data sets include environmental monitoring data, equipment maintenance records, etc., which are processed through a unified constraint scheme to ensure data consistency and comparability.

[0056] The data processing system 3 includes a data quality detection module 7, which is used to detect and mark abnormal data in the data storage system 2 to ensure the accuracy and integrity of the data. The data quality detection module 7 regularly verifies the data stored in the data storage system 2 by setting data quality standards. When abnormal data is detected, the system automatically generates an alarm and takes corresponding processing measures, such as re-collecting data, adjusting data sources, etc.

[0057] Example 2-1: Development and application of constraint scheme

[0058] The data processing system 3 formulates corresponding constraint schemes 5 for each data source provider. These constraint schemes include data format requirements, data quality standards, data update frequency, etc.

[0059] The data processing system 3 first standardizes the data format of each data source provider. The data formats of different sources may be different, and a unified data format is helpful for subsequent data processing and analysis. The data format requirements specifically include:

[0060] Data type definition: clearly define the type of each data field, such as integer, floating point, string, etc.

[0061] Data unit standardization: Convert data in different units into standard units. For example, standardize voltage data into volts (V) and current data into amperes (A).

[0062] Timestamp format: All data must be accompanied by a timestamp. The timestamp format uses the international standard ISO8601, such as YYYY-MM-DDTHH:MM:SSZ.

[0063] Through the above standardization process, the consistency of data format is ensured, which facilitates subsequent analysis and processing.

[0064] Data quality standards are used to ensure the accuracy and reliability of data. The data processing system 3 formulates corresponding quality standards for different data sources, including but not limited to the following aspects:

[0065] Data integrity: Ensure that all required fields in each data record have values.

[0066] Data accuracy: By comparing historical data with expected values, the rationality of the data can be tested. Statistical methods such as mean and standard deviation can be used to determine whether the data is within a reasonable range.

[0067] Data consistency: Ensure the consistency of data from the same source over different time periods. For example, the voltage data of the same device should not fluctuate dramatically in a short period of time.

[0068] The data quality standard can be expressed by the following formula:

[0069]

[0070] Among them, Q i represents the quality score of the ith data source, D ij represents the jth data from the i-th data source, E ij represents the expected value of the jth data from the i-th data source, w j It represents the weight of the j-th data, and n is the number of data.

[0071] The data update frequency refers to the frequency at which the data source provider transmits data to the data storage system 2. Different types of data have different requirements for real-time performance, and the data processing system 3 determines the data update frequency based on actual needs. For example:

[0072] High-frequency data: Key operating parameters such as voltage and current require high-frequency updates, usually once per second.

[0073] Low-frequency data: such as equipment maintenance records, ambient temperature, etc., can be set to update once every hour or every day.

[0074] The data processing system 3 applies the formulated constraint scheme 5 to the main data set 4 to generate a filtered data set 8 .

[0075] Perform format verification on the data obtained from the data storage system 2 to ensure that the data meets the predefined format requirements. The format verification includes data type verification, unit conversion and timestamp format verification.

[0076] Apply data quality standards to check the data for completeness, accuracy, and consistency. If the data does not meet the quality standards, the system will mark it as abnormal data.

[0077] The data is sampled and filtered according to the preset update frequency. High-frequency data is updated by seconds, and low-frequency data is updated by hours or days to ensure the timeliness and rationality of data updates.

[0078] Through the above processing, a screening data set 8 is generated. The screening data set 8 includes a data set of multiple screened node setting files. The specific steps of generating the screening data set 8 are as follows:

[0079] 1. Node setting file definition: A set of key parameters is defined for each substation node, including but not limited to voltage, current, power, temperature, etc. Each node setting file contains data records of these key parameters.

[0080] 2. Application of data screening rules: Screen the data of each node setting file according to the data quality standard in the constraint scheme 5. Only data that meets the quality standard will be retained in the screened data set 8.

[0081] 3. Abnormal data processing: For data marked as abnormal, the data processing system 3 removes it from the screening data set 8 and records the abnormal situation to facilitate subsequent analysis and processing.

[0082] Through the above process, the screening data set 8 generated by the data processing system 3 includes a plurality of screened node setting archive data sets. These data sets are strictly screened and verified to ensure their accuracy and reliability, and can fully reflect the operation status of the substation.

[0083] Example 2-2: Optimization and update mechanism of constraint scheme 5

[0084] The data processing system 3 uses historical data and real-time data to train the machine learning model 6 to identify patterns and anomalies in the data, thereby dynamically adjusting and optimizing the constraint scheme 5. The specific implementation steps are as follows:

[0085] 1. Data collection and preprocessing:

[0086] The data processing system 3 collects historical data and real-time data from the data storage system 2. These data include various key parameters, such as voltage, current, power, temperature, etc. After collecting the data, the system preprocesses the data, including data cleaning, normalization, and feature extraction. Data cleaning is used to remove noise data and outliers, normalization is used to unify the dimensions of the data, and feature extraction is used to extract important features in the data.

[0087] 2. Machine Learning Model 6 Training:

[0088] The system uses a combination of supervised learning and unsupervised learning to train machine learning models6. Supervised learning models are used to predict the quality and trend of data, and unsupervised learning models are used to detect anomalies in data. Supervised learning algorithms include linear regression, support vector machines, decision trees, and neural networks. Unsupervised learning algorithms include K-means clustering, principal component analysis (PCA), etc. The specific training process is as follows:

[0089] -Supervised Learning:

[0090] The training set consists of historical data and labels (such as data quality scores). The model predicts the quality score of new data by learning the relationship between data and labels.

[0091] y=f(X)+∈

[0092] Among them, y is the predicted quality score, X is the input data feature, and ∈ is the error term.

[0093] - Unsupervised Learning:

[0094] Unsupervised learning models detect anomalies in data by analyzing its distribution and structure. For example, the K-means clustering algorithm divides data into multiple clusters and detects outliers as abnormal data.

[0095]

[0096] min: indicates that the goal of this formula is to minimize the sum of squared errors in the entire clustering process, that is, the sum of the squares of the Euclidean distances from all data points to the centroid of their respective clusters.

[0097] k: represents the number of clusters, that is, the data is divided into k different classes.

[0098] C i : represents the i-th cluster, C i Contains all sample data points belonging to this cluster.

[0099] x j : represents the jth data point, x j ∈C i Represents data point x j belongs to the i-th cluster.

[0100] μ i : represents the centroid (center point) of the ith cluster. The centroid is the point obtained by taking the average value of all data points in the cluster in each dimension.

[0101] ‖x j -μ i ‖ 2 : represents the data point x jTo cluster center μ i The square of the Euclidean distance.

[0102] 3. Model verification and evaluation:

[0103] After training is completed, the system verifies and evaluates the model. The validation set is used to evaluate the accuracy and robustness of the model. The evaluation indicators of this application include mean square error (MSE). The model evaluation formula is as follows:

[0104]

[0105] Among them, y i is the true value, is the predicted value, and n′ is the number of samples.

[0106] 4. Optimization and update of constraint scheme:

[0107] Based on the trained model, the data processing system 3 optimizes and updates the current constraint solution 5. The optimization process includes the following aspects:

[0108] A. Data format optimization

[0109] The data processing system 3 dynamically adjusts the data format requirements by analyzing the feedback results of the model. The specific steps are as follows:

[0110] The system performs statistical analysis on data formats from different data sources and evaluates their impact on data quality.

[0111] By comparing the data quality scores of different formats, the optimal data format can be identified.

[0112]

[0113] Among them, Q i is the quality score of the i-th data format, Q ij is the quality score of the jth data in the i-th data format, and n is the number of data.

[0114] Based on the data format analysis results, the system selects the format with the highest data quality score as the standard format. If the data quality of a format is significantly higher than that of other formats, the system automatically applies that format to the corresponding data source.

[0115] For data that does not conform to the standard format, the system automatically converts the format. The format conversion includes unit conversion, data type conversion, timestamp format conversion, etc., to ensure that all data is unified in the standard format.

[0116] For example, to convert temperature from Fahrenheit to Celsius:

[0117]

[0118] Among them, T C is the temperature in degrees Celsius, T F The temperature is in degrees Fahrenheit.

[0119] B. Adjustment of data quality standards

[0120] The data processing system 3 dynamically adjusts the data quality standard according to the quality score predicted by the model. The specific steps are as follows:

[0121] The system predicts the quality score of each data source through the trained supervised learning model. The model input includes the feature values ​​of historical data, and the output is the quality score.

[0122]

[0123] in, is the quality score of the prediction, X is the input data feature matrix, and f is the trained prediction model.

[0124] The system adjusts the quality standards for each data source based on the predicted quality score. For data sources with large fluctuations in quality scores, the system increases the strictness of their quality checks. For example, the threshold for completeness checks is increased from 80% to 90%.

[0125] Quality standard adjustment formula:

[0126] Q ′ =Q×(1+α×σ)

[0127] Among them, Q ′ is the adjusted quality standard, Q is the current quality standard, α is the adjustment coefficient, and σ is the standard deviation of the quality score.

[0128] The system applies the adjusted quality standards to the newly collected data and checks its completeness, accuracy and consistency. For data that does not meet the new standards, the system marks it as abnormal data and records the relevant information.

[0129] C. Data update frequency adjustment

[0130] The data processing system 3 adjusts the data update frequency according to the real-time requirements of the data and the prediction results of the model. The specific steps are as follows:

[0131] The system analyzes the real-time requirements of different types of data and determines the update frequency of each type of data. For example, key operating parameters such as voltage and current require high-frequency updates, while ambient temperature and humidity can be updated at low frequencies.

[0132] Based on the quality score of model prediction and real-time requirements, the system dynamically adjusts the data update frequency. For example, for key parameter data, the system will increase the update frequency from once a minute to once a second.

[0133] Update frequency adjustment formula:

[0134]

[0135] Where f′ is the adjusted update frequency, f is the current update frequency, β is the adjustment coefficient, Q″ is the current quality score, Q min and Q max The minimum and maximum values ​​for the quality score.

[0136] The system controls the frequency of data collection and transmission according to the adjusted update frequency to ensure the balance between data real-time and quality. For example, for voltage data, the system increases the data collection frequency to ensure that it is collected once every second.

[0137] Through the above steps, the data processing system 3 realizes the dynamic optimization and update of the constraint scheme 5. Through data format optimization, data quality standard adjustment and data update frequency adjustment, the system can effectively improve the accuracy, reliability and real-time performance of the data, and provide data support for the efficient operation of the substation.

[0138] 5. Abnormal situation handling:

[0139] When the data quality of a data source provider fluctuates, the system will automatically adjust the constraint scheme of the data source to improve the accuracy and reliability of the data. The specific steps are as follows:

[0140] The system detects anomalies in the data through unsupervised learning models and records the source, type, and time of the abnormal data. The system adjusts the corresponding data quality standards and data update frequency according to the type and frequency of abnormal data. For example, for data sources with frequent anomalies, the system will improve its data quality standards and increase the frequency of data collection. The system applies the adjusted constraint scheme to the data processing process and continuously monitors its effect. Through the feedback mechanism, the constraint scheme is continuously optimized to ensure the high quality and reliability of the data.

[0141] In summary, the present invention automatically optimizes and updates the constraint scheme 5 of the data processing system 3 through the machine learning model 6, thereby realizing dynamic management and efficient processing of multi-source data.

[0142] Example 3: Identification of key data sets

[0143] The data processing system 3 identifies the key data set 9 in the main data set 4 stored in the data storage system 2 by the key data source provider based on the screened data set 8. The key data set 9 refers to the data that has a significant impact on the operation of the substation, and these data are strictly screened and verified to ensure their accuracy and reliability. The following is a detailed description of the process.

[0144] The data processing system 3 first obtains the screening data set 8 from the data storage system 2. The screening data set 8 is generated by applying the constraint scheme 5 to the main data set 4, and these data sets include the screened node setting files. Each node setting file contains key parameters of substation operation, such as voltage, current, temperature, etc. Through steps such as data format standardization, quality detection and frequency control, data that does not meet the requirements has been excluded, ensuring the basic quality of the data.

[0145] Steps to identify key data sets:

[0146] 1. Data feature extraction:

[0147] The data processing system 3 extracts data features of each key parameter from the screening data set 8. These features include statistical features (such as mean, variance, maximum value, minimum value) and time features (such as change rate, periodicity).

[0148] 2. Materiality Assessment:

[0149] The data processing system 3 uses machine learning algorithms (such as random forests and support vector machines) to evaluate the degree of influence of each feature on the operation of the substation. During the training process of the machine learning algorithm, the system associates the data features with the substation operation status (such as equipment failure and performance degradation) and identifies the features that have a significant impact on the operation of the substation.

[0150] 3. Key data identification:

[0151] According to the importance scores of the features, the data processing system 3 identifies the key data set 9. The key data set 9 includes data features that have a significant impact on the operation of the substation. The system sets a threshold based on the scores of these features and selects data with scores higher than the threshold as key data.

[0152] 4. Verification and confirmation:

[0153] The data processing system 3 verifies and confirms the identified key data set 9. The verification process includes data accuracy verification and consistency check. The system verifies the accuracy of the identified key data by comparing historical data with actual operation conditions.

[0154] 5. Dynamic adjustment and optimization:

[0155] The identification of key data sets 9 is a dynamic process. The data processing system 3 continuously adjusts and optimizes the key data sets by continuously monitoring the operating status and data characteristics of the substation. The system regularly updates the machine learning model 6, re-evaluates the importance of data features, and ensures that the key data sets always reflect the latest status of substation operation.

[0156] In practical applications, the key data set 9 provides reliable data support for the operation monitoring and management of the substation. By identifying and monitoring key data, the system can detect potential faults and abnormal conditions in a timely manner, take preventive measures in advance, and improve the operational reliability and safety of the substation. For example, when the voltage and current data in the key data set fluctuate abnormally, the system can quickly issue an alarm and instruct relevant maintenance personnel to check and handle it.

[0157] In summary, the present invention generates a key data set 9 by extracting features, evaluating importance, identifying key data, and verifying the screening data set 8. These key data are strictly screened and verified to ensure their accuracy and reliability, providing data support for the efficient operation of the substation.

[0158] Example 4: Processing of other data sources

[0159] For multiple other data source providers, the data processing system 3 applies the constraint scheme 5 to other data sets in the main data set 4 to generate a filtered data set 10. These data sets include environmental monitoring data, equipment maintenance records, etc. Through unified constraint scheme processing, the consistency and comparability of the data are ensured.

[0160] The data processing system 3 first identifies and classifies other data sources. These data sources include but are not limited to environmental monitoring sensors, equipment maintenance record systems, external power market data, etc. The data types provided by each data source are different, and the data processing system needs to classify and label these data.

[0161] According to the nature and purpose of the data, the system divides the data into categories such as environmental monitoring data, equipment maintenance data, and market data. Environmental monitoring data includes parameters such as temperature, humidity, and air pressure; equipment maintenance data includes equipment maintenance records, fault history, etc.; market data includes electricity prices, supply and demand information, etc.

[0162] In terms of the application of the constraint scheme, the data processing system 3 applies the data format requirements in the constraint scheme 5 to other data sets. For different categories of data, the system adopts a unified format standardization method. For example, all timestamps are converted to the ISO8601 standard format, and data in different units are converted to a unified standard unit.

[0163] The system performs quality checks on other data sets according to the data quality standards in constraint scheme 5. Data quality checks include integrity checks, accuracy verification, and consistency checks. For environmental monitoring data, the system checks the continuity and rationality of the data; for equipment maintenance data, the system verifies the integrity and accuracy of data records.

[0164] The data processing system 3 controls the frequency of other data sets according to the data update frequency requirements in the constraint scheme 5. The system adjusts the frequency of data collection and update according to the real-time requirements of the data. For example, for environmental monitoring data, the system may be set to update once an hour; for equipment maintenance data, the system may be set to update once a day.

[0165] The data processing system 3 applies the data screening rules in the constraint scheme 5 to other data sets. The system filters out the data that does not meet the quality standards and format requirements according to the preset screening rules, and generates a screened data set 10.

[0166] The data processing system 3 ensures the consistency and comparability of data from all data sources through a unified constraint solution. The system compares and analyzes data from different data sources to ensure consistency in format, quality and update frequency.

[0167] The filtered data set 10 can be used for multi-faceted analysis and management of substation operation. Environmental monitoring data can be used to evaluate the impact of the external environment on substation equipment, equipment maintenance data can be used to formulate preventive maintenance plans, and market data can be used to optimize power dispatch and trading decisions. Through unified processing and analysis of these data, substations can achieve more efficient and reliable operation management.

[0168] In practical applications, the data processing system 3 significantly improves the management level of the substation by processing multiple other data sources. For example, by analyzing environmental monitoring data, the system can warn in advance of the impact of severe weather on the operation of the substation; by analyzing equipment maintenance records, the system can predict the failure trend of equipment and formulate a reasonable maintenance plan to avoid losses caused by sudden equipment failure; by analyzing market data, the system can optimize the purchase and sale strategy of electricity and improve economic benefits.

[0169] In summary, the present invention ensures the consistency and comparability of data by unified processing of multiple other data sources, thereby providing reliable data support for efficient management of substations.

[0170] Example 5: Data quality detection module

[0171] The data processing system 3 includes a data quality detection module 7, which is used to detect and mark abnormal data in the data storage system 2 to ensure the accuracy and integrity of the data. The data quality detection module 7 regularly verifies the data stored in the data storage system 2 by setting data quality standards. The following is a detailed description of the specific implementation steps.

[0172] The data quality detection module 7 first sets a set of data quality standards, which include the completeness, accuracy, consistency and timeliness of the data. Specifically:

[0173] Make sure that all required fields in each data record have values, and there are no null or missing values.

[0174] By comparing historical data with expected values, the rationality and accuracy of the data can be tested. For example, voltage data should fluctuate within a reasonable range and no abnormal values ​​beyond the expected range should appear.

[0175] Ensure data consistency from the same data source at different time points. For example, the current data of the same device should not fluctuate dramatically in a short period of time.

[0176] Ensure that data collection and transmission meet the preset update frequency and reflect the operating status of the substation in a timely manner.

[0177] The data quality detection module 7 regularly performs quality verification on the data stored in the data storage system 2. The verification process includes the following steps:

[0178] The system traverses each data record to check whether there are null or missing values. If missing values ​​are found, the system will mark the data as abnormal data.

[0179] The system compares the current value of each data record with historical data and expected values ​​to evaluate its rationality. For example, statistical methods are used to calculate the mean and standard deviation of the data to detect whether the data fluctuates within a reasonable range.

[0180]

[0181] Where x is the current data value, μ is the data mean, and σ is the standard deviation. If the deviation exceeds the preset threshold, the system will mark the data as abnormal data.

[0182] The system compares data from the same data source at different points in time to assess their consistency. For example, time series analysis methods are used to detect the continuity and rate of change of data.

[0183]

[0184] Among them, x t and x t-1are the data values ​​at the current and previous time points, respectively, and Δt is the time interval. If the rate of change exceeds the preset threshold, the system will mark the data as abnormal data.

[0185] The system checks the timestamp of the data to ensure that the data collection and transmission meet the preset update frequency. If the data is not updated on time, the system will mark the data as abnormal data.

[0186] When the data quality detection module 7 detects abnormal data, the system will automatically generate an alarm and take corresponding treatment measures. These measures include: the system generates an abnormal data alarm and notifies the relevant operators. The alarm information includes detailed information such as data source, abnormal type, detection time, etc.

[0187] For data marked as abnormal, the system will automatically request the data source provider to re-collect the data to obtain the latest data that meets quality standards.

[0188] If a data source frequently has quality problems, the system will automatically adjust the constraint scheme 5 of the data source to improve the data quality standard, or switch to an alternative data source.

[0189] The system records all detected abnormal data and establishes abnormal data archives. These records can be used for subsequent analysis and improvement, helping the system to continuously optimize data quality detection standards and methods.

[0190] In practical applications, the data quality detection module 7 ensures the accuracy and integrity of the substation operation data. For example, by timely discovering and processing abnormal values ​​in environmental monitoring data, the system can avoid the potential impact of environmental factors on equipment operation. By checking the quality of equipment maintenance records, the system can ensure the accuracy of maintenance information and avoid maintenance plan errors caused by inaccurate data.

[0191] In summary, the present invention ensures the accuracy and integrity of data through the implementation of the data quality detection module 7. By setting data quality standards, regularly verifying the stored data, and taking timely processing measures for abnormal data, the system can provide reliable data support and ensure the efficient and safe operation of the substation.

[0192] Example 6: Practical application scenario

[0193] In practical applications, the operation of a multi-source data processing system in a substation can significantly improve the operation efficiency and management level of the substation. For example, in a substation, the system monitors the operation status of the substation in real time by integrating data from different devices. When a device fails, the system can quickly identify the cause of the failure and provide a corresponding solution. In addition, by analyzing historical data, the system can predict the maintenance needs of the equipment, perform preventive maintenance in advance, and avoid sudden failures.

[0194] The present invention is not only applicable to data processing in substations, but can also be applied to other fields that require multi-source data processing. For example, the multi-source data processing technology of the present invention can be used in the fields of smart grid, industrial automation, smart transportation, etc. to achieve efficient data integration and intelligent analysis.

[0195] With the development of technology, the system structure and processing method of the present invention can be continuously optimized and upgraded. For example, more advanced machine learning algorithms and big data processing technologies can be introduced to improve the speed and accuracy of data processing. The comprehensive monitoring capability of the system can also be further improved by increasing the types and number of data sources.

[0196] In summary, the present invention provides an efficient and reliable substation operation multi-source data processing system, which integrates multiple data sources and adopts machine learning algorithms and data quality detection technology to achieve comprehensive monitoring and intelligent analysis of substation operation data, thereby significantly improving the operation efficiency and management level of the substation.

[0197] The above description is only a preferred embodiment of the present invention, so all equivalent changes or modifications made according to the structure, characteristics and principles described in the scope of the patent application of the present invention are included in the scope of the patent application of the present invention.

Claims

1. A multi-source data processing system, comprising: Multiple data source providers (1), data storage systems (2) and data processing systems (3); A plurality of data source providers (1) store the collected data in the data storage system (2); the data processing system (3) obtains relevant data from the data storage system (2) and generates a main data set (4); characterized in that: The data processing system (3) formulates a corresponding constraint scheme (5) for each data source provider, and the data processing system (3) dynamically adjusts and optimizes the constraint scheme (5) through a machine learning model (6); the data processing system (3) applies the constraint scheme (5) to the main data set (4) to generate a filtered data set (8); the filtered data set (8) includes a plurality of filtered node setting archive data sets; the data processing system (3) identifies a key data set (9) in the main data set (4) stored in the data storage system (2) by a key data source provider based on the filtered data set (8); for a plurality of other data source providers, the data processing system (3) applies the constraint scheme (5) to other data sets of the main data set (4) to generate a filtered data set (10).

2. A multi-source data processing system as claimed in claim 1, characterized in that: The constraint quality index of the constraint scheme (5) is: Q i =f(D i ,F i ,T i ) Among them, Q i represents the constraint quality index, D i Indicates data format requirements, F i represents the data quality standard, T i Indicates the data update frequency.

3. A multi-source data processing system as claimed in claim 1, characterized in that: The data processing system (3) further comprises a data quality detection module (7), wherein the data quality detection module (7) is used to detect and mark abnormal data in the data storage system (2), and to regularly verify the data stored in the data storage system (2) by setting data quality standards.

4. A multi-source data processing system as claimed in claim 2, characterized in that: The data format requirements of the constraint scheme (5) include data type definition, data unit standardization and timestamp format; the data quality standards include data integrity, data accuracy and data consistency; the data update frequency is set according to the real-time requirements of the data type, including high-frequency data and low-frequency data.

5. A multi-source data processing system as claimed in claim 1, characterized in that: The data processing system (3) applies the constraint scheme (5) to the main data set (4) to generate a screening data set (8), comprising: a) performing format verification on the data obtained from the data storage system, wherein the format verification includes data type verification, unit conversion and timestamp format verification; b) Apply data quality standards, check the completeness, accuracy and consistency of data, and mark and remove data that does not meet the quality standards; c) Adjust the data update frequency according to the real-time requirements of the data type, where high-frequency data is updated by seconds and low-frequency data is updated by hours or days; d) The screening data set (8) includes a data set of multiple screened node setting files, each node setting file contains a set of data records of key parameters.

6. A multi-source data processing system as claimed in claim 1, characterized in that: The data processing system (3) dynamically adjusts and optimizes the constraint scheme (5) through a machine learning model (6), comprising the following steps: a) Collect historical data and real-time data for preprocessing, including data cleaning, normalization and feature extraction; b) training the machine learning model (6) using supervised learning and unsupervised learning methods, wherein the supervised learning is used to predict the quality and trend of the data, and the unsupervised learning is used to detect anomalies in the data; c) Validate and evaluate the model; d) Based on the trained model, optimize and update the current constraint scheme (5), including data format optimization, data quality standard adjustment, and data update frequency adjustment; e) When data quality fluctuations are detected, the constraint scheme of the data source is automatically adjusted.

7. A multi-source data processing system as claimed in claim 1, characterized in that: Identifying the key data set based on the screening data set (9) includes the following steps: a) extracting data features from the screening data set, wherein the data features include statistical features and time features; b) Use machine learning algorithms to evaluate the impact of each data feature on substation operation; c) identifying key data based on the importance scores of the data features; d) Verify and confirm key data through data accuracy verification and consistency checks.

8. A multi-source data processing system as claimed in claim 1, characterized in that: The data processing system (3) applies the constraint scheme (5) to other data sets of the main data set (4) to generate a filtered data set (10) comprising the following steps: a) Identify and categorize other data sources, including environmental monitoring sensors, equipment maintenance record systems, and external electricity market data; b) Apply the data format requirements, data quality standards and data update frequency in the constraint scheme to other data sets to perform format standardization, quality inspection and frequency control; c) According to the preset screening rules, filter out the data that does not meet the quality standards and format requirements, and generate a screened data set.

9. A multi-source data processing system as claimed in claim 3, characterized in that: The steps of data quality detection module (7) for data detection are as follows: a) Set data quality standards, including data completeness, accuracy, consistency and timeliness; b) Regularly perform quality checks on data stored in the data storage system, including checking for null or missing values, comparing current values ​​with historical data and expected values, and evaluating data consistency and update frequency; c) When abnormal data is detected, an alarm is automatically generated and corresponding processing measures are taken, and the corresponding processing measures include requesting re-collection of data, adjusting the constraint scheme of the data source, and recording the abnormal data.

10. A multi-source data processing system according to any one of claims 1 to 9, characterized in that: The multi-source data processing system is applied to data processing of substations or smart grids, industrial automation, and smart transportation.

Citation Information

Cited By

  • Frequency control data transmission system and method for extra-high voltage direct current sending end

    CN122204900A