Power test data quality evaluation method and system based on deep learning

Through deep learning, the completeness, logic and consistency analysis of power test data, and the classification and characteristics are independently analyzed, solving the problem of difficulty in discovering local anomalies and complex logical relationships in the existing technology, and achieving a more accurate and reliable data quality evaluation.

CN120429702APending Publication Date: 2025-08-05GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510384406.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing power test data quality evaluation methods are difficult to find local anomalies and complex logical relationships in the data, resulting in low accuracy and reliability of the evaluation results.

Method used

The power test data is analyzed integrity, logic and consistency by using deep learning methods, and classified into target power data categories. The data quality is evaluated through independent analysis of features and abnormal proportions, and a comprehensive evaluation is conducted based on the correlation coefficients between data features.

Benefits of technology

It improves the accuracy and reliability of power test data quality evaluation, can discover local anomalies and complex logical relationships, and provides more comprehensive data quality information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429702A_ABST
    Figure CN120429702A_ABST
Patent Text Reader

Abstract

The invention provides an electric power test data quality evaluation method and system based on deep learning, and the method comprises the steps: carrying out the integrity, logicality and consistency analysis of to-be-evaluated electric power test data based on deep learning, and obtaining abnormal electric power data and normal electric power data; classifying the normal power data based on the business characteristics to obtain a target power data class; performing feature mutual independent analysis based on the data features of each target power data class to obtain an initial correlation coefficient between the data features in each target power data class; based on the abnormal quantity of the abnormal power data and the total quantity of the power test data, determining a data abnormal proportion of the power test data; and performing quality evaluation on the power test data based on the data exception proportion and the initial correlation coefficient between the data features in each target power data class to obtain a quality evaluation result. According to the invention, the reliability and accuracy of the evaluation result of the power test data quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and system for evaluating power test data quality based on deep learning. Background Art

[0002] During power system operation, assessing the quality of power test data is crucial for ensuring safe and stable operation. Currently, existing methods for assessing power test data quality primarily rely on traditional statistical analysis and rule matching. Traditional statistical analysis methods assess data quality by calculating statistics such as the mean, variance, and standard deviation to assess the central tendency and dispersion of data. For example, statistics such as voltage and current are used to assess data stability and accuracy. Rule matching, on the other hand, uses pre-defined rules and thresholds to examine each data point individually to determine whether it complies with the specified rules.

[0003] However, traditional statistical analysis methods often only reflect the overall characteristics of the data, making it difficult to identify local anomalies and potential problems within the data. Furthermore, they are very strict about data distribution, significantly impacting the accuracy of evaluation results when data does not meet certain criteria. While rule matching can quickly identify obvious anomalies, it struggles to effectively detect complex logical relationships and potential errors. Because rules are pre-set, they cannot adapt to dynamic data changes and complex business scenarios, resulting in low reliability in power test data quality assessments. Summary of the Invention

[0004] The present invention provides a power test data quality assessment method and system based on deep learning, which are used to improve the reliability and accuracy of the assessment results of the power test data quality.

[0005] In a first aspect, the present invention provides a method for evaluating power test data quality based on deep learning, comprising:

[0006] Based on deep learning, the integrity, logic, and consistency of the power test data to be evaluated are analyzed to obtain abnormal power data and normal power data;

[0007] Classifying the normal power data based on business characteristics to obtain a target power data class;

[0008] Based on the data features of each target power data class, independent feature analysis is performed to obtain the initial correlation coefficient between the data features in each target power data class;

[0009] determining a data anomaly ratio of the power test data based on the abnormal number of the abnormal power data and the total number of the power test data;

[0010] The quality of the power test data is evaluated based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class, and the quality assessment results are obtained.

[0011] In a second aspect, the present invention further provides a power test data quality assessment system based on deep learning, which is applied to the power test data quality assessment method based on deep learning as described in the first aspect; the power test data quality assessment system based on deep learning includes:

[0012] A deep learning analysis module is used to analyze the integrity, logic, and consistency of the power test data to be evaluated based on deep learning, and obtain abnormal power data and normal power data;

[0013] A data classification module, configured to classify the normal power data based on business characteristics to obtain a target power data class;

[0014] A feature independence analysis module is used to perform feature independence analysis based on the data features of each target power data class to obtain the initial correlation coefficient between the data features in each target power data class;

[0015] an abnormality analysis module, configured to determine a data abnormality ratio of the power test data based on the abnormal number of the abnormal power data and the total number of the power test data;

[0016] The data quality assessment module is used to perform quality assessment on the power test data based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class to obtain a quality assessment result.

[0017] In a third aspect, the present invention also provides an electronic device, comprising: a memory for storing a computer software program; a processor for reading and executing the computer software program, thereby implementing any of the deep learning-based power test data quality assessment methods described above.

[0018] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, in which a computer software program is stored. When the computer software program is executed by a processor, it implements any of the deep learning-based power test data quality assessment methods described above.

[0019] In a fifth aspect, the present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the deep learning-based power test data quality assessment methods described above.

[0020] The power test data quality assessment method based on deep learning provided by the embodiment of the present invention performs integrity analysis on the power test data through deep learning, so that local omission problems in the data can be discovered, avoiding the shortcomings of traditional statistical analysis methods that are difficult to discover local anomalies, and improving the accuracy of the assessment results of the power test data quality. At the same time, the logic and consistency analysis of the power test data is performed, so that logical conflicts and labeling problems in the data can be more accurately detected, which makes up for the deficiency that the rule matching technology cannot adapt to complex logical relationships and dynamic changes, and improves the reliability of the assessment results of the power test data quality. Further independent analysis of the data features is performed, so that the correlation of the data features can be deeply analyzed, providing more comprehensive data quality information, rather than relying solely on the overall statistical features. Finally, the quality of the power test data is accurately assessed by the data anomaly ratio and the correlation coefficient between the data features, further improving the accuracy of the assessment results of the power test data quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a flowchart of a method for evaluating power test data quality based on deep learning provided by an embodiment of the present invention;

[0022] Figure 2 1 is a schematic diagram of the structure of a power test data quality assessment system based on deep learning provided by an embodiment of the present invention;

[0023] Figure 3 An embodiment diagram of an electronic device provided by an embodiment of the present invention;

[0024] Figure 4 A diagram of an embodiment of a computer-readable storage medium provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and are not to be construed as indicating or implying relative importance or implicitly specifying the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined. The term "for example" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in the present invention as "for example" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is provided to enable anyone skilled in the art to implement and use the present invention. In the following description, details are listed for illustrative purposes. It should be understood that one of ordinary skill in the art will recognize that the present invention can be implemented without these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the present invention with unnecessary detail. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0027] Optional, see Figure 1 , Figure 1 : This is a flow chart of the power test data quality assessment method based on deep learning provided by the present invention. The execution subject of the power test data quality assessment method based on deep learning in the embodiment of the present invention is a data quality assessment system. Therefore, the power test data quality assessment method based on deep learning includes:

[0028] Step 10: Based on deep learning, the integrity, logic and consistency of the power test data to be evaluated are analyzed to obtain abnormal power data and normal power data.

[0029] Optionally, the data quality assessment system obtains the power test data to be evaluated and uses a deep learning model to conduct a comprehensive analysis of the power test data to determine its integrity, logic, and consistency, classifying the power test data into abnormal power data and normal power data, as described in steps 101 to 105. This includes an integrity check to determine if data elements are missing and metadata is missing; a logical analysis to determine if there are any inconsistencies in the logical relationships between data; and a consistency comparison to determine if the data labels are consistent with pre-set database labels. Therefore, it can be understood that data that does not meet the integrity, logic, and consistency requirements is marked as abnormal power data, while data that meets the standards is marked as normal power data.

[0030] In one embodiment, deep learning uses a convolutional neural network (CNN) model, and the power test data to be evaluated is a data matrix containing power parameters (such as voltage, current, power, etc.), as well as corresponding metadata (such as measurement time, device number, etc.). After being trained with a large amount of normal and abnormal data, the model can identify abnormal patterns in the data. For integrity analysis, if the model detects that an element in the data matrix is missing, or the device number in the metadata is empty, it is determined that the data has an integrity problem. In terms of logical analysis, if the voltage value is far higher than the normal range and the current value is abnormally low, which contradicts the normal power calculation formula, the model will identify it as a logical anomaly. For consistency analysis, if the device type marked in the data is inconsistent with the marking in the preset database, the model will determine it as an inconsistent anomaly. After model analysis, the system will mark data that does not meet the integrity, logic, and consistency standards as abnormal power data, and data that meets the standards as normal power data.

[0031] Step 20: classify the normal power data based on the service characteristics to obtain the target power data class.

[0032] Furthermore, the data quality assessment system categorizes normal power data based on business characteristics, including business type, business time range, and power usage pattern, thereby obtaining multiple target power data categories, as described in steps 201 through 204. Different business characteristics correspond to different power usage patterns and requirements, and categorization allows for more targeted data analysis and processing.

[0033] In one embodiment, the business types include industrial electricity consumption, commercial electricity consumption, and residential electricity consumption. The business time range is divided into weekday daytime, weekday evening, weekend daytime, and weekend evening. The electricity consumption mode can be divided into peak electricity consumption mode, off-peak electricity consumption mode, and normal electricity consumption mode. The system first preliminarily classifies the data according to the business type. For industrial electricity consumption data, it is further subdivided according to the business time range. For example, during the daytime on weekdays, industrial electricity consumption is usually in peak electricity consumption mode, and the system classifies this part of the data into the "industrial-weekday daytime-peak electricity consumption" target electricity data class. For commercial electricity consumption data, it may be in off-peak electricity consumption mode on weekend nights, and the system classifies it into the "commercial-weekend evening-off-peak electricity consumption" target electricity data class. Therefore, in this way, the system classifies normal electricity data into multiple target electricity data classes according to different business characteristics, so as to facilitate subsequent more in-depth analysis.

[0034] Step 30 : performing independent feature analysis based on the data features of each target power data class to obtain an initial correlation coefficient between the data features in each target power data class.

[0035] Furthermore, for each target power data class, the data quality assessment system analyzes the mutual independence between the data features in the target power data class, and obtains the initial correlation coefficient between the data features in each target power data class by calculation, wherein the initial correlation coefficient represents the degree of association between the data features, as specifically described in steps 301 to 304.

[0036] In one embodiment, taking the target power data class of "industry-working daytime-peak power consumption" as an example, this type of data includes data features such as voltage, current, power, and equipment operating time. The system uses the Pearson correlation coefficient to calculate the correlation between features. For the two features of voltage and current, the system traverses all data points in the target power data class and calculates the Pearson correlation coefficient between the voltage value and the current value. For example, after calculation, the correlation coefficient of voltage and current is 0.85, indicating that the two features have a strong positive correlation. For the two features of equipment operating time and power, the calculated correlation coefficient is 0.3, indicating that the correlation between them is relatively weak. By calculating the correlation between all data features in each target power data class, the system obtains the initial correlation coefficients between the data features in each target power data class. These coefficients reflect the degree of association between different features.

[0037] Step 40 : determining the data anomaly ratio of the power test data based on the abnormal number of abnormal power data and the total number of power test data.

[0038] Furthermore, the data quality assessment system counts the number of abnormal power data and the total number of power test data, and calculates the data abnormality ratio based on the number of abnormal data and the total number.

[0039] In one embodiment, during a power test data evaluation, the data quality evaluation system identifies 500 pieces of abnormal power data, and the total number of power test data is 5000. The data abnormality ratio obtained by calculation is 500 / 5000×100%=10%.

[0040] Step 50 : Performing a quality assessment on the power test data based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class to obtain a quality assessment result.

[0041] Furthermore, the data quality assessment system comprehensively considers the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class, performs a comprehensive quality assessment on the power test data, and finally obtains a quality assessment result, as specifically described in steps 501 to 504.

[0042] In one embodiment, the data anomaly ratio is 10%. For the target power data category "Industrial - Weekday Daytime - Peak Power Consumption," the initial correlation coefficient between data features indicates a strong correlation between voltage and current (e.g., a correlation coefficient of 0.85), while a weak correlation between equipment operating hours and power (e.g., a correlation coefficient of 0.3). The system first determines the overall degree of anomaly in the data based on the data anomaly ratio. A 10% anomaly ratio indicates that the data has some issues, but is not seriously out of control. The system then analyzes the feature correlations within each target power data category. For feature pairs with strong correlations, such as voltage and current, if such strong correlations are reasonable in actual business operations, they will be given a certain degree of recognition when assessing data quality. However, for feature pairs with weaker correlations, such as equipment operating hours and power, which should have a stronger correlation in actual business operations, the system considers them to be potentially detrimental to data quality. Taking into account the data anomaly ratio and the feature correlations within each target power data category, the system concludes that the quality assessment of this power test data is "medium quality, with some anomalies and some data feature correlations not fully consistent with business expectations."

[0043] The embodiment of the present invention performs integrity analysis on the power test data through deep learning, so that local omissions in the data can be discovered, which solves the shortcoming of difficulty in discovering local anomalies and improves the accuracy of the evaluation results of the power test data quality. The logic and consistency analysis of the power test data makes it possible to more accurately detect logical conflicts and labeling problems in the data, makes up for the inability to adapt to complex logical relationships and dynamic changes, and improves the reliability of the evaluation results of the power test data quality. The data features are further analyzed independently of each other, so that the correlation of data features can be deeply analyzed, providing more comprehensive data quality information, rather than relying solely on overall statistical features. Finally, the quality of the power test data is accurately evaluated through the data anomaly ratio and the correlation coefficient between data features, further improving the accuracy of the evaluation results of the power test data quality.

[0044] In one embodiment, steps 101 to 105 are described as follows:

[0045] Step 101: Analyze the data elements of each data in the power test data to determine whether there are omissions in the data elements of each data; if there are omissions, mark the data as first integrity abnormal data; if there are no omissions, mark the data as first candidate power data.

[0046] Optionally, the data quality assessment system analyzes each data element in the power test data to determine whether any data elements are missing. If a data element is missing, the system marks the data as first integrity abnormal data; if the data element is complete, the system marks the data as first candidate power data.

[0047] In one embodiment, a power test data record contains three data elements: voltage, current, and power. After reading this data, the data quality assessment system checks whether the data elements are complete. If only voltage and current are recorded, but power data is missing, the system identifies missing data elements and marks this data as a first-level completeness anomaly. Conversely, if the data record contains all three data elements: voltage, current, and power, the system marks it as a first-level candidate power data.

[0048] Step 102: Analyze the key information of the metadata of each piece of data in the first candidate power data to determine whether the metadata of each piece of data is missing; if there is a missing, mark the data as second integrity abnormal data; if there is no omission, mark the data as second candidate power data.

[0049] Furthermore, the data quality assessment system conducts an in-depth analysis of the key metadata information for each piece of the first candidate power data to determine whether any metadata is missing. If any metadata is missing, the system marks the data as second-level integrity anomaly data. If the metadata is complete, the system marks the data as second-level candidate power data.

[0050] In one embodiment, for a record marked as first candidate power data, its metadata should include key information such as measurement time and measurement device number. When the system checks the data's metadata and finds the measurement device number missing, even if other data elements are complete, the system will determine that the data's metadata is missing and mark it as second integrity anomaly data. If the data's key metadata information, such as measurement time and measurement device number, is complete, the system will mark it as second candidate power data.

[0051] Step 103: Analyze the logical relationship between each piece of data in the second candidate power data to determine whether there is any contradiction in the logical relationship between each piece of data; if there is a contradiction, mark the data as logical conflict abnormal data; if there is no contradiction, mark the data as third candidate power data.

[0052] Furthermore, the data quality assessment system conducts a detailed analysis of the logical relationship between each piece of data for the second candidate power data to determine whether there is any contradiction in the logical relationship between the data. Once a logical contradiction is found, the system will mark the data as logical conflict abnormal data; if the logical relationship between the data is normal, the system will mark the data as the third candidate power data. In one embodiment, taking a record in the second candidate power data as an example, according to the basic principles of electricity, power is equal to the product of voltage and current (P=U*I). For example, the voltage recorded in the data is 220 volts and the current is 5 amperes. According to logical calculation, the power should be 1100 watts. However, the power recorded in the data is 1500 watts, which is inconsistent with the normal logical calculation result. The system will mark this data as logical conflict abnormal data. If the voltage, current and power values in the data meet the power calculation formula and the logical relationship is normal, the system will mark it as the third candidate power data.

[0053] Step 104: Analyze the data label of each piece of data in the third candidate power data to determine whether the data label of each piece of data is consistent with the label in the preset database; if inconsistent, mark the data as labeled abnormal data; if consistent, mark the data as normal power data.

[0054] Furthermore, the data quality assessment system compares and analyzes the data labels of each piece of third candidate power data to determine whether the data labels of each piece of data are consistent with the labels in the preset database. If the data labels are inconsistent with the preset database labels, the system will mark the data as anomaly data; if the two are consistent, the system will mark the data as normal power data.

[0055] In one embodiment, the third candidate power data contains a data record for a power equipment type labeled "transformer," while the correct label for this equipment in the default database is "distribution cabinet." Upon detecting this discrepancy, the data quality assessment system will flag this data as abnormally labeled data. If the data label matches the label in the default database, both being "distribution cabinet," the system will mark it as normal power data.

[0056] Step 105 : Determine the first integrity abnormal data, the second integrity abnormal data, the logic conflict abnormal data, and the marked abnormal data as abnormal power data.

[0057] Furthermore, the data quality assessment system aggregates the first integrity abnormal data, the second integrity abnormal data, the logic conflict abnormal data and the marked abnormal data, and uniformly determines them as abnormal power data.

[0058] In one embodiment, after analyzing and marking the data in the previous four steps, the system aggregates all data marked as abnormal. For example, in a power test data set, there are 10 first integrity abnormal data, 5 second integrity abnormal data, 8 logical conflict abnormal data, and 3 marked abnormal data. The system identifies these 10 + 5 + 8 + 3 = 26 data items as abnormal power data, while the remaining data not marked as abnormal is normal power data.

[0059] The embodiments of the present invention use deep learning to perform integrity analysis on power test data, enabling the detection of local omissions in the data, addressing the difficulty in detecting local anomalies and improving the accuracy of power test data quality assessments. Logic and consistency analysis of power test data enables more accurate detection of logical conflicts and annotation issues in the data, addressing the inability to adapt to complex logical relationships and dynamic changes, and improving the reliability of power test data quality assessments.

[0060] In one embodiment, steps 201 to 204 are described as follows:

[0061] Step 201 : For each piece of normal power data, a business relevance analysis is performed on the business characteristics of each piece of data based on the business type to obtain a business type characteristic value of each piece of data.

[0062] Optionally, for each piece of normal power data, for different business types, the data quality assessment system extracts specific features closely related to the business characteristics from each piece of data. For example, for industrial electricity business types, the current fluctuation characteristics when the equipment starts are extracted from each piece of data; for residential electricity business types, the power change characteristics during peak hours are extracted from each piece of data.

[0063] Furthermore, for different business types, the data quality assessment system performs business relevance analysis based on specific features closely related to the business characteristics to obtain the business type characteristic value of each piece of data. The specific formula for the characteristic value of each business type of each piece of data is as follows:

[0064]

[0065] Among them, F type Indicates the characteristic value of the business type type, n indicates the number of selected features, S type,i Represents the original data value of the i-th feature of business type type; C i The coefficient representing the i-th feature is set based on the empirical correlation between the business type and the feature.

[0066] Step 202 : performing time mapping analysis on the time characteristics of each piece of data based on the business time range to obtain a business time characteristic value of each piece of data.

[0067] Furthermore, the data quality assessment system maps and converts the time features of each data item based on the business time range. For example, the power data at different times of the day can be divided into peak, off-peak, and off-peak periods according to the business time range. The data features of each period are specifically mapped and processed to obtain the business time feature value of each data item. The specific formula is as follows:

[0068]

[0069] Among them, M time Indicates the business time characteristic value after time characteristic mapping; T range Indicates the business time range identification value, for example, the peak period is set to 3, the off-peak period is set to 2, and the off-peak period is set to 1; D time Indicates the relative position value of the time point corresponding to the data within the business time range, max(D time ) represents the maximum value of the relative position value within the business time range; k represents the preset mapping adjustment coefficient.

[0070] Step 203 : performing a first classification analysis on each piece of data based on the business type characteristic value and the business time characteristic value of each piece of data to obtain an initial power data class.

[0071] Furthermore, the data quality assessment system cross-integrates the business type characteristic value and business time characteristic value of each data to obtain the comprehensive characteristic value C of each data. fusion , the specific formula is as follows:

[0072]

[0073] Furthermore, the data quality assessment system performs a first classification analysis on each piece of data based on the preset classification threshold of each power data class and the comprehensive characteristic value of each piece of data to obtain the initial power data class of each piece of data. In one embodiment, the preset classification threshold of power data class A is Th A , the preset classification threshold of power data class B is Th B , the preset classification threshold of power data class C is Th C Therefore, the initial power data class of each data is init It can be expressed as:

[0074]

[0075] Step 204 : performing a second classification analysis on each initial power data class based on the power consumption pattern of each data in each initial power data class to obtain a target power data class.

[0076] Furthermore, the data quality assessment system obtains the power curve of each data in each initial power data class, and extracts the peak, valley and fluctuation frequency in the power curve, where the peak represents the point where the power reaches a relative maximum value in the local range, and the valley represents the point where the power reaches a relative minimum value in the local range.

[0077] Furthermore, the data quality assessment system matches the peak value, valley value, and fluctuation frequency in the power curve of each data with the preset peak value, preset valley value, and preset fluctuation frequency of each power consumption pattern in the preset power consumption pattern feature template to obtain the pattern similarity between the power curve of each data and each power consumption pattern in the power consumption pattern feature template. The specific formula is as follows:

[0078]

[0079] in, Represents the power curve P(t) and the jth power consumption mode M in the power consumption mode feature template j The pattern similarity of Indicates the peak value in the power curve P(t) and the power consumption mode M j The number of preset peaks in the matching is determined based on the following conditions: the power of the peak is within a certain range and the time position of the peak is within a certain range; Indicates the valley value in the power curve P(t) and the power consumption mode M j The number of preset valley values that match the value in the table. The matching judgment conditions are the same as those for the peak values. represents the number of peaks detected in the power curve P(t), represents the number of valleys detected in the power curve P(t); F P The fluctuation frequency of the power curve P(t) is obtained by counting the total number of rising and falling edges of the power curve per unit time and dividing it by the duration. Indicates power consumption mode M j The preset fluctuation frequency in .

[0080] Furthermore, the data quality assessment system determines the power consumption pattern of each data piece based on the power consumption pattern with the highest similarity to the power curve in the power consumption pattern feature template, and performs a second classification analysis on the initial power data class of each data piece based on the power consumption pattern of each data piece to obtain the target power data class.

[0081] In one embodiment, the power consumption patterns in the power consumption pattern feature template are divided into peak power consumption patterns, off-peak power consumption patterns, and normal power consumption patterns. For the initial power data class "Industry - Daytime on Working Days", the power curve of data X is matched with the power consumption pattern feature template, and it is determined that the pattern similarity with the peak power consumption pattern in the power consumption pattern feature template is the highest. Therefore, data X is further classified from the initial power data class "Industry - Daytime on Working Days" to the target power data class "Industry - Daytime on Working Days - Peak Power Consumption".

[0082] The embodiment of the present invention gradually refines the classification standards from business type, business time range to power consumption mode, and obtains multiple target power data classes with clear business orientation and characteristics. Therefore, the characteristics can be independently analyzed based on the target power data classes, so that the correlation of data characteristics can be deeply analyzed, and more comprehensive data quality information can be provided, rather than relying solely on the overall statistical characteristics. Finally, the quality of the power test data is accurately evaluated through the data anomaly ratio and the correlation coefficient between the data characteristics, further improving the accuracy of the evaluation results of the power test data quality.

[0083] In one embodiment, steps 301 to 304 are described as follows:

[0084] Step 301 : For any two first data features and second data features in each target power data class, determine the feature similarity between the first data features and the second data features according to the feature attributes of the first data features and the second data features in each dimension.

[0085] Optionally, for each target power data class, the data quality assessment system selects any two data features in the target power data class, namely, a first data feature and a second data feature, wherein the data features have dimensional attributes such as numerical value, time change trend, and power fluctuation range.

[0086] Furthermore, the data quality assessment system obtains characteristic attributes of the first data feature and the second data feature in each dimension, and determines the characteristic similarity between the first data feature and the second data feature based on the degree of intersection of the characteristic attributes of the first data feature and the second data feature in each dimension. The specific formula of the characteristic similarity is:

[0087]

[0088] Among them, Sim(f a ,f b ) represents the first data feature f a and the second data feature f b The feature similarity between them, q represents the number of feature attributes, h l(f) represents the value of data feature f on the lth attribute.

[0089] Step 302: cluster the first data feature and the second data feature based on feature similarity to obtain multiple feature groups.

[0090] Optionally, the clustering rule in the embodiment of the present invention is to cluster two data features whose feature similarity is greater than or equal to a preset similarity threshold, where the preset similarity threshold is set according to actual conditions, such as 0.7, 0.85, etc. Therefore, the data quality assessment system clusters the first data feature and the second data feature whose feature similarity is greater than or equal to the preset similarity threshold to obtain multiple feature groups.

[0091] It should be noted that the clustering process in the embodiment of the present invention is hierarchical clustering. In one embodiment, the target power data class of "commercial-weekend evening-off-peak power consumption" includes data features of "store lighting power," "air conditioning system power," "elevator operating power," and "computer equipment power." If the feature similarity between "store lighting power" and "computer equipment power" is greater than a preset similarity threshold, then "store lighting power" and "computer equipment power" are clustered together to obtain Group 1 {"store lighting power"; "computer equipment power"}. If the feature similarity between Group 1 and "air conditioning system power" and "elevator operating power" is less than a preset similarity threshold, then no clustering is performed. If the feature similarity between "air conditioning system power" and "elevator operating power" is greater than the preset similarity threshold, then "air conditioning system power" and "elevator operating power" are clustered together to obtain Group 2 {"air conditioning system power"; "elevator operating power"}.

[0092] Step 303 : For each feature group, analysis is performed based on the feature interaction relationship and feature correlation relationship between the data features to obtain a feature interaction degree index and a feature correlation degree index respectively.

[0093] Furthermore, for each feature group, the data quality assessment system determines the feature interaction degree index between the data features based on the feature interaction relationship between the data features. The specific process is as follows:

[0094] For any two data features f1 and f2, divide the data of features f1 and f2 into different bins. For example, if the data range of feature f1 is [min1, max1], divide it into n bins, and each bin width is Δ1 = (max1-min1) / n. If the data range of feature f2 is [min2, max2], divide it into m bins, and each bin width is Δ2 = (max2-min2) / m. n and m are determined based on the characteristics of the data and the analysis requirements. For example, n and m can be selected so that the amount of data in each bin is roughly equal.

[0095] Furthermore, the number of data points n whose statistical data feature f1 falls into the i-th bin and whose data feature f2 falls into the j-th bin is ij , and then divided by the total number of data features N, we get the probability that data feature f1 falls into the i-th box and data feature f2 falls into the j-th box is P ij , therefore, the specific formula is: P ij =n ij / N.

[0096] Furthermore, the marginal distribution of the data feature f1 is calculated Marginal distribution Indicates the probability that data feature f1 falls into the i-th box. The value of data feature f2 is not considered at this time. Therefore, the marginal distribution It can be expressed as Similarly, the marginal distribution of data feature f2 can be calculated Marginal distribution Indicates the probability that the data feature f2 falls into the jth box,

[0097] Furthermore, the information entropy H(f1) of the data feature f1 is calculated. Similarly, calculate the information entropy H(f2) of data feature f2,

[0098] Furthermore, the joint information entropy H(f1,f2) of the data feature f1 and the data feature f2 is calculated. The specific formula is:

[0099] Furthermore, the mutual information I(f1,f2) of the data feature f1 and the data feature f2 is calculated. The mutual information measures the degree of mutual dependence between the data feature f1 and the data feature f2, and is calculated by the joint information entropy and the edge information entropy. I(f1,f2)=H(f1)+H(f2)-H(f1,f2).

[0100] Furthermore, the feature interaction index C(f1,f2) between the data feature f1 and the data feature f2 is calculated based on the calculated mutual information I(f1,f2). The specific formula is:

[0101]

[0102] Furthermore, for each feature group, the data quality assessment system analyzes the feature correlation relationship between the data features and determines the feature correlation degree index between the data features. The specific process is as follows:

[0103] The data features in the feature group are taken as nodes and the causal relationship between the data features is taken as edges. A dynamic Bayesian network is constructed. For any two data features f1 and f2, the network parameters are learned by the maximum likelihood estimation method to obtain the first conditional probability distribution P(f 1,t+1 |f 1,t ,f 2,t ), the second conditional probability distribution P(f 2,t+1 |f 1,t ,f 2,t ), the third conditional probability distribution P(f 1,t+1 |f 1,t ) and the fourth conditional probability distribution P(f 2,t+1 |f 2,t ), where P(f 1,t+1 |f 1,t ,f 2,t ) represents the data feature f at the given current time t 1,t and data features f 2,t At the next moment t+1, the data feature f 1,t+1 The conditional probability distribution of P(f 2,t+1 |f 1,t ,f 2,t ) represents the data feature f at the given current time t 1,t and data features f 2,t At the next moment t+1, the data feature f 2,t+1 The conditional probability distribution of P(f 1,t+1 |f 1,t ) represents the data feature f at the given current time t 1,t At the next moment t+1, the data feature f 1,t+1 The conditional probability distribution of P(f 2,t+1 |f 2,t ) represents the data feature f at the given current time t 2,t When the next moment t+1 data feature (f 2,t+1 The conditional probability distribution of .

[0104] Furthermore, the joint probability distribution P(f 1,t ,f 2,t ).

[0105] Furthermore, based on the first conditional probability distribution, the second conditional probability distribution, the third conditional probability distribution, the fourth conditional probability distribution and the joint probability distribution, the feature correlation index between the data feature f1 and the data feature f2 is calculated:

[0106] Among them, K(f1,f2) represents the feature interaction degree index between data features, D KL (x||y) represents the KL divergence of x and y, which measures the difference between x and y; t represents the time step, and T represents the total number of time steps of the time series data.

[0107] Step 304 : performing feature independence analysis based on the feature interaction index and the feature correlation index of each feature group to obtain initial correlation coefficients between data features in the target power data class.

[0108] Furthermore, the data quality assessment system performs independent feature analysis based on the feature interaction index and feature correlation index of each feature group to obtain the initial correlation coefficient between data features in the target power data class, as specifically described in steps 3041 to 3044.

[0109] The embodiment of the present invention performs independent feature analysis on data features, enabling in-depth analysis of the correlation between data features and providing more comprehensive data quality information, rather than relying solely on overall statistical features, thereby improving the accuracy of the evaluation results of power test data quality.

[0110] In one embodiment, steps 3041 to 3044 are described as follows:

[0111] Step 3041: For any first target data feature in each feature group, determine, based on the operation log of the power grid system, a second target data feature that has a propagation influence relationship and a third target data feature that does not have a propagation influence relationship when the first target data feature changes.

[0112] Optionally, the data quality assessment system reads the historical operation logs of the power grid system, and for any first target data feature in each feature group, analyzes through the operation logs the second target data feature that has a propagation impact relationship with the first target data feature when the first target data feature changes, and the third target data feature that does not have a propagation impact relationship with the first target data feature.

[0113] In one embodiment, in a feature group of the target power data class of "industry-weekday daytime-peak electricity consumption", the first target data feature is "large motor power". The power grid system operation log records in detail the power changes of various types of equipment and the power changes of other related equipment. Through in-depth mining of the logs, the data quality assessment system found that when the "large motor power" changes, the "transformer load" will change accordingly, because the fluctuation of large motor power will directly affect the load of the transformer, so the "transformer load" is determined as the second target data feature. However, the "factory lighting power" has no direct correlation with the changes in the "large motor power". Its power changes are mainly controlled by the factory lighting demand and have nothing to do with the operating status of the large motor. Therefore, the "factory lighting power" is determined as the third target data feature.

[0114] Step 3042: Perform feature correlation analysis on the first target data feature and the second target data feature based on the feature interaction degree index to obtain a feature correlation degree index.

[0115] Furthermore, the data quality assessment system determines a feature distance metric between the first target data feature and the second target data feature based on the difference in feature attributes of the first target data feature and the second target data feature in each dimension, and obtains a feature correlation index between the first target data feature and the second target data feature. The specific formula is as follows:

[0116]

[0117] Among them, Cov(f A ,f B ) represents the first target data feature f A and the second target data feature f B The characteristic correlation index between A ,f B ) represents the feature interaction index, q represents the number of feature attributes, h l (f) represents the value of data feature f on the lth attribute, Indicates the preset propagation parameters.

[0118] Step 3043: Perform feature redundancy analysis on the first target data feature and the third target data feature based on the feature association degree index to obtain a feature redundancy degree index.

[0119] Furthermore, the data quality assessment system determines whether removing the third target data feature will affect the first target data feature based on the redundancy of the feature attributes of the first target data feature and the third target data feature in each dimension, and obtains a feature redundancy index between the first target data feature and the third target data feature. The specific formula is as follows:

[0120]

[0121] Among them, R(f A ,f C ) represents the first target data feature f A and the third target data feature f C The feature redundancy index between A ,f C ) represents the feature correlation index.

[0122] Step 3044 : Perform feature independence analysis based on the feature correlation index and feature redundancy index of each feature group to obtain initial correlation coefficients between data features in the target power data class.

[0123] Furthermore, the data quality assessment system performs independent feature analysis on the feature correlation index and feature redundancy index of each feature group to obtain the initial correlation coefficient between data features in each target power data class. The specific formula is as follows:

[0124]

[0125] Among them, KV(G k ) represents the initial correlation coefficient between data features in the kth target power data class, g d [Cov(f i ,f j )] represents the feature correlation index between the i-th and j-th data features of the d-th feature group in the k-th target power data class, g d [R(f i ,f j )] represents the feature redundancy index between the i-th and j-th data features of the d-th feature group in the k-th target power data class.

[0126] The embodiment of the present invention performs independent feature analysis on data features, enabling in-depth analysis of the correlation between data features and providing more comprehensive data quality information, rather than relying solely on overall statistical features, thereby improving the accuracy of the evaluation results of power test data quality.

[0127] In one embodiment, steps 501 to 504 are described as follows:

[0128] Step 501: Determine an adapted business power data class based on a business scenario of power test data application.

[0129] Optionally, the data quality assessment system obtains the business scenario in which the power test data is applied, and determines the business power data class that is suitable for the business scenario based on the characteristics and requirements of the business scenario.

[0130] In one embodiment, there is a power test data used to support the load forecasting business scenario of the smart grid. In this scenario, it is necessary to consider the impact of the power consumption patterns of different types of electricity users (such as industry, commerce, and residents) at different times (weekdays, weekends, and different time periods) on the load. Through an in-depth understanding of the load forecasting business, the data quality assessment system determines that the adapted business power data categories are "industry-weekday daytime-peak power consumption", "commercial-weekend evening-valley power consumption", "residential-weekday evening-normal power consumption" and other categories closely related to load forecasting. These categories cover data of different types of electricity users at typical times and power consumption patterns, and can provide comprehensive and targeted data support for load forecasting.

[0131] Step 502 : Perform a comprehensive evaluation on the target power data class based on the business power data class to obtain a data category comprehensiveness index.

[0132] Furthermore, the data quality assessment system conducts a comprehensive assessment of the target power data class through the business power data class, determines the coverage of the target power data class to the business power data class, and obtains the data category comprehensiveness index of the target power data class. The business power data class set in the embodiment of the present invention can be expressed as {b1, b2, ..., b n}, the target power data class can be expressed as {t1,t2,...,t m}, for each business power data class b i , determine whether the target power data type t j To match this, the specific formula of the data category comprehensiveness index F is:

[0133] Among them, δ(b i ,t j ) represents the matching function, b i With t j When matching, δ(b i ,t j )=1, otherwise δ(b i ,t j )=0.

[0134] In one embodiment, n=6 (i.e., the business power data class should have 6 typical categories), 4 of which can be matched in the target power data class, then the data category comprehensiveness index F=4 / 6≈0.67, indicating that the coverage of the target power data class for the business power data class is 67%.

[0135] Step 503 : updating the initial correlation coefficients of the two target power data classes based on the cross-coupling influence degree between any two target power data classes to obtain an updated correlation coefficient of each target power data class.

[0136] Furthermore, the data quality assessment system obtains a power grid physical model, which is a complex network constructed using power data classes in historical data as nodes and the cross-coupling effects between power data classes as edges. Furthermore, the data quality assessment system traverses the power grid physical model based on any two target power data classes to determine the degree of cross-coupling effect between the two target power data classes.

[0137] Furthermore, the data quality assessment system updates the initial correlation coefficients of the two target power data classes according to the degree of cross-coupling influence between the two target power data classes, and obtains the updated correlation coefficient of each target power data class. The specific formula is as follows:

[0138]

[0139] Among them, C i and C j represents the initial correlation coefficient of power data classes i and j, and represents the updated correlation coefficient of power data types i and j, S ij Indicates the degree of cross-coupling influence between power data classes i and j.

[0140] Step 504 : Performing a quality assessment on the power test data based on the data anomaly ratio, the data category comprehensiveness index, and the updated correlation coefficient of each target power data category to obtain a quality assessment result.

[0141] Furthermore, the data quality assessment system performs a quality assessment on the power test data based on the data anomaly ratio, the data category comprehensiveness index and the updated correlation coefficient of each target power data class to obtain a quality assessment result, as specifically described in steps 5041 to 5044.

[0142] The embodiment of the present invention accurately evaluates the quality of power test data through the correlation coefficient between data anomaly ratio and data features, thereby improving the accuracy of the evaluation result of the power test data quality.

[0143] In one embodiment, steps 5041 to 5044 are described as follows:

[0144] Step 5041 : performing a change trend consistency analysis based on the data change trends between each target power data class and its corresponding business power data class to obtain a trend consistency index of the target power data class.

[0145] Optionally, for each target power data class, the data quality assessment system extracts a first power characteristic of each target power data class and a second power characteristic of the corresponding business power data class. Furthermore, the data quality assessment system performs a comparative analysis based on the first power characteristic of each target power data class and the second power characteristic of the corresponding business power data class, and performs a trend consistency analysis based on the changing trends of the power characteristics to obtain a trend consistency index for the target power data class.

[0146] In one embodiment, in a certain power test data set, "Industry - Weekday Daytime - Peak Power Consumption" is a target power data class, and its corresponding business power data class is also "Industry - Weekday Daytime - Peak Power Consumption". The system obtains the power data time series of these two data classes over a period of time (for example, three months). The power time series of the target power data class is recorded as The power time series of business power data is recorded as First, the time series of the target power data class Perform dynamic principal component analysis, capture the dynamic change characteristics of the time series by continuously updating the principal component space, and obtain the principal component score sequence Time series of business power data Perform the same operation to obtain the principal component score sequence Calculate the similarity between two principal component score sequences, where the similarity score is calculated as: Furthermore, the trend consistency index is calculated based on the similarity score. The specific formula is:

[0147] Among them, β i Represents the time decay factor.

[0148] For example, after calculation, S in =750, The trend consistency index is T in =[750*log 10 (750 / 50)] / 1000≈0.88.

[0149] Step 5042: Perform data difference analysis based on the data difference between each target power data class and its corresponding business power data class to obtain a deviation degree index for each target power data class.

[0150] Furthermore, for each target power data class, the data quality assessment system obtains the first Gaussian mixture model parameters of the target power data class and the second Gaussian mixture model parameters of the business power data class corresponding to each target power data class. The Gaussian mixture model parameters include mean, variance, mixing weight, etc.

[0151] Furthermore, the data quality assessment system performs data difference analysis based on the first Gaussian mixture model parameters of each target power data class and the second Gaussian mixture model parameters of its corresponding business power data class to obtain a deviation degree index for each target power data class.

[0152] Let’s continue with the example of the target power data class “Industry-Workday Daytime-Peak Power Consumption” and its corresponding business power data class. The system considers multiple data features, such as voltage amplitude, frequency fluctuation, active power factor, etc. For each feature, a feature distribution model (such as a Gaussian mixture model) is constructed to describe the distribution of the target power data class and the business power data class. For example, the Gaussian mixture model parameters of the voltage amplitude feature of the target power data class can be expressed as The Gaussian mixture model parameters of the voltage amplitude characteristics of the business power data class (mean, variance, mixing weight) can be expressed as

[0153] Computes the Kullback-Leibler divergence D between two distributions KL , the specific formula is:

[0154] Among them, N(p i ; μ, σ 2 ) represents the Gaussian distribution probability density function, p i represents a data point, and k represents the number of data points.

[0155] For M features, the total difference value D is obtained by combining the distribution distance values of multiple features total :

[0156]

[0157] Furthermore, according to the total difference value D total Calculate the deviation index D e :

[0158]

[0159] For example, M = 5 (voltage amplitude, frequency fluctuation, active power factor, reactive power factor, current harmonic content), after calculation, we get Deviation index

[0160] Step 5043 , based on the trend consistency index, deviation index and updated correlation coefficient of each target power data class, an evaluation is performed to obtain the comprehensive adaptability of the power test data to its business scenario.

[0161] Furthermore, the data quality assessment system multiplies the updated correlation coefficient by the trend consistency index and the deviation index, respectively, to obtain a first calculation result and a second calculation result. Furthermore, the data quality assessment system sums the first and second calculation results to obtain the comprehensive adaptability of the power test data to its business scenario. Therefore, the comprehensive adaptability can be expressed as: comprehensive adaptability = updated correlation coefficient * trend consistency index + updated correlation coefficient * deviation index.

[0162] Step 5044 , performing a quality assessment on the power test data based on the data anomaly ratio, the data category comprehensiveness index, and the comprehensive adaptability, and obtaining a quality assessment result.

[0163] Furthermore, the data quality assessment system performs weighted calculation on the power test data based on the preset weights of the data anomaly ratio, data category comprehensiveness index and comprehensive adaptability to obtain a comprehensive assessment value.

[0164] Furthermore, in an embodiment of the present invention, a critical range of scores corresponding to each quality assessment level is set. For example, the quality assessment levels are divided into "good quality", "medium quality" and "poor quality", wherein the critical range of scores for the quality assessment level of "good quality" can be set to be greater than or equal to 0.75, the critical range of scores for the quality assessment level of "medium quality" can be set to be greater than or equal to 0.55 and less than 0.75, and the critical range of scores for the quality assessment level of "poor quality" can be set to be less than 0.55. Therefore, the data quality assessment system determines the critical range of scores in which the comprehensive assessment value is located to perform quality assessment on the power test data and obtain a quality assessment result. For example, the comprehensive assessment value is 0.65, which is greater than or equal to 0.55 and less than 0.75. Therefore, the quality assessment result of the power test data is "medium quality".

[0165] The embodiment of the present invention accurately evaluates the quality of power test data through the correlation coefficient between data anomaly ratio and data features, thereby improving the accuracy of the evaluation result of the power test data quality.

[0166] Furthermore, the power test data quality assessment system based on deep learning provided by the present invention is described below. The power test data quality assessment system based on deep learning described below and the power test data quality assessment method based on deep learning described above can be referenced to each other.

[0167] Reference Figure 2 , Figure 2 This is a structural diagram of the power test data quality assessment system based on deep learning provided by the present invention, and the power test data quality assessment system based on deep learning includes.

[0168] A deep learning analysis module 210 is configured to perform integrity, logic, and consistency analysis on the power test data to be evaluated based on deep learning to obtain abnormal power data and normal power data;

[0169] The data classification module 220 is used to classify normal power data based on business characteristics to obtain target power data class;

[0170] A feature independence analysis module 230 is configured to perform feature independence analysis based on the data features of each target power data class to obtain an initial correlation coefficient between the data features in each target power data class;

[0171] An anomaly analysis module 240 is configured to determine a data anomaly ratio of the power test data based on the number of abnormal power data and the total number of power test data;

[0172] The data quality assessment module 250 performs quality assessment on the power test data based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class to obtain a quality assessment result.

[0173] The embodiment of the present invention performs integrity analysis on the power test data through deep learning, so that local omissions in the data can be discovered, which solves the shortcoming of difficulty in discovering local anomalies and improves the accuracy of the evaluation results of the power test data quality. The logic and consistency analysis of the power test data makes it possible to more accurately detect logical conflicts and labeling problems in the data, makes up for the inability to adapt to complex logical relationships and dynamic changes, and improves the reliability of the evaluation results of the power test data quality. The data features are further analyzed independently of each other, so that the correlation of data features can be deeply analyzed, providing more comprehensive data quality information, rather than relying solely on overall statistical features. Finally, the quality of the power test data is accurately evaluated through the data anomaly ratio and the correlation coefficient between data features, further improving the accuracy of the evaluation results of the power test data quality.

[0174] See also Figure 3 , Figure 3 This is a diagram of an embodiment of an electronic device provided by an embodiment of the present invention. Figure 3 As shown, an embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:

[0175] Based on deep learning, the integrity, logic, and consistency of the power test data to be evaluated are analyzed to obtain abnormal power data and normal power data;

[0176] Classify normal power data based on business characteristics to obtain target power data class;

[0177] Based on the data features of each target power data class, independent feature analysis is performed to obtain the initial correlation coefficient between the data features in each target power data class;

[0178] determining a data anomaly ratio of the power test data based on the abnormal number of abnormal power data and the total number of power test data;

[0179] The quality of the power test data is evaluated based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class, and the quality assessment results are obtained.

[0180] See also Figure 4 , Figure 4 Detailed description of an embodiment of a computer-readable storage medium provided by an embodiment of the present invention. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, the following steps are implemented:

[0181] Based on deep learning, the integrity, logic, and consistency of the power test data to be evaluated are analyzed to obtain abnormal power data and normal power data;

[0182] Classify normal power data based on business characteristics to obtain target power data class;

[0183] Based on the data features of each target power data class, independent feature analysis is performed to obtain the initial correlation coefficient between the data features in each target power data class;

[0184] determining a data anomaly ratio of the power test data based on the abnormal number of abnormal power data and the total number of power test data;

[0185] The quality of the power test data is evaluated based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class, and the quality assessment results are obtained.

[0186] In another aspect, the present invention further provides a computer program product, which includes a computer program. The computer program may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the following steps:

[0187] Based on deep learning, the integrity, logic, and consistency of the power test data to be evaluated are analyzed to obtain abnormal power data and normal power data;

[0188] Classify normal power data based on business characteristics to obtain target power data class;

[0189] Based on the data features of each target power data class, independent feature analysis is performed to obtain the initial correlation coefficient between the data features in each target power data class;

[0190] determining a data anomaly ratio of the power test data based on the abnormal number of abnormal power data and the total number of power test data;

[0191] The quality of the power test data is evaluated based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class, and the quality assessment results are obtained.

[0192] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method of each embodiment or certain parts of the embodiment.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for evaluating power test data quality based on deep learning, characterized in that: include: Based on deep learning, the integrity, logic, and consistency of the power test data to be evaluated are analyzed to obtain abnormal power data and normal power data; Classifying the normal power data based on business characteristics to obtain a target power data class; Based on the data features of each target power data class, independent feature analysis is performed to obtain the initial correlation coefficient between the data features in each target power data class; determining a data anomaly ratio of the power test data based on the abnormal number of the abnormal power data and the total number of the power test data; The quality of the power test data is evaluated based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class, and the quality assessment results are obtained.

2. The power test data quality assessment method based on deep learning according to claim 1 is characterized in that: The quality assessment of the power test data is performed based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class to obtain a quality assessment result, including: Determining an adapted business power data class based on the business scenario of the power test data application; Performing a comprehensive evaluation on the target power data class based on the business power data class to obtain a data category comprehensiveness index; updating the initial correlation coefficients of the two target power data classes based on the cross-coupling influence degree between any two target power data classes to obtain an updated correlation coefficient of each target power data class; The quality assessment of the power test data is performed based on the data anomaly ratio, the data category comprehensiveness index and the updated correlation coefficient of each target power data category to obtain the quality assessment result.

3. The power test data quality assessment method based on deep learning according to claim 2 is characterized in that: The quality assessment of the power test data based on the data anomaly ratio, the data category comprehensiveness index, and the updated correlation coefficient of each target power data category is performed to obtain the quality assessment result, including: Based on the data change trend between each target power data class and its corresponding business power data class, a change trend consistency analysis is performed to obtain the trend consistency index of each target power data class; Perform data difference analysis based on the data difference between each target power data class and its corresponding business power data class to obtain a deviation degree index for each target power data class; Based on the trend consistency index, deviation index and updated correlation coefficient of each target power data class, the comprehensive adaptability of the power test data to its business scenario is obtained; The quality assessment of the power test data is performed based on the data anomaly ratio, the data category comprehensiveness index and the comprehensive adaptability to obtain the quality assessment result.

4. The power test data quality assessment method based on deep learning according to claim 1, characterized in that: The independent feature analysis based on the data features of each target power data class to obtain the initial correlation coefficient between the data features in each target power data class includes: For any two first data features and second data features in each target power data class, determining the feature similarity between the first data feature and the second data feature according to the feature attributes of the first data feature and the second data feature in each dimension; Clustering the first data feature and the second data feature based on the feature similarity to obtain a plurality of feature groups; For each feature group, analysis is performed based on the feature interaction relationship and feature correlation relationship between data features to obtain the feature interaction degree index and feature correlation degree index respectively; Based on the feature interaction degree index and feature correlation degree index of each feature group, the features are analyzed for mutual independence, and the initial correlation coefficient between the data features in each target power data class is obtained.

5. The power test data quality assessment method based on deep learning according to claim 4 is characterized in that: The feature interaction degree index and feature correlation degree index based on each feature group are used to perform independent feature analysis to obtain the initial correlation coefficient between data features in each target power data class, including: For any first target data feature in each feature group, determining, based on the operation log of the power grid system, a second target data feature that has a propagation influence relationship and a third target data feature that does not have a propagation influence relationship when the first target data feature changes; Performing feature correlation analysis on the first target data feature and the second target data feature based on the feature interaction degree index to obtain a feature correlation degree index; Performing feature redundancy analysis on the first target data feature and the third target data feature based on the feature association degree index to obtain a feature redundancy degree index; Based on the feature correlation index and feature redundancy index of each feature group, the features are analyzed independently to obtain the initial correlation coefficient between the data features in each target power data class.

6. The power test data quality assessment method based on deep learning according to claim 1, characterized in that: The business characteristics include business type, business time range and power consumption mode; the normal power data is classified based on the business characteristics to obtain the target power data class, including: For each piece of normal power data, performing a business relevance analysis on the business characteristics of each piece of data based on the business type to obtain a business type characteristic value of each piece of data; Based on the business time range, the time characteristics of each data are time-mapped and analyzed to obtain the business time characteristic value of each data; Perform a first classification analysis on each piece of data based on the business type characteristic value and business time characteristic value of each piece of data to obtain the initial power data class; A second classification analysis is performed on each initial power data class based on the power consumption pattern of each data in each initial power data class to obtain the target power data class.

7. The power test data quality assessment method based on deep learning according to any one of claims 1 to 6, characterized in that: The deep learning-based analysis of the integrity, logic, and consistency of the power test data to be evaluated to obtain abnormal power data and normal power data includes: Analyze the data elements of each piece of data in the power test data to determine whether there is any omission in the data elements of each piece of data; if there is an omission, mark the data as first integrity abnormal data; if there is no omission, mark the data as first candidate power data; Analyze key information of metadata of each piece of data in the first candidate power data to determine whether the metadata of each piece of data is missing; if missing, mark the data as second integrity abnormal data; if not missing, mark the data as second candidate power data; Analyze the logical relationship between each piece of data in the second candidate power data to determine whether there is a contradiction in the logical relationship between each piece of data; if there is a contradiction, mark the data as logical conflict abnormal data; if there is no contradiction, mark the data as third candidate power data; Analyze the data label of each piece of data in the third candidate power data to determine whether the data label of each piece of data is consistent with the label in the preset database; if inconsistent, mark the data as abnormally labeled data; if consistent, mark the data as normal power data; The first integrity abnormal data, the second integrity abnormal data, the logic conflict abnormal data, and the marked abnormal data are determined as the abnormal power data.

8. A power test data quality assessment system based on deep learning, characterized in that: The method for evaluating power test data quality based on deep learning according to any one of claims 1 to 7 is applied; the power test data quality evaluation system based on deep learning comprises: A deep learning analysis module is used to analyze the integrity, logic, and consistency of the power test data to be evaluated based on deep learning, and obtain abnormal power data and normal power data; A data classification module, configured to classify the normal power data based on business characteristics to obtain a target power data class; A feature independence analysis module is used to perform feature independence analysis based on the data features of each target power data class to obtain the initial correlation coefficient between the data features in each target power data class; an abnormality analysis module, configured to determine a data abnormality ratio of the power test data based on the abnormal number of the abnormal power data and the total number of the power test data; The data quality assessment module is used to perform quality assessment on the power test data based on the data anomaly ratio and the initial correlation coefficient between the data features in each target power data class to obtain a quality assessment result.

9. An electronic device comprising: Memory for storing computer software programs; A processor for reading and executing the computer software program, wherein when the computer software program is executed by the processor, the method for evaluating the quality of power test data based on deep learning as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium storing a computer software program, wherein: When the computer software program is executed by a processor, it implements the power test data quality assessment method based on deep learning as described in any one of claims 1 to 7.