Data quality detection method and system based on time series data
By setting time periods and intervals in time series data to detect data missing and deviations, and using historical means to correct abnormal data, the missing and abnormal problems in the time series data collection process are solved, ensuring the integrity and reliability of data quality. It is suitable for industrial monitoring, financial transactions, weather forecasting and other fields.
Patent Information
- Application Number
- CN202510958256.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-11
AI Technical Summary
During the collection process, time series data may be affected by factors such as equipment failure, sensor errors, network transmission delays, and external environmental interference, resulting in data missing, errors, or abnormal fluctuations, affecting the accuracy and reliability of data analysis. Especially in the fields of industrial monitoring, financial transactions, weather forecasting, etc., minor anomalies may cause the system to respond untimely or increase prediction errors.
By setting time periods and intervals to obtain time series data, calculating the time difference between adjacent data points, detecting data missing, calculating data deviations and setting thresholds to separate normal and abnormal data, using historical means to correct abnormal data, and finally detecting and processing duplicate data to ensure data quality.
It enables timely discovery and correction of missing and anomalies in time series data, improves the integrity and accuracy of the data collection process, and ensures high reliability and accuracy at critical moments.
Smart Images

Figure CN120780982A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data detection technology, and in particular to a data quality detection method and system based on time series data. Background Art
[0002] During the collection process, time series data may be affected by various factors such as equipment failure, sensor errors, network transmission delays, and external environmental interference, resulting in missing data, errors, or abnormal fluctuations. Failure to detect and correct these problems in a timely manner is likely to cause deviations in data analysis results and even lead to wrong decisions.
[0003] In applications such as industrial monitoring, financial transactions, and weather forecasting, the real-time and accuracy of data are particularly critical. Any minor anomaly may cause the system to respond untimely or increase the prediction error, thereby affecting the stable operation of the entire process.
[0004] Time series data records the continuous state of a system or device over time. Its data quality is directly related to the accuracy and reliability of subsequent analysis, prediction, and decision-making. Therefore, quality testing of time series data is of great significance. Summary of the Invention
[0005] In view of the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a data quality detection method and system based on time series data, so as to be able to detect the data quality of time series data.
[0006] In order to achieve the above-mentioned purpose, the present invention provides the following technical solutions: a data quality detection method based on time series data, comprising: step one, setting a time period and a time interval, acquiring data at a fixed time interval within the time period to obtain multiple time series data, and based on the multiple time series data, acquiring the time difference between each two adjacent time series data; step two, comparing the time difference with the time interval, if the time difference is equal to the time interval, it means that there is no missing between the two adjacent time series data; if the time difference is greater than the time interval, it means that there is a missing between the two adjacent time series data; step three, counting the number of missing between the two adjacent time series data, if the number of missing is zero, it means that the time series data within the time period is complete and there is no missing; if the number of missing is greater than zero, it means that the time series data within the time period is incomplete and there is missing; step four, on the premise that the time series data is complete, acquiring the time series data value, and acquiring the time series data mean based on the historical records, and subtracting the time series data mean from the time series data value and processing it into a positive value to obtain the time series data. The numerical deviation of the time series data; Step 5, setting the numerical deviation threshold of the time series data, comparing the numerical deviation of each time series data with the numerical deviation threshold of the time series data respectively; if the numerical deviation of the time series data is less than or equal to the numerical deviation threshold of the time series data, it means that the numerical deviation of the time series data is within a reasonable range, and the time series data value represented by it is marked as a normal time series data value; if the numerical deviation of the time series data is greater than the numerical deviation threshold of the time series data, it means that the numerical deviation of the time series data exceeds the reasonable range, and the time series data value represented by it is marked as an abnormal time series data value; Step 6, obtaining the number of abnormal time series data values, and setting the abnormal number threshold, comparing the abnormal number with the abnormal number threshold; if the abnormal number is less than or equal to the abnormal number threshold, it means that the abnormality of the time series data value in the time period is not serious, then for each abnormal time series data value, use the average of the two time series data values before and after it to replace it; if the abnormal number is greater than the abnormal number threshold, it means that the abnormality of the time series data value in the time period is serious, and then re-acquire the time series data.
[0007] In some embodiments, the time interval between two adjacent abnormal time series data values is obtained and recorded as the abnormal time interval. When the number of abnormalities is less than or equal to the abnormal number threshold, the abnormal time interval is compared with the time interval, and different responses are obtained based on the comparison results.
[0008] In some embodiments, if the abnormal time interval is equal to the time interval, it means that the two adjacent abnormal time series data values are continuous, then the abnormality is determined to be serious and the time series data is reacquired; if the abnormal time interval is greater than the time interval, it means that the two adjacent abnormal time series data values are discontinuous, then the abnormality is maintained as not serious.
[0009] In some embodiments, when the time series data within a time period is complete and the time series data values are normal, all the time series data values are obtained again, and it is determined whether there are repeated time series data values in all the time series data values, and different responses are obtained based on the determination results.
[0010] In some embodiments, if there are no repeated time series data values, the time series data is maintained intact and the time series data values are normal; if there are repeated time series data values, the number of repeated time series data value groups is obtained, and at the same time, a repeated time series data value group threshold is set according to historical records, and the number of repeated time series data value groups is compared with the repeated time series data value group threshold, and different responses are obtained according to the comparison results.
[0011] In some embodiments, if the number of repeated time series data value groups is less than or equal to the repeated time series data value group threshold, it means that the number of repeated groups is within the normal range and does not affect the quality of the time series data; if the number of repeated time series data value groups is greater than the repeated time series data value group threshold, it means that the number of repeated groups exceeds the normal range and affects the quality of the time series data, and the repetition ratio is further analyzed.
[0012] In some embodiments, the specific process of further analyzing the repetition ratio is to obtain the number of repeated time series data values, divide the number of repeated time series data values by the total number of time series data values to obtain the repetition ratio, set a repetition ratio threshold, compare the repetition ratio with the repetition ratio threshold, and obtain different responses based on the comparison results.
[0013] In some embodiments, if the repetition ratio is less than or equal to the repetition ratio threshold, it means that the ratio of the number of repeated time series data values is low, and the mean correction method is used to correct one of the numbers of repeated time series data values; if the repetition ratio is greater than the repetition ratio threshold, it means that the ratio of the number of repeated time series data values is high, and the time series data is acquired again.
[0014] The application further provides a data quality detection system based on time series data, which is used for executing the method described above, and comprises: a time difference acquisition module, which sets a time period and a time interval, acquires data in the time period at a fixed time interval to obtain a plurality of time series data, and acquires a time difference between each two adjacent time series data based on the plurality of time series data; a first comparison module, which compares the time difference with the time interval, and if the time difference is equal to the time interval, it represents that there is no missing between the two adjacent time series data; if the time difference is greater than the time interval, it represents that there is missing between the two adjacent time series data; a statistical analysis module, which counts the number of missing between the two adjacent time series data, and if the number of missing is zero, it represents that the time series data in the time period is complete and there is no omission; if the number of missing is greater than zero, it represents that the time series data in the time period is incomplete and there is omission; a deviation acquisition module, which acquires a time series data value on the premise that the time series data is complete, acquires a time series data mean value according to historical records, and obtains a time series data value deviation after the time series data value is subtracted from the time series data mean value and is processed by a positive value; a second comparison module, which sets a time series data value deviation threshold, compares each time series data value deviation with the time series data value deviation threshold respectively, and if the time series data value deviation is less than or equal to the time series data value deviation threshold, it represents that the time series data value deviation is within a reasonable range, and then marks the time series data value represented thereby as a normal time series data value; if the time series data value deviation is greater than the time series data value deviation threshold, it represents that the time series data value deviation exceeds the reasonable range, and then marks the time series data value represented thereby as an abnormal time series data value; and an abnormality analysis module, which acquires an abnormal time series data value number, sets an abnormal number threshold, compares the abnormal number with the abnormal number threshold, and if the abnormal number is less than or equal to the abnormal number threshold, it represents that the time series data value abnormality in the time period is not serious, and then substitutes each abnormal time series data value with the mean value of the two time series data values before and after the abnormal time series data value; if the abnormal number is greater than the abnormal number threshold, it represents that the time series data value abnormality in the time period is serious, and then reacquires the time series data.
[0015] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the data quality detection method based on time series data.
[0016] Compared with the prior art, the technical scheme provided by the application has the following beneficial effects: The present invention uses a multi-level detection process to ensure that all aspects of time series data collection, correction, and duplicate detection meet the expected quality standards. First, data is acquired at fixed intervals within a preset time period. By calculating the time difference between adjacent data points and comparing it with the set interval, the continuity of the data is determined and missing data is promptly detected, ensuring the integrity and reliability of the entire collection process. Second, when the data is confirmed to be complete, the data mean is calculated based on historical records, and the absolute deviation of each data value from the mean is tested. Combined with a pre-set deviation threshold, the data is divided into normal and abnormal categories. For abnormal data, if the number does not exceed the abnormal threshold, the mean of the previous and next data is used for correction to reduce the impact of the abnormality on the overall data. Otherwise, it is considered a serious abnormality and requires re-collection of data to ensure accuracy. Finally, after the data is deemed complete and the values are normal (after necessary corrections), the data is again checked for duplicate values. The number of duplicate data groups and the duplicate ratio are calculated, and these indicators are compared with pre-defined thresholds to determine whether the duplication phenomenon is within a reasonable range. If the number of duplicate groups or the duplicate ratio exceeds the normal threshold, the severity of the duplication is determined based on whether to use correction methods for some of the duplicate data or directly re-acquire the data. Overall, this detection process not only helps to detect omissions and anomalies in the data collection process early and improve data quality, but also ensures that time series data maintains high accuracy and reliability at critical moments through timely correction or re-collection of duplicate data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the thinking logic of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] It is to be understood that the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the elements may be multiple, and the term "one" should not be understood as a limitation on the quantity.
[0020] The data quality detection method based on time series data provided by the present invention is as follows: Figure 1 Shown, including: The first step is to set a time period and a time interval for data acquisition. Data is acquired at fixed time intervals every certain time from the beginning of the time period until the end of the time period, and then the acquisition is stopped to obtain multiple time series data in a time period. The time difference between each two adjacent time series data is obtained, and each time difference is compared with the time interval. Different responses are made according to the comparison results. If the time difference is equal to the time interval, it means that there is no missing between the two adjacent time series data. If the time difference is greater than the time interval, it means that there is a missing between the two adjacent time series data. The number of missing between the two adjacent time series data is counted. If the number of missing is zero, it means that the time series data in the time period is complete and there is no missing. If the number of missing is greater than zero, it means that the time series data in the time period is incomplete and there is missing.
[0021] The second step is to obtain the specific value of each time series data when it is judged that the time series data in a time period is complete, which is recorded as time series data value. The average value of the time series data value is obtained according to the historical record, which is recorded as time series data mean. The time series data value deviation is obtained by subtracting the time series data mean from each time series data value and then taking the absolute value. In addition, the time series data value deviation threshold is set, and each time series data value deviation is compared with the time series data value deviation threshold. Different responses are made according to the comparison results. If the time series data value deviation is less than or equal to the time series data value deviation threshold, it means that the time series data value deviation is within a reasonable range. In this case, the time series data value represented by it is marked as normal time series data value. If the time series data value deviation is greater than the time series data value deviation threshold, it means that the time series data value deviation exceeds the reasonable range. In this case, the time series data value represented by it is marked as abnormal time series data value.
[0022] The third step is to obtain the number of abnormal time series data values, which is recorded as abnormal number. At the same time, the abnormal number threshold is set according to the total number of all time series data values in a time period. If the abnormal number exceeds the abnormal number threshold, it means that the abnormal number is large, which leads to serious abnormality of time series data values in the time period and cannot be used. Under this condition, the abnormal number is compared with the abnormal number threshold, and different responses are made according to the comparison results. If the abnormal number is less than or equal to the abnormal number threshold, it means that the abnormality of time series data values in the time period is not serious. In this case, the mean value of the two time series data values before and after each abnormal time series data value is used to replace it. If the abnormal number is greater than the abnormal number threshold, it means that the abnormality of time series data values in the time period is serious. In this case, the time series data is reacquired.
[0023] For example, assuming the time period is 1 hour and the data acquisition time interval is 10 minutes, that is, it is planned to collect data at 00:00, 00:10, 00:20, 00:30, 00:40, 00:50 and 01:00, a total of 7 times; in the first step, if the actual recorded time points are 00:00, 00:10, 00:20, 00:40, 00:50 and 01:00, the time difference between adjacent records can be calculated, and it is found that the interval between 00:20 and 00:40 is 20 minutes (greater than the expected 10 minutes), indicating that there is missing data here, and the number of missing data in this time period is 1; enter the second step, when judging the data records selected in the period (here it is assumed that only the complete data period is tested or the missing part is ignored), the specific value of each time is obtained, for example, the recorded values are 10, 12, 11 respectively , 30, 13, 12. The mean value is 16 calculated based on historical data. The absolute deviation of each data from 16 is then calculated in sequence, resulting in deviation values of 6, 4, 5, 14, 3, and 4. The time series data value deviation threshold is set to 5. Then, deviation values greater than 5 (i.e., 6 and 14) are marked as abnormal data, while deviation values less than or equal to 5 are considered normal. The third step is to count the number of abnormal data. In this example, the number of abnormal data is 2. Assuming that the total number of data in the period is 6 and the abnormal number threshold is set to 2, the number of abnormal data is equal to the threshold, and the abnormality is judged to be not serious. At this time, each abnormal data can be corrected by the mean of the data values before and after it. For example, the abnormal data represented by a deviation of 14 is 30, and the mean of the two data before and after is (11+13) / 2=12. If the number of abnormal data exceeds the threshold, it means that the data is seriously abnormal and the time series data for the period needs to be re-acquired.
[0024] When the number of anomalies is less than or equal to the anomaly threshold, the time interval between two adjacent anomaly time series data values is obtained, recorded as the anomaly time interval. The anomaly time interval is compared with the time interval, and different responses are determined based on the comparison result. If the anomaly time interval is equal to the time interval, it means that the two adjacent anomaly time series data values are continuous. In this case, the anomaly is determined to be serious and the time series data needs to be re-acquired. If the anomaly time interval is greater than the time interval, it means that the two adjacent anomaly time series data values are discontinuous. In this case, the anomaly is maintained as not serious. For example, if two adjacent anomaly data points appear at 00:20 and 00:30, respectively, the anomaly time interval between them is 10 minutes, which is equal to the expected time interval. This indicates that the two anomaly data points appear consecutively, indicating a serious anomaly. In this case, the anomaly is determined to be serious and the time series data needs to be re-acquired. On the other hand, if the anomaly data points appear at 00:20 and 00:40, respectively, the interval is 20 minutes, which exceeds the preset 10 minutes. This indicates that the anomaly data is discontinuous and the anomaly is still not serious. The current anomaly status is maintained and subsequent correction processing can be performed on this part of the data.
[0025] When the timing data in a time period is complete, and the timing data values are normal (abnormal non-serious timing data values are equivalent to normal after correction), all timing data values are obtained again, and it is judged whether there are repeated timing data values in all timing data values. If there are no repeated timing data values, the timing data is complete and the timing data values are normal. If there are repeated timing data values, the number of repeated timing data value groups is obtained, and at the same time, the number of repeated timing data value groups is set according to the historical record, and the number of repeated timing data value groups is compared with the number of repeated timing data value groups threshold, and different responses are obtained according to the comparison result. If the number of repeated timing data value groups is less than or equal to the number of repeated timing data value groups threshold, it represents that the number of repeated groups is within the normal range, and does not affect the timing data quality. If the number of repeated timing data value groups is greater than the number of repeated timing data value groups threshold, it represents that the number of repeated groups exceeds the normal range, and affects the timing data quality. In this case, the number of repeated timing data values is obtained, and the repeated proportion is obtained by dividing the number of repeated timing data values by the total number of timing data values. The repeated proportion threshold is set, and the repeated proportion is compared with the repeated proportion threshold. Different responses are obtained according to the comparison result. If the repeated proportion is less than or equal to the repeated proportion threshold, it represents that the number of repeated timing data values is low. In this case, one of the repeated timing data values is modified by using the above-mentioned correction method. If the repeated proportion is greater than the repeated proportion threshold, it represents that the number of repeated timing data values is high. In this case, the timing data is reacquired. For example, the number of repeated timing data value groups threshold is set to 3 groups. If the actual detected number of repeated groups does not exceed 3 groups, it is considered that the repetition is within the normal range and does not significantly affect the data quality. Otherwise, if the number of repeated groups exceeds 3 groups, the severity of the repetition needs to be further evaluated, that is, the number of all repeated values is counted, and the repeated proportion is obtained by dividing the number by the total number of timing data values in the entire period. Assuming that the total number is 100 data points, if 20 repeated data points are detected, the repeated proportion is 20%. Assuming that the repeated proportion threshold is set to 25%, the repeated proportion 20% is less than the repeated proportion threshold 25%, and the repeated data is adjusted according to the mean value correction method of the data before and after the time. If the repeated proportion is 30%, it exceeds 25%, and it is determined that a high proportion of repeated data affects the accuracy and reliability of the timing data. At this time, it should be considered that the data quality has been greatly affected, and the timing data needs to be reacquired.
[0026] In summary, first, a fixed time period and data acquisition interval are preset, data is acquired at intervals within the period, and the time difference between adjacent data points is calculated. If the time difference is equal to the set interval, it indicates that the data is continuous, and if it is greater than the interval, there is a missing value. The number of missing values is counted to determine the data integrity. On the premise of data integrity, the value at each time is detected. First, the overall mean is calculated based on historical data, then the absolute deviation of each data from the mean is calculated, and a deviation threshold is set. Data within a reasonable range of deviation is marked as normal, and data outside the range is considered abnormal. Next, the number of abnormal data is counted and compared with the preset abnormal number threshold. If the number of abnormalities does not exceed the threshold, the mean of the two points before and after the abnormal value is used for correction, and the time interval between adjacent abnormal data is checked. If consecutive abnormalities occur, the abnormalities are serious and need to be re-collected. Otherwise, continue to correct. When the data is considered complete and normal after the above processing, the full cycle data is acquired again and detected for repeated values. If there is no repetition, the status is maintained. If there is a repetition, the number of repeated groups is counted and compared with the threshold of the number of repeated groups. If the number of groups exceeds the threshold, the proportion of repeated data to the total data is calculated and compared with the set threshold. If the proportion is low, adjust one group of data using the correction method. Otherwise, if the repeated proportion is high, the data quality is damaged and the data needs to be re-collected.
[0027] The processes described above with reference to the flowcharts can be implemented as computer software programs in accordance with embodiments of the present disclosure. Embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication section, and / or installed from a detachable medium. When the computer program is executed by a central processing unit, the above-described functions defined in the methods of the present application are performed. It should be noted that the computer readable medium of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but not limited to, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, in which a computer readable program code is carried. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that can send, propagate or transfer a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to, wireless, wire, optical cable, RF, or any suitable combination of the above.
[0028] The computer program product of the present application can be a computer program product, which is a machine-readable medium (or media) having stored therein some code (i.e., some computer code or software) that, when executed by a machine, causes the machine to perform any of the functions disclosed herein. Note that the computer program product can be a non-transitory computer program product. The term "non-transitory" does not mean that the computer program product is entirely non-transitory during the entire period of time that the computer program product exists or is in use. The term "non-transitory" means that the computer program product is not maintained in a transitory signal form for any duration of time. In other words, the computer program product is maintained in a non-transitory, tangible form for at least some duration of time while the computer program product is in use.
[0029] Those skilled in the art will understand that the above description is only one implementation of the application in view of the teachings of the present application. Therefore, many changes and modifications can be made by those skilled in the art to the functions and implementations described herein without departing from the scope of the present application which is set forth in the following claims.
Claims
1. A data quality detection method based on time series data, characterized in that: include: Step 1: Set a time period and time interval, acquire data at fixed time intervals within the time period, obtain multiple time series data, and obtain the time difference between each two adjacent time series data based on the multiple time series data; Step 2: Compare the time difference with the time interval. If the time difference is equal to the time interval, it means there is no missing data between the two adjacent time series data. If the time difference is greater than the time interval, it means that there is a gap between the two adjacent time series data; Step 3: Count the number of missing data between two adjacent time series data. If the number of missing data is zero, it means that the time series data within the time period is complete and there is no omission. If the number of missing items is greater than zero, it means that the time series data within the time period is incomplete and there are omissions; Step 4: Under the premise that the time series data is complete, obtain the time series data value, and obtain the time series data mean based on historical records. Subtract the time series data mean from the time series data value and perform positive value processing to obtain the time series data value deviation; Step 5: Set a time series data numerical deviation threshold, and compare each time series data numerical deviation with the time series data numerical deviation threshold. If the time series data numerical deviation is less than or equal to the time series data numerical deviation threshold, it means that the time series data numerical deviation is within a reasonable range, and the time series data value represented by it is marked as a normal time series data value; if the time series data numerical deviation is greater than the time series data numerical deviation threshold, it means that the time series data numerical deviation exceeds the reasonable range, and the time series data value represented by it is marked as an abnormal time series data value; Step 6: Obtain the number of abnormal time series data values and set an abnormal number threshold. Compare the abnormal number with the abnormal number threshold. If the abnormal number is less than or equal to the abnormal number threshold, it means that the abnormality of the time series data values in the time period is not serious. Then, for each abnormal time series data value, use the average of the two time series data values before and after it to replace it. If the number of anomalies is greater than the anomaly threshold, it means that the time series data values within the time period are seriously abnormal, and the time series data is retrieved again.
2. The data quality detection method based on time series data according to claim 1 is characterized in that: Obtain the time interval between two adjacent abnormal time series data values, record it as the abnormal time interval, and when the number of abnormalities is less than or equal to the abnormal number threshold, compare the abnormal time interval with the time interval, and come up with different responses based on the comparison results.
3. The data quality detection method based on time series data according to claim 2 is characterized in that: If the abnormal time interval is equal to the time interval, it means that the two adjacent abnormal time series data values are continuous, then the abnormality is judged to be serious and the time series data is acquired again; if the abnormal time interval is greater than the time interval, it means that the two adjacent abnormal time series data values are discontinuous, then the abnormality is maintained as not serious.
4. The data quality detection method based on time series data according to claim 3 is characterized in that: When the time series data within a time period is complete and the time series data values are normal, all the time series data values are obtained again, and it is determined whether there are repeated time series data values in all the time series data values, and different responses are obtained based on the judgment results.
5. The data quality detection method based on time series data according to claim 4 is characterized in that: If there are no repeated time series data values, the time series data is kept intact and the time series data values are normal; if there are repeated time series data values, the number of repeated time series data value groups is obtained. At the same time, a threshold for the number of repeated time series data value groups is set according to historical records, and the number of repeated time series data value groups is compared with the threshold for the number of repeated time series data value groups. Different responses are obtained based on the comparison results.
6. The data quality detection method based on time series data according to claim 5 is characterized in that: If the number of repeated time series data value groups is less than or equal to the repeated time series data value group threshold, it means that the number of repeated groups is within the normal range and does not affect the quality of the time series data. If the number of repeated time series data value groups is greater than the repeated time series data value group threshold, it means that the number of repeated groups exceeds the normal range and affects the quality of the time series data. In this case, the repetition ratio should be further analyzed.
7. The data quality detection method based on time series data according to claim 6 is characterized in that: The specific process of further analyzing the repetition ratio is to obtain the number of repeated time series data values, divide the number of repeated time series data values by the total number of time series data values to obtain the repetition ratio, set the repetition ratio threshold, compare the repetition ratio with the repetition ratio threshold, and come up with different responses based on the comparison results.
8. The data quality detection method based on time series data according to claim 7 is characterized in that: If the repetition ratio is less than or equal to the repetition ratio threshold, it means that the ratio of the number of repeated time series data values is low, and the mean correction method is used to correct one of the numbers of repeated time series data values; if the repetition ratio is greater than the repetition ratio threshold, it means that the ratio of the number of repeated time series data values is high, and the time series data is obtained again.
9. A data quality detection system based on time series data, used to execute the method according to any one of claims 1 to 8, characterized in that: include: A time difference acquisition module sets a time period and a time interval, acquires data at fixed time intervals within the time period, obtains multiple time series data, and acquires the time difference between each two adjacent time series data based on the multiple time series data; The first comparison module compares the time difference with the time interval. If the time difference is equal to the time interval, it means there is no gap between the two adjacent time series data; if the time difference is greater than the time interval, it means there is a gap between the two adjacent time series data; The statistical analysis module counts the number of missing data between two adjacent time series data. If the number of missing data is zero, it means that the time series data within the time period is complete and there are no omissions. If the number of missing data is greater than zero, it means that the time series data within the time period is incomplete and there are omissions. Deviation acquisition module, which obtains the time series data value under the premise of complete time series data, and obtains the time series data mean based on historical records. The time series data value minus the time series data mean is processed into a positive value to obtain the time series data value deviation; The comparison module sets a time series data numerical deviation threshold, and compares each time series data numerical deviation with the time series data numerical deviation threshold. If the time series data numerical deviation is less than or equal to the time series data numerical deviation threshold, it means that the time series data numerical deviation is within a reasonable range, and the time series data value represented by it is marked as a normal time series data value; if the time series data numerical deviation is greater than the time series data numerical deviation threshold, it means that the time series data numerical deviation exceeds a reasonable range, and the time series data value represented by it is marked as an abnormal time series data value; The anomaly analysis module obtains the number of abnormal time series data values, sets an anomaly threshold, and compares the number of anomalies with the anomaly threshold. If the number of anomalies is less than or equal to the anomaly threshold, it means that the anomaly of the time series data values within the time period is not serious. Then, for each abnormal time series data value, the average of the two time series data values before and after it is used to replace it. If the number of anomalies is greater than the anomaly threshold, it means that the time series data values within the time period are seriously abnormal, and the time series data is retrieved again.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the data quality detection method based on time series data as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Time series data anomaly detection method and device
CN117574298A
Power time series data anomaly detection method and system, medium and processor
CN118551887A
KR20220040659A
Cited By
Industrial equipment health state prediction system based on multi-modal data fusion
CN121659222A