Self-adaptive dynamic data cleaning method for real-time streaming data

By setting up the real-time streaming data access window and data link channel in the real-time streaming data processing, performing classification processing and abnormality evaluation, and adaptively planning the cleaning path, the problem of low cleaning efficiency in the existing technology is solved, and more efficient and accurate real-time streaming data cleaning is achieved.

CN120045848AInactive Publication Date: 2025-05-27GUIZHOU CRAFTSMAN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510523694.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art lacks adaptability when processing real-time streaming data, and cannot select appropriate cleaning methods based on streaming data in different formats, resulting in low cleaning efficiency.

Method used

By setting up a real-time streaming data access window, obtaining business demand data and performing classification processing, selecting corresponding data link channels and cleaning nodes based on the classification processing results, performing secondary classification processing and data abnormality evaluation, adaptively planning the cleaning path, and dynamically adjusting the cleaning process.

Benefits of technology

It improves the accuracy and efficiency of the real-time streaming data cleaning process, and can make real-time adjustments according to the needs of different data, reduce resource waste, and improve the satisfaction of personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045848A_ABST
    Figure CN120045848A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive dynamic data cleaning method for real-time streaming data, which relates to the field of data cleaning and comprises the following steps of: acquiring corresponding service demand data and real-time streaming data; classifying the obtained real-time streaming data for multiple times according to the service demand data, and inputting the real-time streaming data into corresponding data link channels; performing data anomaly evaluation on the corresponding real-time streaming data according to the classification processing result corresponding to the corresponding cleaning node in the data link channel to obtain corresponding node abnormal data, and setting a cleaning path of the corresponding real-time streaming data according to the node abnormal data; performing data cleaning on the real-time streaming data according to the corresponding cleaning path, performing real-time feedback on the cleaning process, and generating feedback sampling adjustment information; dynamically adjusting the acquisition process of the corresponding real-time streaming data according to the feedback sampling adjustment information, and outputting the real-time streaming data to complete the data cleaning process; the efficiency in the data cleaning process is improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data cleaning, and in particular, to an adaptive dynamic data cleaning method for real-time stream data. Background Art

[0002] With the rapid development of information technology, real-time stream data has been widely used in many fields; for example, in the Internet of Things scenario, hundreds of millions of sensor devices continuously collect various types of data, such as environmental data like temperature, humidity, pressure, and the operating status data of devices, etc., and transmit them in real time in the form of data streams. The sources of real-time stream data are extensive, the formats are diverse, and during the transmission process, it may be affected by factors such as noise interference and network failures, resulting in uncertainties in the accuracy and integrity of the data. Therefore, it is very necessary to perform data cleaning on real-time stream data.

[0003] After retrieval, the invention patent with the Chinese patent number CN117453668A discloses a processing method, device, and computer equipment for spatio-temporal stream data. The method includes: constructing a spatio-temporal stream data processing framework. When performing data processing, receiving multi-source spatio-temporal stream data through an access node, and saving the spatio-temporal stream data in the spatio-temporal stream data format. Performing data cleaning on the spatio-temporal stream data in the spatio-temporal stream data format through a cleaning node to generate a cleaning identifier, adding the cleaning identifier to the spatio-temporal stream data format to obtain a cleaned spatio-temporal stream data format. After attaching a conversion identifier to the cleaned spatio-temporal stream data format through a conversion node, storing it in a result library. Performing machine learning on the data in the result library as needed through a machine learning node, and generating an analysis result, and converting the analysis result into a result data format and storing it in the result library. Using this method can achieve the consistency and integrity of data behavior during the process of stream data from collection to use.

[0004] Compared with the prior art, the invention patent with the Chinese patent number CN117453668A can perform duplicate checks and error checks on the corresponding stream data by setting corresponding cleaning identifiers, thereby improving the accuracy and cleaning efficiency to a certain extent during the stream data cleaning process.

[0005] However, in the actual use process of the above method, only the same cleaning node can be used to perform data cleaning on stream data in different formats, and set corresponding cleaning identifiers for the stream data after data cleaning is completed. It is impossible to select corresponding methods for data cleaning according to the differences in stream data. Therefore, the problem that we need to solve is to perform data cleaning on the corresponding stream data through an adaptive dynamic data cleaning method, thereby improving the data cleaning efficiency of real-time stream data; Summary of the Invention

[0006] The object of the present invention is to solve the disadvantages of poor self - adaptability and low cleaning efficiency in the prior art, and to propose an adaptive dynamic data cleaning method for real - time stream data.

[0007] To achieve the above object, the present invention adopts the following technical solutions: An adaptive dynamic data cleaning method for real - time stream data, comprising the following steps: Step 1: Set a real - time stream data access window, and obtain corresponding service requirement data and real - time stream data through the real - time stream data access window; Step 2: Perform a first - level classification process on the obtained real - time stream data according to the service requirement data, and input the real - time stream data into the corresponding data link channels according to the results of the first - level classification process. The corresponding cleaning nodes in the data link channels perform a second - level classification process on the corresponding real - time stream data; Step 3: Perform a data abnormality assessment on the corresponding real - time stream data according to the results of the second - level classification process corresponding to the cleaning nodes in the data link channels, obtain corresponding node abnormal data, and perform an adaptive cleaning path planning according to the node abnormal data corresponding to each cleaning node in the data link channels to obtain the cleaning path of the corresponding real - time stream data; Step 4: Perform data cleaning on the corresponding real - time stream data according to the corresponding cleaning path, and perform real - time feedback on the data cleaning results to generate corresponding feedback sampling adjustment information; Step 5: Perform dynamic adaptive adjustment on the acquisition process of the corresponding real - time stream data according to the feedback sampling adjustment information until the corresponding real - time stream data completes the data cleaning process, and output the fully cleaned link data The above - mentioned technical solution further includes: The process of obtaining the corresponding service requirement data and real - time stream data includes: Set a real - time stream data access window, and a data verification unit and a data acquisition unit are arranged in the real - time stream data access window; The data verification unit is used for corresponding users to enter service requirement information, and the service requirement information includes service attribute data, service decision data, processing quality standard data, and classification system data, verify the service requirement information, and set a data cleaning management space according to the verification results; The data acquisition unit is associated with the corresponding data cleaning management space, and the data acquisition unit is used to collect the real - time stream data corresponding to the corresponding service requirement information.

[0008] Further, the process of performing a first - level classification process on the obtained real - time stream data according to the service requirement data includes: Obtain the business decision-making data and classification system data corresponding to the business requirement information, set a primary classification processing framework according to the obtained data, and set classification nodes corresponding to the data link channels of the corresponding classification results within the primary classification processing framework. Set the classification evaluation standard data corresponding to the corresponding classification nodes according to the business requirement information; Set a statistical evaluation period, and perform dynamic framework adjustment on each classification node within the primary classification processing framework according to the statistical evaluation period; Traverse and compare the obtained real-time stream data at each classification node within the primary classification processing framework, obtain the data link channel corresponding to the classification node to which the real-time stream data belongs, and obtain the primary classification processing result.

[0009] Furthermore, the process of performing secondary classification processing on the corresponding real-time stream data includes: The data link channel includes a main channel, and cleaning nodes corresponding to various data cleaning operation steps are set on the main channel. The cleaning nodes are connected to the corresponding branch channels; Obtain the processing quality standard data corresponding to the business requirement information, map the obtained processing quality standard data into the corresponding cleaning nodes, obtain the real-time stream data corresponding to the corresponding cleaning nodes, generate a real-time stream data set, analyze and process the obtained real-time stream data set, obtain the corresponding mean deviation data JP, standard deviation data BC, and moving deviation data YC, and compare the obtained deviation data with the corresponding processing quality standard data respectively to obtain the comprehensive deviation data of the real-time stream data set; Preset classification evaluation indicators, compare and analyze the obtained comprehensive deviation data with the classification evaluation indicators, obtain the secondary classification processing result, and map the secondary classification processing result into the corresponding branch channel for visual display.

[0010] Furthermore, the process of performing data abnormality evaluation on the corresponding real-time stream data according to the secondary classification processing result corresponding to the corresponding cleaning node within the data link channel includes: Obtain the cleaning abnormality results of the historical real-time stream data corresponding to the corresponding cleaning node, set a cleaning data set, divide the corresponding cleaning data set into a training set and a validation set, perform model training on the training set based on a machine learning algorithm, construct a corresponding abnormality evaluation model, and perform evaluation and verification processing on the abnormality evaluation model corresponding to the corresponding cleaning node with the corresponding validation set to output the abnormality evaluation model; Input the real-time stream data into the abnormality evaluation model corresponding to the corresponding cleaning node to obtain the corresponding abnormality evaluation data; Obtain the visual display result of the real-time stream data at the corresponding cleaning node within the corresponding data link channel, and obtain the corresponding abnormality evaluation coefficient; Integrate the anomaly evaluation coefficient and anomaly evaluation data corresponding to the corresponding cleaning nodes to obtain the node anomaly data corresponding to the corresponding cleaning nodes.

[0011] Further, the process of obtaining the cleaning path of the corresponding real-time stream data includes: Perform adaptive cleaning path planning according to the node anomaly data corresponding to each cleaning node in the corresponding data link channel for the real-time stream data; Obtain the node anomaly data and node cleaning data corresponding to the corresponding cleaning nodes in the data link channel, and set node comparison tables for the different node anomaly data and node cleaning data corresponding to each cleaning node in the data link channel; Perform correlation analysis on the node anomaly data and node cleaning data corresponding to the corresponding cleaning nodes according to the node comparison table to obtain the cleaning path correlation data; Integrate the cleaning path correlation data of the node anomaly data corresponding to each cleaning node in the node comparison table to generate a cleaning node association table; Obtain the node anomaly data obtained by each cleaning node in the data link channel for the corresponding real-time stream data, retrieve and compare according to the node anomaly data in the cleaning node association table to obtain the cleaning path correlation data corresponding to the corresponding cleaning nodes, and perform sorting analysis on the cleaning path correlation data between each cleaning node to obtain the cleaning path of the corresponding real-time stream data.

[0012] Further, the process of performing real-time feedback on the data cleaning result and generating the corresponding feedback sampling adjustment information includes: Input the real-time stream data into the corresponding cleaning nodes in sequence according to the corresponding cleaning path, and let the corresponding cleaning nodes complete the corresponding data cleaning process according to the corresponding operation standards and output the real-time cleaning result; Perform disassembly analysis on the real-time cleaning result to obtain the proportion data of the corresponding cleaning data, preset a multi-dimensional comparison curve of the proportion data corresponding to the cleaning data with respect to the sampling rate, compare the corresponding proportion data with the corresponding multi-dimensional comparison curve, output the sampling rate corresponding to the corresponding cleaning node, and generate the feedback sampling adjustment information corresponding to the corresponding cleaning node according to the corresponding sampling rate.

[0013] Further, the process of outputting the full cleaning link data includes: Dynamically and adaptively adjust the acquisition process of real-time stream data according to the feedback sampling adjustment information corresponding to the corresponding cleaning nodes, acquire the real-time stream data corresponding to the corresponding sampling rate, complete the data cleaning process for the acquired real-time stream data according to the corresponding cleaning path, and mark the sorting results and cleaning processes corresponding to each path node in the cleaning path, output the corresponding full cleaning link data, and store the cleaning process of the corresponding real-time stream data through the full cleaning link data.

[0014] The present invention has the following beneficial effects: In the present invention, the real-time stream data is classified once through the corresponding service requirement information, and the corresponding data link channels are set according to the classification result, so that different real-time stream data can complete the data cleaning process according to the corresponding requirements. The real-time stream data is classified twice through the corresponding cleaning nodes in the data link channels, so that the corresponding real-time stream data can be adjusted in real time according to the corresponding requirements throughout the data cleaning process, thereby improving the accuracy and data processing efficiency in the data cleaning process. In the present invention, the data abnormality is evaluated according to the visualization display result of the secondary classification result corresponding to the real-time stream data, the cleaning path of the corresponding real-time stream data is planned according to the data abnormality evaluation result, the operation steps corresponding to the cleaning nodes are associated according to the correlation data between the cleaning nodes corresponding to different abnormality degrees, and the cleaning process of the real-time stream data is adjusted through the corresponding cleaning path, so that different cleaning operation steps in the data cleaning process are associated with each other, reducing the resource waste in the data cleaning process and also improving the personalized requirements in the data cleaning process.

[0015] In the present invention, by giving real-time feedback on the data cleaning result and generating the corresponding feedback sampling adjustment information according to the real-time feedback result, the real-time stream data is adjusted in real time according to the corresponding requirements in the corresponding data cleaning process, thereby reducing the resource waste caused by sampling while ensuring the data cleaning accuracy rate. Description of the Drawings

[0016] Figure 1 It is a schematic structural diagram of an adaptive dynamic data cleaning method for real-time stream data proposed by the present invention. Detailed Embodiment

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment

[0018] As Figure 1 shown, an adaptive dynamic data cleaning method for real-time stream data proposed by the present invention includes the following steps: Step 1: Set a real-time stream data access window, and obtain corresponding service requirement data and real-time stream data through the real-time stream data access window; Step 2: Perform a first classification process on the obtained real-time stream data according to the service requirement data, and input the real-time stream data into the corresponding data link channels according to the results of the first classification process. The corresponding cleaning nodes in the data link channels perform a second classification process on the corresponding real-time stream data; Step 3: Evaluate the data abnormality of the corresponding real-time stream data according to the results of the second classification process corresponding to the corresponding cleaning nodes in the data link channels, obtain the corresponding node abnormal data, and plan an adaptive cleaning path according to the node abnormal data corresponding to each cleaning node in the data link channels to obtain the cleaning path of the corresponding real-time stream data; Step 4: Perform data cleaning on the corresponding real-time stream data according to the corresponding cleaning path, and perform real-time feedback on the data cleaning results to generate corresponding feedback sampling adjustment information; Step 5: Dynamically and adaptively adjust the acquisition process of the corresponding real-time stream data according to the feedback sampling adjustment information until the corresponding real-time stream data completes the data cleaning process and outputs the fully cleaned link data.

[0019] In this embodiment, it should be further noted that in the specific implementation process, the process of setting the real-time stream data access window, obtaining the corresponding real-time stream data through the real-time stream data access window, and performing marking processing on the obtained real-time stream data includes: Set up a data cleaning management platform, which is used to manage the whole process of data cleaning for the real-time stream data in the platform, so that the process of completing the data cleaning of the corresponding real-time stream data is more efficient and reasonable; A real-time stream data access window is set in the data cleaning management platform, and a data verification unit and a data acquisition unit are set in the real-time stream data access window; The data verification unit is used for the corresponding user to enter service requirement information, and the service requirement information includes corresponding service attribute data, service decision data, processing quality standard data, and classification system data, where: The service attribute data includes the fields to which the corresponding real-time stream data belongs. For example, there are multiple fields such as the Internet e-commerce field, the financial field, the Internet of Things field, the transportation field, etc., and sub-fields within each field; The business decision data is the basis data for the business type decision corresponding to the field to which the corresponding real-time stream data belongs; The classification system data is the classification standards and specifications formulated according to the business logic and experience corresponding to the field to which the corresponding real-time stream data belongs; The processing quality standard data is the process optimization standard data in the data cleaning process corresponding to the field to which the corresponding real-time stream data belongs; Verify and analyze the corresponding business requirement data, set the corresponding data cleaning management space according to the verification and analysis results, and perform data cleaning processing on the real-time stream data corresponding to the corresponding business requirement information through the data cleaning management space; The data acquisition unit is used to obtain the corresponding data access terminal according to the corresponding business requirement information in the data cleaning management space, and collect the corresponding real-time stream data through the corresponding data access terminal; Set up a data cleaning storage repository in the data cleaning management space, mark the collected real-time stream data according to the corresponding data access terminal and collection time, and store the corresponding real-time stream data into the corresponding data cleaning storage repository for data storage according to the marking processing results.

[0020] It should be further noted that in the specific implementation process, the process of performing a first classification process on the real-time stream data after the marking process and inputting the real-time stream data into the corresponding data link channel according to the first classification process result, and performing a second classification process on the corresponding real-time stream data by the corresponding cleaning node in the data link channel includes: Obtain the data cleaning storage repository corresponding to the data cleaning management space and the corresponding business requirement information, and set the first classification evaluation criteria according to the corresponding business requirement information; Obtain the business decision data and classification system data corresponding to the business requirement information, set the first classification processing framework according to the corresponding business decision data and classification system data. In the first classification processing framework, classification nodes corresponding to the data link channels of the corresponding classification results are sequentially set according to the classification system data. The classification nodes are connected to the corresponding data link channels, and the corresponding business decision data and classification system data are mapped into the corresponding classification nodes; Set the statistical evaluation period, obtain the corresponding historical real-time stream data in the corresponding cleaning management space, perform quantitative statistical evaluation on the historical real-time stream data within the corresponding statistical evaluation period, and obtain the proportion data of the real-time stream data corresponding to each classification node; Perform sorting processing according to the proportion data corresponding to each classification node obtained within the corresponding statistical evaluation period, and perform dynamic framework adjustment on the first classification processing framework according to the sorting processing results of each statistical evaluation period; Traverse and compare the real-time stream data stored in the data cleaning repository with the corresponding classification nodes in the current primary classification processing framework in sequence to obtain the classification nodes to which the corresponding real-time stream data belongs, and input the real-time stream data obtained in the data cleaning repository into the data link channels corresponding to the corresponding classification nodes according to the classification nodes to which the corresponding real-time stream data belongs; The process of the corresponding cleaning nodes in the data link channel performing secondary classification processing on the corresponding real-time stream data includes: There is a main channel set in the data link channel, and corresponding cleaning nodes are sequentially set on the main channel. The cleaning nodes include cleaning nodes corresponding to various types such as data access, duplicate value processing, missing value processing, outlier processing, format standardization processing, data verification processing, etc., and the cleaning nodes corresponding to each type are connected to the corresponding branch channels. The branch channels are used for the corresponding cleaning nodes to perform data pre-cleaning on the corresponding real-time stream data and perform classification processing according to the data pre-cleaning results; Map the corresponding processing quality standard data to the corresponding cleaning nodes. When the real-time stream data is input into the cleaning node corresponding to data access through the corresponding sampling rate, the cleaning node compares and analyzes the corresponding real-time stream data with the corresponding processing quality standard data to obtain the corresponding node deviation data. The specific implementation process includes: Set a real-time stream data set, which includes n real-time stream data, corresponding to , and obtain the corresponding mean data in the real-time stream data set , and obtain the corresponding mean deviation data JP, standard deviation data BC, and moving deviation data YC according to the corresponding real-time stream data set, where: , , , where, is the corresponding numerical value of the real-time stream data corresponding to the corresponding unit time in the real-time stream data set; Compare and analyze the mean deviation data JP, standard deviation data BC, and moving deviation data YC corresponding to the real-time stream data set with the corresponding processing quality standard data respectively to obtain the comprehensive deviation data ZP, where: , where, , and are the processing quality standard data corresponding to the mean deviation data, standard deviation data, and moving deviation data respectively, , and are the weight factors of the processing quality standard data corresponding to the mean deviation data, standard deviation data, and moving deviation data respectively; Preset classification evaluation indicators, compare the obtained comprehensive deviation data with the corresponding classification evaluation indicator data, obtain the proportion of the corresponding comparison result between the comprehensive deviation data corresponding to the real-time flow data and the classification evaluation indicator data in the branch channel, and map the corresponding real-time flow data to the corresponding branch channel for visual display according to the corresponding proportion.

[0021] It should be further noted that in the specific implementation process, the process of evaluating the data abnormality according to the secondary classification processing result of the corresponding real-time flow data by the corresponding cleaning node in the data link channel and obtaining the corresponding node abnormal data includes: Obtain the visual display result of the secondary classification processing result corresponding to the corresponding real-time flow data at the corresponding cleaning node in the corresponding data link channel, and evaluate the data abnormality of the corresponding real-time flow data based on the obtained visual display result. The specific implementation process includes: An abnormal evaluation mapping scale is set in each branch channel according to the visual display result corresponding to the corresponding proportion. The corresponding position of the abnormal evaluation mapping scale includes the abnormal evaluation coefficient value corresponding to the corresponding visual display result. Obtain the position information corresponding to the visual display result of the corresponding branch channel in the abnormal evaluation mapping scale, and obtain the corresponding abnormal evaluation coefficient a; Extract the features of the real-time flow data corresponding to the corresponding cleaning node, obtain the feature data corresponding to the corresponding real-time flow data, obtain the historical processing process data of the corresponding real-time flow data in each cleaning node of the corresponding data link channel in the corresponding data cleaning management space, and obtain the historical real-time flow data corresponding to the corresponding historical processing process data; Mark the cleaning abnormal results and the corresponding feature data of the obtained historical real-time flow data at the corresponding cleaning node, and set the cleaning data set according to the marking processing result; Divide the corresponding cleaning data set into a training set and a validation set, train a model based on the machine learning algorithm for the training set, construct a corresponding abnormal evaluation model, evaluate the abnormal evaluation model corresponding to the corresponding cleaning node with the corresponding validation set, and obtain the corresponding best model; Input the feature data corresponding to the real-time flow data corresponding to the corresponding cleaning node into the abnormal evaluation model. The abnormal evaluation model performs data analysis on the corresponding feature data and outputs the predicted cleaning abnormal result corresponding to the corresponding real-time flow data. Set corresponding quantization statistical methods according to different cleaning nodes. The quantization statistical methods include but are not limited to various methods such as the standard deviation method and the proportion statistical method. Quantify the obtained predicted cleaning abnormal result based on the corresponding quantization statistical method, and obtain the corresponding abnormal evaluation data according to the quantization processing result; Integrate the obtained abnormal evaluation data with the abnormal evaluation coefficients corresponding to the respective cleaning nodes to obtain the corresponding node abnormal data.

[0022] It should be further noted that in the specific implementation process, the process of performing adaptive cleaning path planning on the node abnormal data corresponding to each cleaning node in the data link channel to obtain the cleaning path of the corresponding real-time stream data includes: Perform adaptive cleaning path planning based on the node abnormal data corresponding to each cleaning node in the corresponding data link channel for the real-time stream data; Obtain the node abnormal data and node cleaning data corresponding to the corresponding cleaning nodes in the data link channel, and set node comparison tables for the different node abnormal data and node cleaning data corresponding to each cleaning node in the data link channel; Obtain the node comparison table, and set the node abnormal data and node cleaning data corresponding to the corresponding cleaning nodes in the node comparison table as a comparison array; Combine the comparison arrays corresponding to each cleaning node in the node comparison table pairwise in sequence, obtain a set of comparison arrays according to the pairwise combination results, and mark the corresponding different cleaning nodes U and V in the obtained set of comparison arrays, where the corresponding node abnormal data are respectively marked and , and the corresponding node cleaning data are respectively marked as and , and h is the data marking result corresponding to the corresponding node abnormal data and node cleaning data; Perform cross-correlation analysis on the comparison elements in the set of comparison arrays to respectively obtain the corresponding cross-correlation data 、 , corresponding to the node abnormal data of cleaning node U and the node abnormal data of cleaning node V, the node abnormal data of cleaning node U and the node cleaning data of cleaning node V, the node abnormal data of cleaning node V and the node cleaning data of cleaning node U, and the node cleaning data of cleaning node U and the node cleaning data of cleaning node V respectively. The process includes: Respectively obtain the corresponding data information in the set of comparison arrays according to the cross results, and obtain the corresponding cross-correlation data based on the Pearson correlation coefficient algorithm. H is the total number of corresponding data, where: , ; ; ; Integrate the obtained cross-correlation data and set the corresponding association weight coefficients , , mark the cleaning path correlation data corresponding to the corresponding comparison array set as , where: ; Integrate the cleaning path correlation data of the node abnormal data corresponding to each cleaning node in the node comparison table to generate a cleaning node association table; Obtain the node abnormal data obtained by the corresponding real-time stream data in each cleaning node in the data link channel, retrieve and compare according to the node abnormal data in the cleaning node association table, obtain the cleaning path correlation data corresponding to the corresponding cleaning node, sort and analyze the cleaning path correlation data between each cleaning node, and obtain the cleaning path of the corresponding real-time stream data.

[0023] It should be further noted that in the specific implementation process, the process of the corresponding cleaning path obtaining the corresponding real-time stream data, performing data cleaning on the real-time stream data, and generating corresponding feedback sampling adjustment information according to the data cleaning result includes: Input the real-time stream data into the corresponding cleaning node in sequence according to the corresponding cleaning path, and the corresponding cleaning node completes the corresponding data cleaning process according to the corresponding operation standard and outputs the real-time cleaning result; Decompose and analyze the cleaning process of the real-time cleaning result to obtain the proportion data of the corresponding cleaning data, preset a multi-dimensional comparison curve of the proportion data corresponding to the cleaning data with respect to the sampling rate, compare and analyze the corresponding proportion data with the corresponding multi-dimensional comparison curve, output the sampling rate corresponding to the corresponding cleaning node, and generate the feedback sampling adjustment information corresponding to the corresponding cleaning node according to the corresponding sampling rate.

[0024] It should be further noted that in the specific implementation process, the process of dynamically and adaptively adjusting the acquisition process of the corresponding real-time stream data according to the feedback sampling adjustment information until the corresponding real-time stream data completes the data cleaning process and outputs the full cleaning link data includes: Dynamically and adaptively adjust the acquisition process of the real-time stream data according to the feedback sampling adjustment information corresponding to the corresponding cleaning node, obtain the real-time stream data corresponding to the corresponding sampling rate, complete the data cleaning process for the obtained real-time stream data according to the corresponding cleaning path, mark the sorting result and cleaning process corresponding to each path node in the cleaning path, output the corresponding full cleaning link data, and store the cleaning process of the corresponding real-time stream data through the full cleaning link data.

[0025] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive dynamic data cleaning method for real-time streaming data, characterized in that: The following steps are involved: Step 1: Set up a real-time streaming data access window, and obtain corresponding business demand data and real-time streaming data through the real-time streaming data access window; Step 2: Perform a classification process on the obtained real-time stream data according to the business demand data, input the real-time stream data into the corresponding data link channel according to the result of the first classification process, and the corresponding cleaning node in the data link channel performs a second classification process on the corresponding real-time stream data; Step 3: According to the secondary classification processing results corresponding to the corresponding cleaning nodes in the data link channel, the corresponding real-time stream data is evaluated for data anomaly, and the corresponding node anomaly data is obtained. According to the node anomaly data corresponding to each cleaning node in the data link channel, adaptive cleaning path planning is performed to obtain the cleaning path of the corresponding real-time stream data; Step 4: Clean the corresponding real-time stream data according to the corresponding cleaning path, provide real-time feedback on the data cleaning results, and generate corresponding feedback sampling adjustment information; Step 5: Dynamically and adaptively adjust the acquisition process of the corresponding real-time stream data according to the feedback sampling adjustment information until the corresponding real-time stream data completes the data cleaning process and outputs the fully cleaned link data.

2. The adaptive dynamic data cleaning method for real-time streaming data according to claim 1, characterized in that: The process of obtaining corresponding business demand data and real-time streaming data includes: Setting a real-time streaming data access window, wherein a data verification unit and a data acquisition unit are provided in the real-time streaming data access window; The data verification unit is used for corresponding users to input business demand information, the business demand information includes business attribute data, business decision data, processing quality standard data and classification system data, to verify the business demand information, and to set the data cleaning management space according to the verification processing result; The data cleaning management space is associated with the corresponding data collection unit, and the data collection unit is used to collect real-time stream data corresponding to the corresponding business demand information.

3. The adaptive dynamic data cleaning method for real-time streaming data according to claim 2, characterized in that: The process of classifying the acquired real-time stream data according to the business demand data includes: Acquire business decision data and classification system data corresponding to the business demand information, set a classification processing framework according to the acquired data, set classification nodes of data link channels corresponding to corresponding classification results in the classification processing framework, and set classification evaluation standard data corresponding to corresponding classification nodes according to the business demand information; Set a statistical evaluation cycle, and dynamically adjust the framework of each classification node within a classification processing framework according to the statistical evaluation cycle; The obtained real-time stream data is sequentially traversed and compared at each classification node within a classification processing framework to obtain the data link channel corresponding to the classification node to which the real-time stream data belongs, and obtain a classification processing result.

4. The adaptive dynamic data cleaning method for real-time streaming data according to claim 3, characterized in that: The process of performing secondary classification processing on the corresponding real-time stream data includes: The data link channel includes a trunk channel, on which cleaning nodes corresponding to various data cleaning operation steps are arranged, and the cleaning nodes are connected to corresponding branch channels; Obtain processing quality standard data corresponding to the business demand information, map the obtained processing quality standard data to the corresponding cleaning node, obtain the real-time stream data corresponding to the corresponding cleaning node, generate a real-time stream data set, analyze and process the obtained real-time stream data set, obtain the corresponding mean deviation data JP, standard deviation data BC and moving deviation data YC, compare and analyze the obtained deviation data with the corresponding processing quality standard data, and obtain comprehensive deviation data of the real-time stream data set; The classification evaluation index is preset, the obtained comprehensive deviation data is compared and analyzed with the classification evaluation index, the secondary classification processing result is obtained, and the secondary classification processing result is mapped to the corresponding branch channel for visual display.

5. The adaptive dynamic data cleaning method for real-time streaming data according to claim 4, characterized in that: The process of evaluating the data abnormality of the corresponding real-time stream data according to the secondary classification processing results corresponding to the corresponding cleaning nodes in the data link channel includes: Obtain the cleaning anomaly results of the historical real-time stream data corresponding to the corresponding cleaning node, set the cleaning data set, divide the corresponding cleaning data set into a training set and a validation set, perform model training on the training set based on the machine learning algorithm, build the corresponding anomaly assessment model, evaluate and verify the anomaly assessment model corresponding to the corresponding cleaning node with the corresponding validation set, and output the anomaly assessment model; Input the real-time stream data into the anomaly assessment model corresponding to the corresponding cleaning node to obtain the corresponding anomaly assessment data; Obtain the visual display results of the real-time stream data corresponding to the cleaning node in the corresponding data link channel, and obtain the corresponding abnormality assessment coefficient; The abnormality assessment coefficient and the abnormality assessment data corresponding to the corresponding cleaning node are integrated to obtain the node abnormality data corresponding to the corresponding cleaning node.

6. The adaptive dynamic data cleaning method for real-time streaming data according to claim 5, characterized in that: The process of obtaining the cleaning path of the corresponding real-time streaming data includes: Perform adaptive cleaning path planning based on node abnormality data corresponding to each cleaning node in the corresponding data link channel according to real-time stream data; Obtain node abnormality data and node cleaning data corresponding to the corresponding cleaning node in the data link channel, and set node comparison tables for different node abnormality data and node cleaning data corresponding to each cleaning node in the data link channel; According to the node comparison table, the node abnormal data corresponding to the corresponding cleaning node is correlated with the node cleaning data to obtain the corresponding cleaning path correlation data; Integrate the cleaning path correlation data of the node abnormal data corresponding to each cleaning node in the node comparison table to generate a cleaning node correlation table; Obtain the node abnormality data obtained in each cleaning node in the data link channel of the corresponding real-time stream data, search and compare in the cleaning node association table according to the node abnormality data, obtain the cleaning path correlation data corresponding to the corresponding cleaning node, sort and analyze the cleaning path correlation data between each cleaning node, and obtain the corresponding cleaning path of the corresponding real-time stream data.

7. The adaptive dynamic data cleaning method for real-time streaming data according to claim 6, characterized in that: The process of providing real-time feedback on the data cleaning results and generating corresponding feedback sampling adjustment information includes: The real-time stream data is sequentially input into the corresponding cleaning nodes according to the corresponding cleaning paths, and the corresponding cleaning nodes complete the corresponding data cleaning process according to the corresponding operation standards and output the real-time cleaning results; Perform cleaning process disassembly and analysis on the real-time cleaning results, obtain the proportion data of the corresponding cleaning data, preset a multi-dimensional comparison curve of the corresponding cleaning data corresponding to the proportion data with respect to the sampling rate, compare and analyze the corresponding proportion data with the corresponding multi-dimensional comparison curve, output the sampling rate corresponding to the corresponding cleaning node, and generate feedback sampling adjustment information corresponding to the corresponding cleaning node according to the corresponding sampling rate.

8. The adaptive dynamic data cleaning method for real-time streaming data according to claim 7, characterized in that: The process of outputting full cleaning link data includes: According to the feedback sampling adjustment information corresponding to the corresponding cleaning node, the acquisition process of the real-time stream data is dynamically and adaptively adjusted to obtain the real-time stream data corresponding to the corresponding sampling rate, and the obtained real-time stream data completes the data cleaning process according to the corresponding cleaning path, and the sorting results and cleaning processes corresponding to each path node in the cleaning path are marked, and the corresponding full cleaning link data is output, and the cleaning process of the corresponding real-time stream data is stored through the full cleaning link data.

Citation Information

Patent Citations

  • Stream data processing method and device in power internet of things, and terminal equipment

    CN113486063A

  • Processing method and device for space-time stream data and computer equipment

    CN117453668A

  • Data cleaning method for intelligent big data platform

    CN118733572A

  • Real-time data stream cleaning and monitoring system

    CN118861024A

  • Efficient data processing system based on cloud network fusion and implementation method thereof

    CN118964470A

Cited By

  • Real-time streaming data cleaning and anomaly detection method

    CN122045183A