Data cleaning method and device, electronic equipment and readable storage medium

By preprocessing and optimizing the cleaning rules of manufacturing enterprises' data, and using reinforcement learning and genetic algorithms for dynamic adjustment, the problem of fixed cleaning rules being unable to adapt to rapidly changing business scenarios has been solved, thus achieving flexibility and accuracy in the cleaning rules.

CN120994647APending Publication Date: 2025-11-21GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511039762.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, the data types of various business departments in manufacturing enterprises are diverse, the business logics are different, and the needs change rapidly. This makes it difficult for fixed cleaning rules to dynamically adapt to complex and rapidly changing business scenarios, resulting in poor flexibility of cleaning rules.

Method used

By acquiring the data to be cleaned at the current moment for preprocessing, optimizing the cleaning rules to adapt to data changes, and dynamically adjusting the cleaning rules using reinforcement learning and genetic algorithms, the cleaning is performed based on the features of the preprocessed data.

Benefits of technology

It improves the flexibility of cleaning rules, enabling timely responses to data changes, enhancing the accuracy and efficiency of data cleaning, and adapting to the dynamic characteristics of multiple departments in manufacturing enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994647A_ABST
    Figure CN120994647A_ABST
Patent Text Reader

Abstract

The invention provides a data cleaning method and device, electronic equipment and a readable storage medium, and the method is applied to cooking equipment, and comprises the steps: obtaining to-be-cleaned data at a current moment, carrying out the preprocessing of the to-be-cleaned data, obtaining the preprocessed data, obtaining a current cleaning rule, and carrying out the cleaning of the to-be-cleaned data; the method comprises the steps of preprocessing data, optimizing a current cleaning rule based on the preprocessed data to obtain a target cleaning rule, and cleaning the preprocessed data based on the target cleaning rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data cleaning technology, specifically relating to a data cleaning method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] In the era of big data, data has become a key cornerstone for decision-making and innovation across various fields. Data cleaning, as an important prerequisite for data analysis and mining, ensures data quality and usability, providing a solid foundation for subsequent data analysis and model training.

[0003] In the prior art, data from various business departments of manufacturing enterprises is cleaned using fixed cleaning rules in order to build a knowledge base for intelligent question answering and knowledge sharing.

[0004] However, due to the diverse data types, significant differences in business logic, and rapid changes in requirements across various business departments of manufacturing enterprises, fixed cleaning rules are difficult to dynamically adapt to complex and rapidly changing business scenarios, resulting in poor flexibility of cleaning rules. Summary of the Invention

[0005] This application aims to provide a data cleaning method, apparatus, electronic device, and readable storage medium, which at least solves the problem that fixed cleaning rules in the prior art are difficult to dynamically adapt to complex and rapidly changing business scenarios, and that the cleaning rules have poor flexibility.

[0006] In a first aspect, embodiments of this application disclose a data cleaning method, the method comprising:

[0007] The data to be cleaned at the current moment is obtained, and the data to be cleaned is preprocessed to obtain preprocessed data; the data to be cleaned is the data added at the current moment compared to the data at the previous moment.

[0008] Obtain the current cleaning rules and optimize them based on the preprocessed data to obtain the target cleaning rules;

[0009] Based on the target cleaning rules, the preprocessed data is cleaned.

[0010] Secondly, embodiments of this application disclose a data cleaning apparatus, the apparatus comprising:

[0011] The preprocessing module is used to acquire the data to be cleaned at the current moment and preprocess the data to be cleaned to obtain preprocessed data; the data to be cleaned is the data added at the current moment compared to the data at the previous moment.

[0012] An optimization module is used to obtain the current cleaning rules and optimize the current cleaning rules based on the preprocessed data to obtain the target cleaning rules;

[0013] The cleaning module is used to clean the preprocessed data based on the target cleaning rules.

[0014] Thirdly, embodiments of this application also disclose an electronic device, including a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0015] Fourthly, embodiments of this application also disclose a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the method described in the first aspect.

[0016] In summary, in this embodiment, the data added at the current time compared to the previous time is preprocessed, and the current cleaning rules are optimized based on the preprocessed data to obtain the target cleaning rules. Each time, the cleaning rules are first optimized based on the data that needs to be cleaned, and the optimized cleaning rules are used to clean the data. Since the preprocessed data reflects the changes in business, and the target cleaning rules are optimized based on the preprocessed data, the target cleaning rules can still adapt to the characteristics of the preprocessed data when the preprocessed data changes. The cleaning rules can respond to data changes in a timely manner, which can improve the flexibility of the cleaning rules. Attached Figure Description

[0017] In the attached diagram:

[0018] Figure 1 This is a flowchart illustrating the steps of a data cleaning method provided in an embodiment of this application;

[0019] Figure 2 This is a flowchart of another data cleaning method provided in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of a system architecture provided in an embodiment of this application;

[0021] Figure 4 This is a flowchart illustrating the steps of another data cleaning method provided in this application embodiment;

[0022] Figure 5 This is a block diagram of a data cleaning apparatus provided in an embodiment of this application;

[0023] Figure 6This is a block diagram of an electronic device provided in one embodiment of this application;

[0024] Figure 7 This is a block diagram of an electronic device according to another embodiment of the present application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] Figure 1 This is a flowchart illustrating the steps of a data cleaning method provided in an embodiment of this application. See also... Figure 1 The method may include the following steps:

[0028] Step 101: Obtain the data to be cleaned at the current time and preprocess the data to be cleaned to obtain the preprocessed data; the data to be cleaned is the data added to all data at the current time compared to all data at the previous time.

[0029] For example, the data to be cleaned at the current moment refers to the data newly added to the systems of various business departments of a manufacturing enterprise at a specific point in time (e.g., 10:00 AM daily). This is the difference between the total data volume at the current moment and the total data volume at the previous moment (e.g., 9:00 AM daily), including newly added records. For instance, at the previous moment (9:00 AM), there were 1000 "Material Production Plan" records and 800 "Material Inventory" records. The data to be cleaned at the current moment (10:00 AM) includes the data newly added compared to 9:00 AM, namely 20 "Material Production Plan" records (e.g., "Material Code A001, Planned Output 500 Units, Production Date 2025-07-10") and 15 newly added "Material Receipt Records" in the inventory records (e.g., "Material Code A001, Receipt Quantity 450 Units, Receipt Date 07 / 10 / 2025", "Material Code B002, Receipt Quantity -30 Units (obviously incorrect)").

[0030] Preprocessing is the initial processing of the newly added data, including basic operations such as removing obvious errors (e.g., negative inventory quantities, incorrectly formatted dates), standardizing data formats (e.g., converting "MM / DD / YYYY" to "YYYY-MM-DD"), and simple deduplication (e.g., duplicate order numbers), preparing for subsequent deep cleaning. Preprocessing removes obvious errors (e.g., negative numbers, duplicate records) and data with disordered formats in advance, preventing low-quality data from entering the subsequent cleaning process and reducing the complexity of deep cleaning. The 35 newly added data entries underwent preliminary processing: removing obvious errors: deleting the abnormal record "Quantity received - 30 units"; standardizing data formats: converting "Date received 07 / 10 / 2025" to "2025-07-10"; simple deduplication: if there are two duplicate records of "Material Code A001, Quantity received 450 units", retaining only one. Since the preprocessed data has a uniform format and no obvious errors, it can directly proceed to the next stage of cleaning.

[0031] Step 102: Obtain the current cleaning rules and optimize them based on the preprocessed data to obtain the target cleaning rules.

[0032] For example, the current cleaning rules refer to the set of cleaning rules preset based on business needs in the data cleaning process, including: data range verification (numerical thresholds), format verification (regular expressions), logical verification (inter-field constraints), null value handling, simple association matching, etc. Preprocessed data refers to newly added incremental data that has undergone preliminary processing. Rule optimization can be a process based on reinforcement learning algorithms, analyzing the comparison results between the features of preprocessed data samples and preset quality thresholds, and automatically adjusting the parameters of the current cleaning rules (such as expanding the data verification range and revising the anomaly judgment criteria). Alternatively, genetic algorithms can be used to optimize the cleaning rules, simulating the biological evolution process to select cleaning rules with high fitness. The target cleaning rule is the optimized cleaning rule that is more suitable for the current business scenario and data characteristics, and can dynamically adapt to data changes to improve cleaning accuracy.

[0033] Taking the current cleaning rule with a temperature range of 40℃-50℃ as an example, anything outside this range is considered abnormal. The preprocessed data packet includes 10 records, of which 3 have a temperature of 38℃ and 2 have a temperature of 51℃. During the rule optimization process, reinforcement learning algorithms can be used to analyze the preprocessed data, automatically optimize the temperature range, and form the target cleaning rule.

[0034] Step 103: Clean the preprocessed data based on the target cleaning rules.

[0035] For example, target-based cleaning rules utilize optimized target rules to perform in-depth processing on preprocessed data, accurately identifying valid data and outliers, and ensuring that the data meets preset quality standards.

[0036] Preprocessed data: 1000 air conditioner test records (preprocessed, formatted uniformly, including fields such as "temperature (°C)," and obvious errors such as negative temperatures and incorrectly formatted data have been removed). Target cleaning rules: Rules optimized using reinforcement learning algorithms, specifically "temperature range 38°C-52°C." The target cleaning rules were applied to each of the 1000 preprocessed data records: Data 1: Temperature 40°C - meets the rule - retained as valid data; Data 2: Temperature 55°C - temperature exceeds the 38°C-52°C range - marked as an outlier; Data 3: Temperature 45°C - meets the rule - retained as valid data; Data 4: Temperature 50°C - meets the rule - retained as valid data. Finally, 850 valid data records conforming to the target cleaning rules were obtained.

[0037] In this embodiment, the data added at the current time compared to the previous time is preprocessed, and the current cleaning rules are optimized based on the preprocessed data to obtain target cleaning rules. Each time, the cleaning rules are first optimized based on the data that needs to be cleaned, and the optimized cleaning rules are used to clean the data. Since the preprocessed data reflects the changes in business, and the target cleaning rules are optimized based on the preprocessed data, the target cleaning rules can still adapt to the characteristics of the preprocessed data when the preprocessed data changes. The cleaning rules can respond to data changes in a timely manner, which can improve the flexibility of the cleaning rules.

[0038] Figure 2 This is a flowchart of another data cleaning method provided in this application; see [link / reference]. Figure 2 The method may include the following steps:

[0039] Step 201: Obtain the data to be cleaned at the current time and preprocess the data to be cleaned to obtain the preprocessed data; the data to be cleaned is the data added to all data at the current time compared to all data at the previous time.

[0040] For details, please refer to step 101 above, which will not be repeated here.

[0041] Optionally, step 201 may specifically include:

[0042] The data to be cleaned is filtered using filtering rules; and / or converted according to a preset format; and / or deduplicated according to preset fields.

[0043] For example, filtering rules are used to filter the data to be cleaned. This means filtering out valid data that meets the requirements and removing invalid data that does not meet the conditions (such as negative inventory or orders with contradictory dates) based on preset filtering conditions (such as range validation: production speed must be >0; logic validation: order date ≤ delivery date; format validation: material code must conform to regular expression; null value validation: customer name must not be empty, etc.).

[0044] The data to be cleaned is converted according to a preset format: This refers to uniformly adjusting the format of the data to conform to a standard data format (e.g., converting date format from "MM / DD / YYYY" to "YYYY-MM-DD", removing units / separators from numerical values ​​and standardizing precision; unifying capitalization and terminology mapping for text, such as converting "R&D Department" to "R&D"), to ensure data format consistency. For example, the "order date" of the sales department is uniformly converted to the "YYYY-MM-DD" format to facilitate cross-departmental data integration.

[0045] Deduplication of data to be cleaned based on preset fields: This refers to identifying and deleting duplicate records based on preset key fields (such as "order number" and "material code") to avoid data redundancy. For example, for the "inbound records" of the inventory department, using "material code + inbound date" as preset fields, duplicate entries of the same batch of material inbound data can be deleted.

[0046] Format conversion unifies the heterogeneous data formats from different departments, laying the foundation for subsequent cross-departmental data association and cleaning. Filtering rules can accurately remove invalid data (such as outliers exceeding reasonable limits), and deduplication can delete duplicate records, reducing data storage costs and computational resource consumption for subsequent processing, and avoiding interference from noisy data on the quality of the knowledge base. The data to be cleaned after filtering, format conversion, and deduplication has clearer features and less noise, and can be used as high-quality samples for optimizing subsequent cleaning rules, improving the efficiency and accuracy of rule optimization in reinforcement learning algorithms.

[0047] Step 202: Obtain the current cleaning rules, determine the training data from the preprocessed data according to the preset ratio, and clean the training data based on the current cleaning rules in an optimization round to obtain the first cleaned data.

[0048] For example, the preset ratio is a pre-defined proportion (e.g., 30%, 50%) of training samples extracted from the preprocessed data to balance training effectiveness and computational efficiency, ensuring the representativeness of the training data. Training data consists of sample data extracted from the preprocessed data according to the preset ratio, used for iterative training of rule optimization. An optimization round refers to the process of performing a complete optimization of the current cleaning rule using the training data, including a closed loop of "rule application - result evaluation - parameter adjustment". The first cleaning data is the data obtained after cleaning the training data using the current cleaning rule in the optimization round; it serves as the basis for evaluating the performance of the current rule and initiating parameter optimization.

[0049] Training data is extracted from preprocessed data according to a preset ratio, which ensures that the sample size is sufficient and avoids the waste of resources in training with the full amount of data, making rule optimization more efficient.

[0050] Optionally, the current cleaning rules include: the lower limit of the temperature range is a first value, and the upper limit is a second value. Step 202 may specifically include:

[0051] Sub-step 2021: For any data point in the training data, compare the temperature included in the data with the first value and the second value;

[0052] Sub-step 2022: If the temperature included in the data is greater than or equal to the first value and less than or equal to the second value, then the data is determined as the first cleaning data.

[0053] For sub-steps 2021 and 2022, for example, the current cleaning rules are temperature-related and are set based on temperature data. These rules include two core parameters: a lower limit (first value) and an upper limit (second value) for the temperature range, used to determine the validity of the temperature data (e.g., "first value 40℃, second value 50℃", meaning the temperature must be within the range of 40℃-50℃). The first cleaning data is the valid data selected after the above comparison, conforming to the current temperature range rules. It serves as the basis for subsequent evaluation of rule adaptability and initiation of rule optimization.

[0054] For example, the temperature value of each record in the training data is compared one by one with the first value (lower limit) and the second value (upper limit) to determine whether it is within the valid range. If the temperature value is greater than or equal to the first value and less than or equal to the second value, it is considered valid data and included in the first cleaned data.

[0055] By establishing clear temperature range comparison and judgment logic, standardized cleaning of temperature data was achieved, which not only ensured the validity of the first cleaning data, but also provided a quantifiable evaluation basis for the dynamic optimization of subsequent rules.

[0056] Step 203: Calculate the reward value corresponding to the first cleaned data in the optimization round based on the preset reward function and the current cleaning rules.

[0057] For example, the preset reward function is a function used in reinforcement learning algorithms to quantify the performance of the current cleaning rule. It is a quantitative reward value output by comprehensively evaluating the degree of improvement in data quality of the cleaning result and the cleaning cost (such as computational resource consumption and human intervention requirements). A positive number indicates good rule performance, and a negative number indicates poor performance.

[0058] The current cleaning rule refers to the rule currently used to clean the data (e.g., the lower limit of the temperature range is the first value, and the upper limit is the second value), which serves as the benchmark for calculating the reward value. The reward value is a numerical value calculated using a preset reward function, reflecting the performance of the current cleaning rule in this optimization round, and is the core basis for subsequent adjustments to rule parameters (e.g., expanding the temperature range).

[0059] The preset reward function formula is: R = w1 × quality improvement degree + w2 × (-cleaning cost). Where w1 represents the quality improvement degree coefficient, which reflects the percentage improvement in accuracy, and w2 represents the cleaning cost coefficient.

[0060] By quantifying the performance of rules, precise basis is provided for dynamic adjustment of parameters, enabling the cleaning rules to continuously adapt to the dynamic characteristics of data from multiple departments of manufacturing enterprises, ultimately improving the accuracy and efficiency of data cleaning. In addition, through the automatic calculation and feedback of reward values, the adjustment of rule parameters does not require manual judgment, realizing data-driven self-optimization and reducing labor costs.

[0061] Step 204: Optimize the current cleaning rule based on the reward value, and after optimization, return to the step of cleaning the training data based on the current cleaning rule to obtain the first cleaned data, and start the next optimization round.

[0062] For example, rule optimization adjusts the parameters of the current cleaning rule based on the reward value (such as expanding the temperature range) to make the rule more suitable for the characteristics of the training data. The next optimization round is to repeat the cycle of cleaning the training data based on the new rule, calculating the new reward value, and optimizing the rule again after rule optimization.

[0063] The current cleaning rule sets the lower temperature range to 40℃ and the upper temperature range to 50℃. The training data includes 200 temperature data points, with a current average accuracy score of 70%. Cleaning the training data using the current rule reveals that 15 data points were mistakenly deleted due to an overly narrow threshold (poor quality). The reward R is calculated as: R = w1 × (quality change) + w2 × (-processing cost) = 0.7 × (-5%) + 0.3 × (-0) = -3.5 (assuming weight w1 = 0.7). In the next optimization round, the lower and upper temperature ranges are updated to 39℃ and 50℃ respectively, and the training data is cleaned again. After cleaning, the accuracy improves to 75% (R = +4). The process of optimizing the rule and then cleaning in the next round is repeated until the iteration ends.

[0064] Taking reinforcement learning as an example, the Q-learning algorithm calculates the reward value corresponding to the cleaning result based on a preset reward function, and then updates the Q-value table of the cleaning rules, gradually optimizing the parameter selection of the cleaning rules. In this way, the algorithm can automatically learn and adjust the rule parameters to adapt to different data characteristics and business needs. After the rule optimization is completed, the optimized cleaning rules are applied to all data to be cleaned for formal data cleaning. Q-learning requires multiple rounds of iterative updates to the rule parameters. The first step is to test the current rule with sample data (threshold 40℃-50℃) and calculate the reward R1 (e.g., -3.5). After updating the Q-table, new actions (e.g., adjusting to 38℃-52℃) are explored to clean the same batch of samples. A new reward R2 (e.g., +7) is calculated, and the Q-table is updated again. The above steps are repeated until the Q-table converges (e.g., reward fluctuation <1% for 10 consecutive times) or the maximum number of iterations is reached. Finally, the parameters corresponding to the highest historical reward (e.g., 38℃-52℃) are selected and applied to the preprocessed data.

[0065] Each round of optimization adjusts parameters based on the reward value from the previous round (e.g., increasing the upper limit of temperature) to reduce the accidental deletion of valid data and the retention of noisy data, thereby improving the quality of the first cleaned data round by round. Through reward value feedback, parameter adjustment, and round-by-round iteration, the cleaning rules gradually align with the characteristics of the training data (e.g., the newly added valid range in the temperature data), solving the problem that traditional fixed rules cannot cover new business scenarios.

[0066] Step 205: If the preset deadline is met, select the target cleaning rule from the cleaning rules corresponding to each optimization round based on the reward value corresponding to each optimization round.

[0067] For example, the preset cutoff condition refers to the criterion for stopping the optimization cycle. It is used to determine that the rule optimization has achieved the expected effect. It can include "the fluctuation of the reward value for N consecutive rounds is ≤ the preset threshold (e.g., 5%)", "the reward value reaches the preset target (e.g., ≥0.9)", "the number of optimization rounds reaches the maximum limit (e.g., 20 rounds)", etc.

[0068] The reward value corresponding to each optimization round is a quantitative indicator calculated using a preset reward function after cleaning the training data based on the current cleaning rules in each optimization round. It reflects the adaptability of the rules in that round. The cleaning rules corresponding to each optimization round are the cleaning rules used in each optimization round (with parameters adjusted from the previous round), such as "temperature 40℃-50℃" for round 1, "39℃-50℃" for round 2, ..., "39℃-52℃" for round 5, etc. The target cleaning rule is selected from all the cleaning rules corresponding to all optimization rounds when the preset cutoff condition is met. This rule has the highest reward value, the best stability, and is best suited to the current business scenario, and is used as the final rule for cleaning the entire dataset.

[0069] Based on the reward value corresponding to each optimization round, a target cleaning rule is selected from the cleaning rules corresponding to each optimization round. Specifically, based on the reward value corresponding to each optimization round, the optimization round with the highest reward value is determined as the target round, and the optimized cleaning rule corresponding to the target optimization round is determined as the target cleaning rule. For example, if the optimization round with the highest reward value is 5, then the optimized cleaning rule for round 5 with a temperature of 39℃-52℃ is determined as the target cleaning rule.

[0070] Through multiple rounds of optimization iterations and screening based on preset cutoff conditions, the final target cleaning rule is selected as one that is suitable for the current data characteristics and has stable performance. Each optimized cleaning rule corresponds to a unique reward value; the higher the reward value, the better the rule's suitability. Selecting the rule with the highest reward value as the target rule ensures that it performs optimally among all iterative rules.

[0071] Step 206: Clean the preprocessed data based on the target cleaning rules.

[0072] For details of this step, please refer to step 103 above, which will not be repeated here.

[0073] Optionally, after step 206, the method further includes:

[0074] Step 207: Calculate multiple index values ​​for the second cleaned data; the second cleaned data is the data obtained after cleaning the preprocessed data based on the target cleaning rules; multiple index values ​​are used to reflect the cleaning quality of the second cleaned data;

[0075] Step 208: Calculate a weighted average of multiple indicator values ​​to obtain a comprehensive quality score.

[0076] For steps 207-208, the second cleaned data refers to the data obtained after cleaning the preprocessed data using the target cleaning rules (the final rules determined after multiple rounds of optimization). This data is effective data that has been precisely screened and meets preset quality standards. Multiple metrics are used to quantitatively evaluate the quality of the second cleaned data from different dimensions, including accuracy metrics (error / outlier ratio), completeness metrics (missing rate of required fields), consistency metrics (cross-source / cross-field contradiction rate), timeliness metrics (proportion of data generated at the time meeting business requirements), and relevance metrics (field information gain / chi-square value). The weighted average is a method of calculating a comprehensive score by assigning different weights to each metric based on its importance in the business scenario (e.g., accuracy weight 0.3, consistency weight 0.2), and summing the "metric value × weight". The comprehensive quality score is a single value (usually 0-1) obtained through the weighted average, comprehensively reflecting the overall quality of the second cleaned data. It is the core basis for measuring the cleaning effect and judging whether the data meets the knowledge base's inclusion standards.

[0077] The comprehensive quality score is a value obtained by weighting and averaging multiple quality indicators of the second cleaned data (data cleaned based on the target cleaning rules). Its core reflects the overall quality level of the cleaned data and is a comprehensive consideration of the data's performance in multiple key dimensions. It can provide a quantifiable quality benchmark for the high-quality construction of multi-department knowledge bases in manufacturing enterprises.

[0078] Multiple indicator values ​​include: accuracy indicator value, completeness indicator value, consistency indicator value, timeliness indicator value, and relevance indicator value;

[0079] The number of data in the second cleaned data that is consistent with the target cleaned data, the number of data with complete preset field values, and the number of data whose generation time matches the preset time are determined respectively, to obtain the first quantity, the second quantity, and the third quantity.

[0080] For example, the first quantity is the amount of data that matches the second cleaned data with the target cleaned data. For instance, if 490 out of 500 records have "detection values" and "product models" that perfectly match the standard records, then the first quantity is 490. Preset fields: These are mandatory fields defined in the business scenario (such as "material code," "production date," "responsible person," etc.) used to assess data completeness. The second quantity is the amount of data where all three preset fields are complete. For instance, if 480 out of 500 records contain all three fields, then the second quantity is 480. The preset timeframe is the data generation time range required by the business (such as "production data must be generated within the last 7 days" or "order data must not be earlier than the system's online time") used to assess data timeliness. The third quantity is the amount of data where the "detection date" is within the last 7 days. For instance, if 475 out of 500 records meet the time range, and 25 are data from 7 days ago, then the third quantity is 475.

[0081] The association rule mining algorithm is used to analyze the association degree of each field in the second cleaned data, and to determine the correlation between each field in the second cleaned data and the preset fields.

[0082] For example, association rule mining algorithms (such as the Apriori algorithm) are used to analyze the relationships between different fields in the second-stage cleaned data. They can identify dependencies between fields (such as the fixed correspondence between "material code" and "supplier") and quantify the degree of association. The degree of association can be analyzed using association rule mining algorithms (such as Apriori). For instance, it was found that the association degree between "product model = KFR-35GW" and "test value normal" is 0.92 (meaning the probability of a normal test value for this product model is 92%), and the association degree between "test number format" and "test personnel" is 0.85 (fixed personnel correspond to fixed numbers).

[0083] Preset fields: Core fields of the business theme (such as the test result field in the product quality theme), used to evaluate the relevance of other fields to the theme (the reference benchmark for "information gain and chi-square test"). The relevance can be calculated using the information gain / chi-square test to determine the correlation between other fields and the test values ​​of the preset fields. For example, the chi-square value between "product model" and "test value" is 28.7 (highly correlated), and the information gain between "test date" and "test value" is 0.65 (moderately correlated).

[0084] By calculating the first, second, and third quantities, as well as the correlation and relevance, abstract quality dimensions such as "accuracy" and "completeness" can be transformed into directly comparable numerical values, which can solve the problems of ambiguity and subjectivity in traditional evaluation.

[0085] Based on the first quantity, second quantity, third quantity, correlation, and relevance, the accuracy index, completeness index, timeliness index, consistency index, and relevance index are determined respectively. Specifically, the ratio of the first quantity to the total quantity of data in the second cleaned data, the ratio of the second quantity to the total quantity, and the ratio of the third quantity to the total quantity are determined as the accuracy index, completeness index, and timeliness index. The ratio of the third quantity of data with a correlation greater than or equal to the preset correlation to the total quantity and the relevance corresponding to each field are determined as the consistency index and relevance index.

[0086] For example, the accuracy index value = first quantity / second total cleaned data = 490 / 500 = 0.98; the completeness index value = second quantity / second total cleaned data = 480 / 500 = 0.96; the timeliness index value = third quantity / second total cleaned data = 475 / 500 = 0.95; the consistency index value = correlation = 0.88; and the relevance index value = correlation = 0.76.

[0087] In data quality assessment, this application employs multiple algorithms to ensure the accuracy, completeness, consistency, timeliness, and relevance of the data. Accuracy assessment utilizes data comparison algorithms to compare the data to be cleaned with standard reference data line by line, accurately identifying errors and outliers. Simultaneously, validation algorithms verify the data's reasonableness, ensuring accuracy. Completeness assessment leverages integrity constraint rule checking algorithms to rigorously check for missing required fields in data records, quantifying data completeness and prompting data completion. Consistency assessment uses the Apriori data association rule mining algorithm to identify related fields between different data sources and perform consistency checks, correcting cross-departmental data inconsistencies. Timeliness assessment employs a time decay model algorithm, combining data generation time and business timeliness requirements to select valid data that meets timeliness standards. Relevance assessment utilizes feature selection algorithms such as information gain and chi-square tests to measure the relevance of data fields to business themes, selecting highly relevant data fields and optimizing the knowledge base structure. These algorithms work together to comprehensively improve the accuracy and efficiency of data quality assessment.

[0088] The five indicators—accuracy, completeness, timeliness, consistency, and relevance—correspond to the evaluation requirements for accuracy, completeness, timeliness, consistency, and relevance, respectively. This avoids the shortcomings of existing technologies in insufficient consideration of data quality and ensures that the second-cleaned data meets business needs in all dimensions.

[0089] Optionally, the method also includes:

[0090] Step 209: According to the preset data association rules, obtain the first value corresponding to the first field and the second value corresponding to the second field from the data to be cleaned from different data platforms; the first field and the second field are fields with a correlation relationship.

[0091] Step 210: If the first value and the second value are inconsistent, then one of the first value and the second value shall be corrected according to the preset standard, or the first value and the second value shall be marked respectively, and the correction result or the marking result shall be fed back to the data platform of the data to be cleaned.

[0092] For steps 209-210, the preset data association rules refer to predefined rules used to identify fields with business relationships in different data platforms, typically based on cross-departmental business logic within an enterprise. For example, production plans and inventory records are associated through material codes and production / receiving dates, verifying the consistency between planned output and received quantity (allowing reasonable fluctuations); discrepancies are corrected or flagged. Sales orders and shipping records are associated through order numbers, verifying whether order quantities match shipping quantities (e.g., ≤ order quantity); significant discrepancies are flagged for further investigation. Product BOMs and material inventory are associated through product models and material codes, verifying whether the quantity of materials required for production is ≤ the current available inventory; insufficient quantities trigger a procurement alert. Customer information and after-sales service are associated through customer IDs, verifying whether the customer contact information in the sales system matches the latest contact information in the after-sales system; in case of conflict, the higher-priority source is used for updating.

[0093] Different data platforms can include: product development relationship systems, production planning and scheduling systems, sales automation platforms, and after-sales service feedback systems. Data to be cleaned refers to incremental data from various data platforms that has not yet undergone final cleansing (data that has been preprocessed but has not undergone cross-departmental correlation verification). The first and second fields are fields with related relationships across different data platforms. For example, the planned output in the production plan and the quantity received in the inventory platform are both associated with the same material code and are the objects of cross-departmental data consistency verification. The correlation is the inherent relationship between the first and second fields formed by business logic. For example, the planned output determines the reasonable range of the quantity received, which is a dependency relationship in the business process. The first and second values ​​are the specific values ​​of the first and second fields in the data to be cleaned. For example, in the production plan, "the planned output of material A is 1000 units" is the first value, and in the inventory, "the quantity received of material A is 950 units" is the second value. Preset standards are criteria used to determine the direction of correction or marking rules when the first and second values ​​are inconsistent. For example, "correct inventory data based on production plan data," or "correct directly when the difference rate is ≤5%, and mark as pending manual review when the difference rate is >5%." Correction involves adjusting the inconsistent values ​​according to the preset standards to make the first and second values ​​consistent. For example, correcting the inventory receipt quantity from 950 units to 1000 units. Marking indicates an "inconsistent" status for the first and second values ​​when the value difference does not meet the conditions for direct correction. For example, marking the planned production and receipt quantity of material A as having a 5% difference, pending review, to facilitate subsequent manual intervention. Feedback synchronizes the correction or marking results to the corresponding data source platform. For example, feeding back the corrected receipt quantity to the inventory management platform to update the data status and ensure data synchronization across platforms.

[0094] The preset data association rule can be that there is a relationship between the material code-planned output in the production plan and the material code-inbound quantity in the inventory management platform, requiring consistency verification (based on the "material code" as the association key). The first field can be the "planned output" in the production planning platform (corresponding to material A); the second field can be the "inbound quantity" in the inventory management platform (corresponding to the same material A). The preset standard can be "when the difference between the planned output and the inbound quantity is ≤5%, the inventory data is corrected based on the production plan data; when the difference is >5%, both are marked as 'pending review' and feedback is provided."

[0095] The first value of the "Planned Production" field is retrieved from the data to be cleaned on the production planning platform, which is 1000 units (material A). The second value of the "Quantity Inbound" field is retrieved from the data to be cleaned on the inventory management platform, which is 960 units (same material A). The difference is calculated as (1000-960) / 1000 = 4% (≤5%), which meets the direct correction condition in the preset standard. Therefore, based on the production planning data, the "Quantity Inbound" (second value) on the inventory platform is corrected from 960 units to 1000 units, and the correction result is fed back to the inventory management platform to update the status of this data.

[0096] By comparing the first and second values, business-related data between different departments is identified, and consistency verification and supplementary cleaning are performed based on these relationships. This breaks down data silos between departments, resolves cross-departmental data inconsistencies, improves the consistency and coherence of overall enterprise data, provides more accurate data support for cross-departmental business analysis and decision-making, and further improves the knowledge base system. Feeding the correction or labeling results back to the data platform for the data to be cleaned ensures the consistency and accuracy of data across departments, providing a basis for subsequent data quality analysis.

[0097] In this embodiment, the data added at the current time compared to the previous time is preprocessed, and the current cleaning rules are optimized based on the preprocessed data to obtain target cleaning rules. Each time, the cleaning rules are first optimized based on the data that needs to be cleaned, and the optimized cleaning rules are used to clean the data. Since the preprocessed data reflects the changes in business, and the target cleaning rules are optimized based on the preprocessed data, the target cleaning rules can still adapt to the characteristics of the preprocessed data when the preprocessed data changes. The cleaning rules can respond to data changes in a timely manner, which can improve the flexibility of the cleaning rules.

[0098] See Figure 3 It illustrates a system architecture diagram provided in an embodiment of this application. Figure 3 The system demonstrates a three-layer architecture for intelligent data processing and dialogue agent construction, including a data collection and preprocessing layer, a data cleaning and quality assessment layer, and a cross-departmental data association and cleaning module.

[0099] 1. Data Collection and Preprocessing Layer: This layer interfaces with data source systems from various business departments, such as product development management systems and production planning and scheduling systems. It uses tools to extract heterogeneous data from these data sources into a unified data warehouse and performs preliminary cleaning, including removing obviously erroneous records, standardizing format conversion, and performing simple deduplication.

[0100] 2. Data Cleaning and Quality Assessment Layer: Deploy a dynamic rule adjustment module, use reinforcement learning algorithms to train and optimize data cleaning rules, construct a multi-dimensional data quality assessment index system, including five dimensions: accuracy, completeness, consistency, timeliness, and relevance, and establish a cross-departmental data association cleaning mechanism to identify data fields with relationships between different departments, and perform data consistency verification and supplementary cleaning.

[0101] 3. Cross-departmental Data Association and Cleaning Module: Within the data cleaning and quality assessment layer, the cross-departmental data association and cleaning module extracts relevant data fields from the data warehouses of different departments based on preset data association rules. These fields include "material code," "production date," and "planned output" from production plan data, and "material code," "inbound date," and "inventory quantity" from inventory records. By matching these fields, it identifies data with inter-departmental relationships and performs consistency checks according to rules. For example, it checks whether the "planned output" in the production plan matches the corresponding "inbound quantity" in the inventory record. If inconsistencies are found, they are corrected or marked according to business logic and data quality priorities. For instance, it might correct the "inbound quantity" in the inventory record based on the production plan data, or mark the data as inconsistent for subsequent manual review. Finally, the corrected data or marked information is fed back to the corresponding department's data warehouse to update the data status, ensuring the consistency and accuracy of data across departments. Simultaneously, a data association and cleaning log is recorded, providing a basis for subsequent data quality analysis and rule optimization.

[0102] See Figure 4 It illustrates a flowchart of another data cleaning method provided in an embodiment of this application, the steps of which include:

[0103] Step S1: Begin.

[0104] Step S2: Input the data to be cleaned.

[0105] Step S3: Retrieve the current cleaning rule library.

[0106] Step S4: Apply the current rules to clean the data.

[0107] Step S5: Extract a portion of the cleaned data samples.

[0108] Step S6: Feed the samples back to the rule optimization submodule.

[0109] Step S7: Determine whether the sample data quality meets the preset threshold. If it does, proceed to step S11; otherwise, proceed to step S8.

[0110] Step S8: Optimize rule parameters using reinforcement learning algorithms.

[0111] Step S9: Update the cleaning rule base.

[0112] Step S10: Reapply the rules to clean the data, and proceed to step S5.

[0113] Step S11: Output the target cleaning rules.

[0114] Step S12, End.

[0115] See Figure 5 This application illustrates a data cleaning device 30 provided in an embodiment of the present application, applied to smart home devices. The data cleaning device 30 includes:

[0116] The preprocessing module 301 is used to acquire the data to be cleaned at the current time and preprocess the data to be cleaned to obtain preprocessed data; the data to be cleaned is the data added to all data at the current time compared to all data at the previous time.

[0117] The optimization module 302 is used to obtain the current cleaning rules and optimize the current cleaning rules based on the preprocessed data to obtain the target cleaning rules;

[0118] The cleaning module 303 is used to clean the preprocessed data based on the target cleaning rules.

[0119] Optionally, the optimization module is specifically used for:

[0120] According to a preset ratio, training data is determined from the preprocessed data, and in an optimization round, the training data is cleaned based on the current cleaning rules to obtain the first cleaned data.

[0121] Based on the preset reward function and the current cleaning rules, calculate the reward value corresponding to the first cleaned data in the optimization round;

[0122] The current cleaning rule is optimized based on the reward value, and after optimization, the step of cleaning the training data based on the current cleaning rule to obtain the first cleaned data is returned, and the next optimization round begins.

[0123] If the preset deadline is met, the target cleaning rule is selected from the cleaning rules corresponding to each optimization round based on the reward value corresponding to each optimization round.

[0124] Optionally, the current cleaning rules include: a lower limit of the temperature range being a first value, and an upper limit being a second value; the optimization module is further specifically used for:

[0125] For any data point in the training data, the temperature included in the data is compared with the first value and the second value;

[0126] If the temperature included in the data is greater than or equal to the first value and less than or equal to the second value, then the data is determined as the first cleaning data.

[0127] Optionally, the optimization module is further used for:

[0128] Based on the reward value corresponding to each optimization round, the optimization round with the highest reward value is determined as the target round;

[0129] The optimized cleaning rule corresponding to the target optimization round is determined as the target cleaning rule.

[0130] Optionally, the device further includes:

[0131] The calculation module is used to calculate multiple indicator values ​​of the second cleaned data; the second cleaned data is the data obtained after cleaning the preprocessed data based on the target cleaning rule; the multiple indicator values ​​are used to reflect the cleaning quality of the second cleaned data.

[0132] The weighted average module is used to perform a weighted average of the multiple indicator values ​​to obtain a comprehensive quality score.

[0133] Optionally, the plurality of indicator values ​​include: accuracy indicator value, completeness indicator value, consistency indicator value, timeliness indicator value, and relevance indicator value; the calculation module is specifically used for:

[0134] The number of data in the second cleaned data that is consistent with the target cleaned data, the number of data with complete preset field values, and the number of data whose generation time matches the preset time are determined respectively to obtain the first quantity, the second quantity, and the third quantity;

[0135] The association rule mining algorithm is used to analyze the association degree of each field in the second cleaned data, and to determine the correlation between each field in the second cleaned data and the preset fields;

[0136] Based on the first quantity, the second quantity, the third quantity, the correlation degree, and the relevance degree, the accuracy index value, the completeness index value, the timeliness index value, the consistency index value, and the relevance index value are determined respectively.

[0137] Optionally, the computing module is further used for:

[0138] The ratios of the first quantity to the total quantity of data in the second cleaned data, the second quantity to the total quantity, and the third quantity to the total quantity are respectively determined as the accuracy index value, the completeness index value, and the timeliness index value;

[0139] The ratio of the third quantity of data whose correlation degree is greater than or equal to the preset correlation degree to the total quantity, and the correlation degree corresponding to each field are respectively determined as the consistency index value and the correlation index value.

[0140] Optionally, the preprocessing module is specifically used for:

[0141] The data to be cleaned is filtered using filtering rules;

[0142] And / or, perform format conversion on the data to be cleaned according to a preset format;

[0143] And / or, according to preset fields, the data to be cleaned is deduplicated.

[0144] Optionally, the device further includes:

[0145] The acquisition module is used to acquire, according to preset data association rules, the first value corresponding to the first field and the second value corresponding to the second field from the data to be cleaned from different data platforms; the first field and the second field are fields with a correlation relationship.

[0146] The feedback module is used to correct one of the first value and the second value according to a preset standard if the first value and the second value are inconsistent, or to mark the first value and the second value respectively, and to feed back the correction result or the marking result to the data platform of the data to be cleaned.

[0147] In this embodiment, the data added at the current time compared to the previous time is preprocessed, and the current cleaning rules are optimized based on the preprocessed data to obtain target cleaning rules. Each time, the cleaning rules are first optimized based on the data that needs to be cleaned, and the optimized cleaning rules are used to clean the data. Since the preprocessed data reflects the changes in business, and the target cleaning rules are optimized based on the preprocessed data, the target cleaning rules can still adapt to the characteristics of the preprocessed data when the preprocessed data changes. The cleaning rules can respond to data changes in a timely manner, which can improve the flexibility of the cleaning rules.

[0148] See Figure 6The electronic device 400 may include one or more of the following components: processing component 402, memory 404, power supply component 406, multimedia component 408, audio component 410, input / output (I / O) interface 412, sensor component 414, and communication component 416.

[0149] Processing component 402 typically controls the overall operation of electronic device 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.

[0150] Memory 404 is used to store various types of data to support the operation of electronic device 400. Examples of this data include instructions for any application or method operating on electronic device 400, contact data, phonebook data, messages, pictures, multimedia, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0151] Power supply component 406 provides power to various components of electronic device 400. Power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 400.

[0152] Multimedia component 408 includes an interface that provides an output interface between electronic device 400 and a user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When electronic device 400 is in an operating mode, such as a shooting mode or a multimedia mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0153] Audio component 410 is used to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) used to receive external audio signals when electronic device 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.

[0154] Input / output (I / O) interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0155] Sensor assembly 414 includes one or more sensors for providing state assessments of various aspects of electronic device 400. For example, sensor assembly 414 may detect the on / off state of electronic device 400, the relative positioning of components such as the display and keypad of electronic device 400, changes in position of electronic device 400 or a component of electronic device 400, the presence or absence of user contact with electronic device 400, orientation or acceleration / deceleration of electronic device 400, and temperature changes of electronic device 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0156] Communication component 416 facilitates wired or wireless communication between electronic device 400 and other devices. Electronic device 400 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0157] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement a data storage method provided in the embodiments of this application.

[0158] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of an electronic device 400 to perform the above-described method. For example, the non-transitory storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0159] Figure 7 This is a block diagram of an electronic device 500 according to another embodiment of the present invention. For example, the electronic device 500 may be provided as a server. See also Figure 7 The electronic device 500 includes a processing component 522, which further includes one or more processors, and memory resources represented by memory 532 for storing instructions, such as application programs, that can be executed by the processing component 522. The application programs stored in memory 532 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 522 is configured to execute instructions to perform a data storage method provided in embodiments of this application.

[0160] Electronic device 500 may also include a power supply component 526 configured to perform power management of electronic device 500, a wired or wireless network interface 550 configured to connect electronic device 500 to a network, and an input / output (I / O) interface 558. Electronic device 500 may operate on an operating system stored in memory 532, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0161] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0162] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A data cleaning method, characterized in that, The method includes: The data to be cleaned at the current moment is obtained, and the data to be cleaned is preprocessed to obtain preprocessed data; the data to be cleaned is the data added at the current moment compared to the data at the previous moment. Obtain the current cleaning rules and optimize them based on the preprocessed data to obtain the target cleaning rules; Based on the target cleaning rules, the preprocessed data is cleaned.

2. The method according to claim 1, characterized in that, The optimization of the current cleaning rule based on the preprocessed data to obtain the target cleaning rule includes: According to a preset ratio, training data is determined from the preprocessed data, and in an optimization round, the training data is cleaned based on the current cleaning rules to obtain the first cleaned data. Based on the preset reward function and the current cleaning rules, calculate the reward value corresponding to the first cleaned data in the optimization round; The current cleaning rule is optimized based on the reward value, and after optimization, the step of cleaning the training data based on the current cleaning rule to obtain the first cleaned data is returned, and the next optimization round begins. If the preset deadline is met, the target cleaning rule is selected from the cleaning rules corresponding to each optimization round based on the reward value corresponding to each optimization round.

3. The method according to claim 2, characterized in that, The current cleaning rules include: the lower limit of the temperature range is a first value, and the upper limit is a second value; The step of cleaning the training data based on the current cleaning rules to obtain first cleaned data includes: For any data point in the training data, the temperature included in the data is compared with the first value and the second value; If the temperature included in the data is greater than or equal to the first value and less than or equal to the second value, then the data is determined as the first cleaning data.

4. The method according to claim 2, characterized in that, The step of selecting the target cleaning rule from the cleaning rules corresponding to each optimization round based on the reward value corresponding to each optimization round includes: Based on the reward value corresponding to each optimization round, the optimization round with the highest reward value is determined as the target round; The optimized cleaning rule corresponding to the target optimization round is determined as the target cleaning rule.

5. The method according to claim 1, characterized in that, After cleaning the preprocessed data based on the target cleaning rules, the method further includes: Calculate multiple index values ​​for the second cleaned data; the second cleaned data is the data obtained after cleaning the preprocessed data based on the target cleaning rule; the multiple index values ​​are used to reflect the cleaning quality of the second cleaned data; The weighted average of the multiple indicator values ​​is used to obtain the comprehensive quality score.

6. The method according to claim 5, characterized in that, The multiple indicator values ​​include: accuracy indicator value, completeness indicator value, consistency indicator value, timeliness indicator value, and relevance indicator value; the multiple indicator values ​​for calculating the second cleaned data include: The number of data in the second cleaned data that is consistent with the target cleaned data, the number of data with complete preset field values, and the number of data whose generation time matches the preset time are determined respectively to obtain the first quantity, the second quantity, and the third quantity; The association rule mining algorithm is used to analyze the association degree of each field in the second cleaned data, and to determine the correlation between each field in the second cleaned data and the preset fields; Based on the first quantity, the second quantity, the third quantity, the correlation degree, and the relevance degree, the accuracy index value, the completeness index value, the timeliness index value, the consistency index value, and the relevance index value are determined respectively.

7. The method according to claim 6, characterized in that, The step of determining the accuracy index value, the completeness index value, the timeliness index value, the consistency index value, and the relevance index value based on the first quantity, the second quantity, the third quantity, the correlation degree, and the relevance degree, respectively, includes: The ratios of the first quantity to the total quantity of data in the second cleaned data, the second quantity to the total quantity, and the third quantity to the total quantity are determined respectively to obtain the accuracy index value, the integrity index value, and the timeliness index value. The consistency index value and the correlation index value are obtained by determining the ratio of the third quantity of data whose correlation degree is greater than or equal to the preset correlation degree to the total quantity, and the correlation degree corresponding to each field.

8. The method according to claim 1, characterized in that, The preprocessing of the data to be cleaned to obtain preprocessed data includes: The data to be cleaned is filtered using filtering rules; And / or, perform format conversion on the data to be cleaned according to a preset format; And / or, according to preset fields, the data to be cleaned is deduplicated.

9. The method according to claim 1, characterized in that, The method further includes: According to preset data association rules, the first value corresponding to the first field and the second value corresponding to the second field are obtained from the data to be cleaned from different data platforms; the first field and the second field are fields with a correlation relationship. If the first value and the second value are inconsistent, then one of the first value and the second value is corrected according to the preset standard, or the first value and the second value are marked respectively, and the correction result or the marking result is fed back to the data platform of the data to be cleaned.

10. A data cleaning apparatus, characterized in that, The device includes: The preprocessing module is used to acquire the data to be cleaned at the current moment and preprocess the data to be cleaned to obtain preprocessed data; the data to be cleaned is the data added at the current moment compared to the data at the previous moment. An optimization module is used to obtain the current cleaning rules and optimize the current cleaning rules based on the preprocessed data to obtain the target cleaning rules; The cleaning module is used to clean the preprocessed data based on the target cleaning rules.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1 to 9.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 9.