A data cleaning quality evaluation method and system, a storage medium and a computing device

By detecting the operating status of the data cleaning equipment in the data evaluation device, the differences in operating indicators between the first and second data cleaning processes are obtained, and the execution order and operator selection of the data cleaning process are optimized. This solves the problem of low efficiency of data cleaning equipment in the prior art and realizes efficient operation and reasonable configuration of the equipment.

CN117009335BActive Publication Date: 2025-12-30PIPECHINA SOUTH CHINA CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310984051.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-07
Publication Date
2025-12-30
Estimated Expiration
2043-08-07

AI Technical Summary

Technical Problem

In existing technologies, the evaluation process for data cleaning mainly focuses on the verification of cleaned documents, which makes it difficult to improve the operating efficiency of data cleaning equipment and optimize equipment configuration.

Method used

By detecting the operating status of the data cleaning equipment in the data evaluation device, the differences in operating indicators between the first and second data cleaning processes are obtained, the cleaning quality evaluation results are determined, and the execution order and operator selection of the data cleaning process are optimized.

Benefits of technology

This avoids the inefficiency caused by evaluating only the cleaned documents, allowing for the selection of the most suitable equipment for data cleaning, preventing equipment overload and damage, and improving the operating efficiency of the data cleaning equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117009335B_ABST
    Figure CN117009335B_ABST
Patent Text Reader

Abstract

The application discloses a data cleaning quality evaluation method and system, a storage medium and a computing device. Through application of the technical scheme of the embodiment of the application, after the data evaluation device detects that the data cleaning device finishes cleaning a to-be-processed file, the data evaluation device compares a data cleaning process implemented by itself on the file with a data cleaning process implemented by the data cleaning device on the file, determines whether the data cleaning device performs data cleaning on the file by using a suitable data cleaning process according to a comparison result, and generates a data cleaning quality evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data cleaning quality assessment method, system, storage medium, and computing device. Background Technology

[0002] Data cleaning is the process of re-examining and verifying data, with the aim of removing duplicate information, correcting existing errors, and providing data consistency.

[0003] In related technologies, data cleaning primarily involves sequentially cleaning dirty data through one or more data cleaning projects based on certain rules, thereby obtaining clean data that meets the cleaning requirements. However, the evaluation process for data cleaning in these technologies mainly focuses on verifying the cleaned files to obtain a corresponding cleaning quality assessment result. Understandably, this evaluation method is relatively simplistic and makes it difficult to improve the efficiency of the equipment performing data cleaning.

[0004] In summary, designing an evaluation method that can improve the efficiency of data cleaning has become a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a data cleaning quality assessment method, system, storage medium and computing device to address the problems existing in the prior art.

[0006] To address the aforementioned technical problems, embodiments of the present invention provide a data cleaning quality assessment method, applied to an assessment client, comprising:

[0007] After detecting that the file to be processed has already undergone the first data cleaning process in the cleaning client, a first operating indicator is obtained to reflect the running status of the first data cleaning process in the cleaning client.

[0008] The file to be processed is subjected to a second data cleaning process on the evaluation client to obtain a second operating indicator that reflects the operating status of the second data cleaning process on the evaluation client, wherein the first data cleaning process and the second data cleaning process include the same multiple data cleaning items.

[0009] Based on the degree of difference between the first operating indicator and the second operating indicator, the cleaning quality assessment result of the cleaning client for the file to be processed is determined;

[0010] The cleaning quality assessment result is used to reflect whether the cleaning client uses a data cleaning process that matches the file to be processed to clean the data.

[0011] To address the aforementioned technical problems, embodiments of the present invention also provide a data cleaning quality assessment system, applied to an assessment client, comprising:

[0012] The detection module is configured to, after detecting that the file to be processed has undergone the first data cleaning process in the cleaning client, acquire a first operating indicator that reflects the running status of the first data cleaning process in the cleaning client.

[0013] The generation module is configured to perform a second data cleaning process on the file to be processed on the evaluation client to obtain a second operating indicator that reflects the operating status of the second data cleaning process on the evaluation client, wherein the first data cleaning process and the second data cleaning process include the same multiple data cleaning items.

[0014] The determination module is configured to determine the cleaning quality assessment result of the cleaning client on the file to be processed based on the degree of difference between the first operating indicator and the second operating indicator.

[0015] The cleaning quality assessment result is used to reflect whether the cleaning client uses a data cleaning process that matches the file to be processed to clean the data.

[0016] To address the aforementioned technical problems, embodiments of the present invention also provide a computer-readable storage medium, including instructions that, when executed on a computer, cause the computer to perform the data cleaning quality assessment method provided by the above-described technical solution.

[0017] To address the aforementioned technical problems, this invention also provides a computing device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the data cleaning quality assessment method provided by the above technical solution.

[0018] In this embodiment of the invention, after detecting that the file to be processed has undergone a first data cleaning process on the cleaning client, a first operating indicator reflecting the operating status of the first data cleaning process on the cleaning client is obtained; the file to be processed is then subjected to a second data cleaning process on the evaluation client to obtain a second operating indicator reflecting the operating status of the second data cleaning process on the evaluation client, wherein the first data cleaning process and the second data cleaning process include the same multiple data cleaning items; based on the degree of difference between the first operating indicator and the second operating indicator, the cleaning quality evaluation result of the cleaning client for the file to be processed is determined; wherein the cleaning quality evaluation result is used to reflect whether the cleaning client uses a data cleaning process that matches the file to be processed to clean the file.

[0019] By applying the technical solution of the present invention, after the data evaluation device detects that the data cleaning device has finished cleaning the file to be processed, it can compare its own data cleaning process for the file with the data cleaning process implemented by the data cleaning device for the file, and determine whether the data cleaning device has used a suitable data cleaning process to clean the file based on the comparison results, and generate a data cleaning quality evaluation result accordingly.

[0020] This approach avoids the problem in some technologies where only the cleaned files are evaluated without assessing the overall cleansing process, which prevents subsequent optimization of the cleansing equipment's efficiency. Furthermore, it allows for the selection of the most suitable cleansing equipment based on the quality evaluation results of each device, preventing overloading and potential damage caused by using incompatible equipment.

[0021] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of a data cleaning quality assessment method proposed in an embodiment of the present invention;

[0023] Figure 2 This is a flowchart of a data cleaning quality assessment method proposed in an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of the data cleaning quality assessment system proposed in an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of the computing device proposed in an embodiment of the present invention. Detailed Implementation

[0026] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0027] It should be noted that various aspects of the embodiments described below are within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.

[0028] The following is combined with Figures 1-2 This document describes a data cleaning quality assessment method according to exemplary embodiments of the present invention. It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of the embodiments of the present invention, and the implementation of the embodiments of the present invention is not limited in any way. Rather, the implementation of the embodiments of the present invention can be applied to any applicable scenario.

[0029] This invention also proposes a data cleaning quality assessment method and system.

[0030] Figure 1 A schematic flowchart illustrating a data cleaning quality assessment method according to an embodiment of the present invention is shown. Figure 1 As shown, the method includes:

[0031] S101, after detecting that the file to be processed has already undergone the first data cleaning process on the cleaning client, obtain the first running indicator to reflect the running status of the first data cleaning process on the cleaning client.

[0032] S102, the file to be processed is subjected to a second data cleaning process on the evaluation client to obtain a second operating indicator that reflects the operating status of the second data cleaning process on the evaluation client, wherein the first data cleaning process and the second data cleaning process include the same multiple data cleaning items.

[0033] S103, based on the degree of difference between the first operating indicator and the second operating indicator, determine the cleaning quality assessment result of the file to be processed by the cleaning client; wherein, the cleaning quality assessment result is used to reflect whether the cleaning client uses a data cleaning process that matches the file to be processed to clean the file.

[0034] In related technologies, data cleaning mainly involves sequentially cleaning dirty data through one or more data cleaning projects formulated according to certain rules, thereby obtaining clean data that meets the cleaning requirements.

[0035] Data cleaning projects can include data preprocessing, data transformation, data integration, data validation, data deduplication, handling missing values, handling outliers, data normalization and standardization, data segmentation, and so on.

[0036] Data preprocessing refers to operations such as formatting, normalizing, deduplicating, removing noise and outliers from the collected raw data to ensure data quality and accuracy.

[0037] Data transformation refers to converting raw data into an analyzable data format, typically using structured data formats such as CSV, JSON, or XML. This step involves defining fields, converting data types, and changing encodings to facilitate subsequent data analysis and mining.

[0038] Data integration refers to combining datasets from multiple data sources into a single dataset. This step involves identifying and selecting data sources, and performing operations such as data extraction, cleaning, transformation, and loading. Furthermore, it's crucial to address data duplication and conflicts during data integration.

[0039] Data validation refers to verifying the data after initial cleaning to ensure its quality and integrity. This step requires the use of various techniques and algorithms, such as statistical analysis, logical verification, rule checking, data comparison, and visualization, to promptly identify and correct data anomalies and errors.

[0040] Data deduplication refers to the removal of duplicate data. Methods for data deduplication include identifier-based deduplication, text content-based deduplication, and timestamp-based deduplication. Understandably, deduplication reduces interference and noise, improving the reliability and efficiency of data analysis.

[0041] Data missing value handling refers to the process by which missing values ​​in data affect subsequent calculations and analysis. Methods for handling missing values ​​include deleting rows / columns containing missing values, filling missing values ​​with the mean, median, or mode, and using regression or clustering models to predict missing values. The appropriate method should be selected based on the specific missing value and the data type.

[0042] Outlier handling refers to data points that significantly differ from the majority of the data, possibly due to measurement or recording errors. Outlier handling methods include deleting outliers, replacing them with the mean, replacing them with the median, and treating outliers as new categories. The appropriate method should be chosen based on the specific circumstances to avoid excessive interference with the data.

[0043] However, current data cleaning assessment methods primarily focus on verifying the cleaned files to obtain a quality evaluation result. Understandably, this approach is rather simplistic and makes it difficult to improve the efficiency of the equipment used for data cleaning.

[0044] Data partitioning refers to dividing a dataset into multiple subsets according to certain rules to facilitate analysis and processing.

[0045] Data normalization and standardization refer to standardizing data to concentrate it within a specific range, thus avoiding the impact of differences between dependent variables on the results. Normalization and standardization are two different standardization methods, and the appropriate method should be selected based on the data type and business requirements.

[0046] Furthermore, the evaluation process for data cleaning in related technologies mainly focuses on verifying the cleaned files to obtain corresponding cleaning quality assessment results. Understandably, this evaluation method is relatively simplistic and makes it difficult to improve the efficiency of the equipment performing the data cleaning.

[0047] Based on this, this invention proposes a data cleaning quality assessment method, which can compare its own data cleaning process for the file with the data cleaning process implemented by the data cleaning device after the data evaluation device detects that the data cleaning device has finished cleaning the file, and determine whether the data cleaning device has used an appropriate data cleaning process to clean the file based on the comparison results, and generate a data cleaning quality assessment result accordingly.

[0048] It is understandable that the technical solution of this invention can, on the one hand, avoid the problem in related technologies where only the cleaned files are evaluated without evaluating the operational status of the data cleaning process, which prevents subsequent optimization of the data cleaning equipment's operating efficiency. On the other hand, it can also select the most suitable equipment for data cleaning based on the cleaning quality evaluation results of each data cleaning device, thereby avoiding the problem of using mismatched equipment for data cleaning, which could lead to overloading the cleaning equipment and potentially causing equipment damage.

[0049] Furthermore, the embodiments of the present invention are combined herein. Figure 2 The plan will be explained in detail:

[0050] Step 1: After detecting that the file to be processed has already undergone the first data cleaning process in the cleaning client, obtain the first runtime of the first data cleaning process, the first execution order of each data cleaning item executed by the cleaning client during the first data cleaning process, and the first running operator of the first data cleaning process in the cleaning client.

[0051] The first runtime is the total time spent by the cleaning client in running the first data cleaning process.

[0052] In one approach, the first execution order is the order in which the cleaning client executes each data cleaning project.

[0053] Understandably, this is because data cleaning projects involve multiple steps, such as data preprocessing, data transformation, data integration, data validation, data deduplication, handling missing values, handling outliers, data normalization and standardization, data merging and splitting, etc.

[0054] Understandably, for the same file to be processed, executing cleaning tasks in different orders may affect the efficiency of the data cleaning process.

[0055] For example, when processing image files, which often contain many repeated images, placing the data deduplication step before the missing value handling and outlier handling steps can significantly reduce the complexity of subsequent missing value handling (and outlier handling) tasks. This, in turn, improves efficiency and reduces runtime.

[0056] For example, when the files to be processed are text files, they typically have data sources that are quite scattered. Therefore, placing the data integration project before other processing projects can avoid problems such as erroneous deduplication and normalization of data with the same name but from different sources. This reduces the need to restart the data cleaning process due to incomplete data, thus avoiding unnecessary increases in runtime.

[0057] For example, when processing audio files, which are easily affected by noise, they often contain a large number of outliers. Therefore, placing outlier handling before other processing tasks can reduce the unnecessary waste of runtime resources caused by frequently processing noisy or other non-business data.

[0058] Therefore, the execution order of different data cleaning projects can lead to different data cleaning efficiencies. Based on this, this embodiment of the invention needs to obtain the first execution order of each data cleaning project executed by the cleaning client during the first data cleaning process, so that it can be compared with the execution order generated by the evaluation client to determine whether the project execution order needs to be improved in subsequent cleaning processes.

[0059] In another approach, the first running operator is the operator unit in the cleaning client that supports the execution of the first data cleaning process.

[0060] In this process, a client may deploy different operator units, and these different operator units may have varying capabilities in supporting the data cleaning process (the differences mainly lie in computational efficiency). Therefore, how to select the most suitable operator unit for the data cleaning process on this cleaning client to achieve efficient data cleaning operation has become a problem that this invention aims to solve.

[0061] Understandably, for real-world data cleaning scenarios, the best way to ensure the efficiency of the data cleaning process is not to configure it with the most powerful operators. Instead, it's necessary to select operators that can best support the data cleaning process for business processing.

[0062] In other words, even if an operator unit has high computational power, if the operator is incompatible with the application support coefficients of the data cleaning process, it will not be able to exert its maximum computational power. This will naturally affect the operating efficiency of the data cleaning process.

[0063] Therefore, using different operator units to run the first data cleaning process may result in different data cleaning efficiencies. Based on this, this embodiment of the invention needs to obtain the first running operator of the cleaning client during the first data cleaning process, so that it can be compared with the second running operator of the evaluation client to determine whether the selection of running operators needs to be improved in subsequent cleaning processes.

[0064] Step 2: Use the first runtime, the first execution order, and the first execution operator as the first runtime index.

[0065] Step 3 selects a second execution order and a second running operator for the file to be processed, and uses the second execution order and the second running operator to perform a second data cleaning process on the file to be processed.

[0066] The second execution order reflects the standard order in which the evaluation client executes each data cleaning project, and the second running operator reflects the operator used by the evaluation client to run the second data cleaning process; furthermore, the first execution order is different from the second execution order, and / or the first running operator is different from the second running operator.

[0067] Understandably, in this embodiment of the invention, it is necessary to evaluate whether the client also uses the second execution order and the second running operator to run the second data cleaning process on the file to be processed.

[0068] It should be noted that the first and second data cleaning processes must include the same data cleaning items, but the execution order of the data cleaning items may be different.

[0069] For example, the first data cleaning process includes three data cleaning items in sequence: data preprocessing, data transformation, and data integration (i.e., the first execution order). The second data cleaning process also needs to perform these three data cleaning items, but in the order of data transformation, data integration, and data preprocessing (i.e., the second execution order).

[0070] Similarly, the second execution operator may also be different from the first execution operator. When they are different, it can be manifested in different ways, such as different numbers of execution operators, different execution logic of execution operators, different execution languages ​​of execution operators, different models of execution operators, etc.

[0071] Understandably, in this embodiment of the invention, the first execution order and the second execution order need to be different, and / or the first running operator and the second running operator need to be different (that is, at least one of the two needs to be different in terms of execution order and running operator). Only in this way can the evaluation client perform data cleaning work on the file to be processed in a running state that is different from the first data cleaning process, and then determine which data cleaning process is more suitable for cleaning the file to be processed based on the comparison between the second running index and the first running index.

[0072] Step 4: Obtain the second runtime of the second data cleaning process, and use the second runtime, second execution order, and second execution operator as the second runtime metric. Then proceed to either step 5a or step 5b.

[0073] Step 5a: If the first runtime is detected to be longer than the second runtime, determine that the cleaning client's cleaning quality assessment result for the file to be processed is unqualified.

[0074] In one embodiment of the present invention, the runtime of two data cleaning processes can be used as an evaluation index for whether the first data cleaning process is qualified.

[0075] Understandably, a shorter second runtime indicates that the second data cleaning process has achieved data cleaning of the file to be processed with higher efficiency. Therefore, this embodiment of the invention can determine that the cleaning client's assessment result of the cleaning quality of the file to be processed is unqualified (i.e., it is determined that the cleaning client did not use a data cleaning process that matches the file to be processed).

[0076] Step 6a: Feed back the cleaning quality assessment results, which include the second operating indicators, to the cleaning client so that the cleaning client can subsequently use the second execution order and the second operating operator to clean the data of the file to be processed.

[0077] In one approach, in order to optimize the operating efficiency of the data cleaning equipment, this embodiment of the invention needs to feed back a second operating indicator to the cleaning client, so that the client knows that using a second execution order and a second operating operator to perform the data cleaning process on the file to be processed is more efficient.

[0078] Therefore, when the cleaning client subsequently obtains other files of the same type as the file to be processed, it can perform data cleaning on those other files using a more optimized data cleaning process. This achieves a technical effect that enables the client to learn better data cleaning processes on its own.

[0079] Step 5b: If the first runtime is less than or equal to the second runtime, the cleaning client is deemed to have passed the cleaning quality assessment of the files to be processed.

[0080] Understandably, a shorter first run time indicates that the first data cleaning process has achieved data cleaning of the files to be processed with higher efficiency.

[0081] Therefore, embodiments of the present invention can determine that the cleaning client's cleaning quality assessment result for the file to be processed is qualified (i.e., it is determined that the cleaning client has used a data cleaning process that matches the file to be processed), and mark it. This allows other files of the same type as the file to be processed to be directly assigned to the cleaning client in the future, without the need for further comparison of operating metrics.

[0082] By applying the technical solution of the present invention, after the data evaluation device detects that the data cleaning device has finished cleaning the file to be processed, it can compare its own data cleaning process for the file with the data cleaning process implemented by the data cleaning device for the file, and determine whether the data cleaning device has used a suitable data cleaning process to clean the file based on the comparison results, and generate a data cleaning quality evaluation result accordingly.

[0083] This approach avoids the problem in some technologies where only the cleaned files are evaluated without assessing the overall cleansing process, which prevents subsequent optimization of the cleansing equipment's efficiency. Furthermore, it allows for the selection of the most suitable cleansing equipment based on the quality evaluation results of each device, preventing overloading and potential damage caused by using incompatible equipment.

[0084] Optionally, in another embodiment of the method described above based on the present invention, obtaining a first operating indicator reflecting the operating status of the first data cleaning process on the cleaning client includes:

[0085] The first runtime of the first data cleaning process is obtained; the first execution order of each data cleaning item is executed by the cleaning client during the execution of the first data cleaning process; and the first running operator of the first data cleaning process is run in the cleaning client.

[0086] The first runtime, the first execution order, and the first execution operator are used as the first runtime metric.

[0087] Optionally, in another embodiment of the method based on the present invention, after taking the first runtime, the first execution order, and the first execution operator as the first runtime indicator, the method further includes:

[0088] A second execution order and a second running operator are selected for the file to be processed, and the second data cleaning process is performed on the file to be processed using the second execution order and the second running operator;

[0089] Wherein, the second execution order is used to reflect the standard order in which the evaluation client executes each data cleaning project, and the second running operator is used to reflect the operator of the evaluation client running the second data cleaning process; and the first execution order is different from the second execution order, and / or the first running operator is different from the second running operator.

[0090] Optionally, in another embodiment of the method based on the present invention, after performing the second data cleaning process on the file to be processed using the second execution order and the second running operator, the method further includes:

[0091] Obtain the second runtime of the second data cleaning process;

[0092] The second runtime, the second execution order, and the second execution operator are used as the second runtime indicators. Based on the degree of difference between the first runtime indicator and the second runtime indicator, the cleaning client determines the cleaning quality assessment result of the file to be processed.

[0093] Optionally, in another embodiment of the method described above based on the present invention, determining the cleaning quality assessment result of the cleaning client for the file to be processed based on the degree of difference between the first operating indicator and the second operating indicator includes:

[0094] If the first runtime is detected to be greater than the second runtime, it is determined that the cleaning client's cleaning quality assessment result for the file to be processed is unqualified.

[0095] If the first runtime is detected to be less than or equal to the second runtime, it is determined that the cleaning client's cleaning quality assessment result for the file to be processed is unqualified.

[0096] Optionally, in another embodiment of the method based on the present invention, after determining that the cleaning quality assessment result of the cleaning client for the file to be processed is unqualified, the method further includes:

[0097] The cleaning quality assessment result, which includes the second operating indicator, is fed back to the cleaning client, so that the cleaning client can subsequently use the second execution order and the second operating operator to perform data cleaning on other files of the same type as the file to be processed.

[0098] Optionally, in another embodiment of the present invention, such as Figure 3 As shown in the figure, this embodiment of the invention also provides a data cleaning quality assessment system. It includes:

[0099] The detection module 201 is configured to, after detecting that the file to be processed has undergone the first data cleaning process in the cleaning client, acquire a first operating indicator that reflects the running status of the first data cleaning process in the cleaning client.

[0100] The generation module 202 is configured to perform a second data cleaning process on the file to be processed on the evaluation client to obtain a second operating indicator that reflects the operating status of the second data cleaning process on the evaluation client, wherein the first data cleaning process and the second data cleaning process include the same multiple data cleaning items.

[0101] The determination module 203 is configured to determine the cleaning quality assessment result of the cleaning client on the file to be processed based on the degree of difference between the first operating indicator and the second operating indicator.

[0102] The cleaning quality assessment result is used to reflect whether the cleaning client uses a data cleaning process that matches the file to be processed to clean the data.

[0103] By applying the technical solution of the present invention, after the data evaluation device detects that the data cleaning device has finished cleaning the file to be processed, it can compare its own data cleaning process for the file with the data cleaning process implemented by the data cleaning device for the file, and determine whether the data cleaning device has used a suitable data cleaning process to clean the file based on the comparison results, and generate a data cleaning quality evaluation result accordingly.

[0104] This invention also provides a computing device for performing the above-described data cleaning quality assessment method. Please refer to... Figure 4 This illustrates a schematic diagram of a computing device provided by some embodiments of the present invention. For example... Figure 4 As shown, the computing device 3 includes: a processor 300, a memory 301, a bus 302, and a communication interface 303. The processor 300, the communication interface 303, and the memory 301 are connected through the bus 302. The memory 301 stores a computer program that can run on the processor 300. When the processor 300 runs the computer program, it executes the data cleaning quality assessment method provided in any of the foregoing embodiments of the present invention.

[0105] The memory 301 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this device network element and at least one other network element is achieved through at least one communication interface 303 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.

[0106] Bus 302 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory 301 is used to store programs. After receiving an execution instruction, the processor 300 executes the program. The data cleaning quality assessment method disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 300, or implemented by the processor 300.

[0107] The processor 300 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 300 or by instructions in software form. The processor 300 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 301. The processor 300 reads the information in memory 301 and, in conjunction with its hardware, completes the steps of the above method.

[0108] The computing device provided in this embodiment of the invention and the data cleaning quality assessment method provided in this embodiment of the invention are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.

[0109] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0110] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0111] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0112] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0113] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A data cleaning quality assessment method, characterized in that, The application is applied to evaluate a client, comprising: After detecting that a to-be-processed file has been processed by a cleaning client in a first data cleaning process, obtaining a first running index reflecting a running state of the first data cleaning process in the cleaning client; The first running index reflecting the running state of the first data cleaning process in the cleaning client comprises: obtaining a first running duration of the first data cleaning process, a first execution sequence of the cleaning client in the process of running the first data cleaning process, and a first running operator of the cleaning client running the first data cleaning process; and taking the first running duration, the first execution sequence and the first running operator as the first running index; After taking the first running duration, the first execution sequence and the first running operator as the first running index, further comprising: selecting a second execution sequence and a second running operator for the to-be-processed file, and using the second execution sequence and the second running operator to process the to-be-processed file in a second data cleaning process in the evaluation client to obtain a second running index reflecting a running state of the second data cleaning process in the evaluation client; wherein the second execution sequence reflects a standard sequence of the evaluation client executing each data cleaning item, and the second running operator reflects an operator of the evaluation client running the second data cleaning process; and the first execution sequence and the second execution sequence are different, and / or the first running operator and the second running operator are different; the first data cleaning process and the second data cleaning process comprise the same plurality of data cleaning items; Based on the difference between the first running index and the second running index, determine the cleaning quality evaluation result of the cleaning client for the to-be-processed file; The cleaning quality evaluation result reflects whether the cleaning client uses a data cleaning process matched with the to-be-processed file to clean the to-be-processed file.

2. The method of claim 1, wherein, After using the second execution sequence and the second running operator to process the to-be-processed file in the second data cleaning process, further comprising: Obtain a second running duration of the second data cleaning process; Take the second running duration, the second execution sequence and the second running operator as the second running index, and determine the cleaning quality evaluation result of the cleaning client for the to-be-processed file based on the difference between the first running index and the second running index.

3. The method of claim 2, wherein, The determination of the cleaning quality evaluation result of the cleaning client for the to-be-processed file based on the difference between the first running index and the second running index comprises: If the first running duration is greater than the second running duration, determine that the cleaning quality evaluation result of the cleaning client for the to-be-processed file is unqualified; If it is detected that the first running time is less than or equal to the second running time, it is determined that the cleaning quality evaluation result of the cleaning client on the to-be-processed file is qualified.

4. The method of claim 3, wherein, After the determination that the cleaning quality evaluation result of the cleaning client on the to-be-processed file is unqualified, the method further includes: The cleaning quality evaluation result containing the second running index is fed back to the cleaning client, so that the cleaning client subsequently adopts the second execution sequence and the second running operator to perform data cleaning on other files of the same type as the to-be-processed file.

5. A data cleansing quality assessment system, characterized by, The application is applied to an evaluation client, and includes: A detection module is configured to detect that a to-be-processed file has been subjected to a first data cleaning process in a cleaning client, and obtain a first running index reflecting a running state of the first data cleaning process in the cleaning client; The first running index reflecting the running state of the first data cleaning process in the cleaning client includes: a first running time of the first data cleaning process, a first execution sequence of the cleaning client in executing each data cleaning item in the process of running the first data cleaning process, and a first running operator of the cleaning client in running the first data cleaning process; and the first running time, the first execution sequence, and the first running operator are taken as the first running index; A generation module is configured to, after taking the first running time, the first execution sequence, and the first running operator as the first running index, select a second execution sequence and a second running operator for the to-be-processed file, and perform a second data cleaning process on the to-be-processed file in the evaluation client by using the second execution sequence and the second running operator, to obtain a second running index reflecting a running state of the second data cleaning process in the evaluation client; the second execution sequence reflects a standard sequence of the evaluation client in executing each data cleaning item, and the second running operator reflects an operator of the evaluation client in running the second data cleaning process; the first execution sequence is different from the second execution sequence, and / or the first running operator is different from the second running operator; and the first data cleaning process and the second data cleaning process include the same plurality of data cleaning items; A determination module is configured to determine a cleaning quality evaluation result of the cleaning client on the to-be-processed file based on a difference degree between the first running index and the second running index. The cleaning quality evaluation result reflects whether the cleaning client adopts a data cleaning process matched with the to-be-processed file to perform data cleaning on the to-be-processed file.

6. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are run on a computer, the computer is caused to perform the data cleaning quality evaluation method according to any one of claims 1 to 4.

7. A computing device, comprising: The application includes: Computer program product comprising a memory, a processor and a computer program stored on the memory and loadable into the processor, characterized in that the processor implements the method for assessing the quality of data cleaning according to any one of claims 1 to 4 when executing the program.

Citation Information

Patent Citations

  • Data quality improvement method and system based on distributed system

    CN110737640A

  • Data quality evaluation management method and device

    CN111427974A