Data cleaning method and server
By automatically cleaning the data to be cleaned on the server using preset cleaning operators and data cleaning models, the problems of low efficiency and high cost in the prior art are solved, and more efficient and accurate data cleaning is achieved.
Patent Information
- Application Number
- CN202510059355.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively solve the data accuracy problem, resulting in low data cleaning efficiency and high cost.
By applying the data cleaning method on the server, the preset cleaning operator and data cleaning model are used to automatically clean the cleaning data, identify and process different types of problem data.
Improves the accuracy and efficiency of data cleaning, reduces the workload of manual cleaning, reduces costs, and improves data availability.
Smart Images

Figure CN120067081A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a data cleaning method and a server. Background Art
[0002] With the rapid development of information technology and the advent of the big data era, all walks of life have begun to establish information systems and accumulate a large amount of data. The accuracy of data is the basic condition for various data analyses. However, in reality, due to various reasons in the processes of data collection, transmission, storage, and processing, the problem of data accuracy is widespread. The purpose of data cleaning is to detect incorrect data in the data, eliminate or correct the incorrect data, so as to improve the accuracy and quality of the data. Therefore, a data cleaning method is needed. Summary of the Invention
[0003] To solve the above problems, embodiments of this application provide a data cleaning method and a server, which improve the accuracy of data cleaning.
[0004] In a first aspect, an embodiment of this application provides a data cleaning method applied to a server, including: obtaining data to be cleaned; identifying, according to at least one data quality rule, the data to be cleaned to obtain at least one problematic data, where the data quality rule is used to represent a rule that conforms to the problematic data; in the case where the problematic data is a first type of problematic data, using a preset cleaning operator to clean the first type of problematic data to obtain a first cleaning result, where the first type of problematic data represents problematic data that satisfies cleaning using the preset cleaning operator; in the case where the problematic data is a second type of problematic data, using a data cleaning model to clean the second type of problematic data to obtain a second cleaning result, where the second type of problematic data represents problematic data that satisfies cleaning using the data cleaning model.
[0005] In the embodiments of this application, a data cleaning method is provided. Through the preset cleaning operator and the data cleaning model, the automatic modification and cleaning of problematic data are realized, the workload of manually cleaning data in the data cleaning process is reduced, the efficiency of data cleaning is improved, and the labor cost is saved. Additionally, the method of the embodiments of this application can clean and modify problematic data more thoroughly, the proportion of modified and corrected data is higher, the possibility of data discarding is lower, the possibility of important data loss is reduced, and the data availability is ensured. This is of extremely important significance in scenarios where the data is of high importance, such as government affairs scenarios. By using the preset cleaning operator and the data cleaning model to perform data cleaning operations on different types of data respectively, the data is classified and cleaned, and the cleaning efficiency is higher. And the amount of data that is not cleaned is lower. In scenarios where the amount of data that is not cleaned needs to be manually modified by the user, the amount of data that needs to be manually modified is lower, the workload of manual modification is reduced, and the user experience is improved.
[0006] In a possible implementation, a data quality rule represents a quality rule of at least one rule type for at least one dimension; the at least one dimension includes at least one of integrity, uniqueness, timeliness, validity, accuracy, and consistency, and the at least one rule type includes at least one of a regular expression, an enumerated value, a range value, a Structured Query Language (SQL) expression, and an SQL statement.
[0007] In this implementation, an exemplary description of the specific implementation of the data quality rule is given. It should be noted that the content of this implementation does not constitute a limitation on the preset data quality rules in the embodiments of the present application.
[0008] In a possible implementation, identifying the data to be cleaned according to at least one data quality rule to obtain at least one problem data includes: judging the integrity of the data to be cleaned according to the quality rule of the integrity dimension to obtain an integrity identification result; judging whether the data to be cleaned is unique according to the quality rule of the uniqueness dimension to obtain a uniqueness identification result; judging whether the update of the data to be cleaned meets the timeliness requirement according to the quality rule of the timeliness dimension to obtain a timeliness identification result; judging whether the data to be cleaned meets the validity requirement according to the quality rule of the validity dimension to obtain a validity identification result; judging whether the data to be cleaned is accurate according to the quality rule of the accuracy dimension to obtain an accuracy identification result; judging whether the data to be cleaned meets the consistency requirement according to the quality rule of the consistency dimension to obtain a consistency identification result; judging whether the data to be cleaned is problem data according to at least one of the integrity identification result, the uniqueness identification result, the timeliness identification result, the validity identification result, the accuracy identification result, and the consistency identification result.
[0009] In this implementation, the implementation of identifying problem data through different dimensions is described. Identifying and judging the data to be cleaned from different dimensions ensures the accuracy of the problem data and reduces the possibility of missing problem data identification.
[0010] In a possible implementation manner, the above-mentioned identification of the data to be cleaned according to at least one data quality rule to obtain at least one problematic data includes: judging whether the data to be cleaned meets the requirements of the regular expression according to the regular expression to obtain a regular expression discrimination result; judging whether the data to be cleaned is within the range of the enumerated value according to the enumerated value to obtain an enumerated value discrimination result; judging whether the data to be cleaned is within the range represented by the range value according to the range value to obtain a range value discrimination result; judging whether the data to be cleaned meets the requirements of the structured query language expression according to the structured query language expression to obtain a structured query language expression discrimination result; judging whether the data to be cleaned meets the requirements of the structured query language statement according to the structured query language statement to obtain a structured query language statement discrimination result; judging whether the data to be cleaned is problematic data according to at least one of the regular expression discrimination result, the enumerated value discrimination result, the range value discrimination result, the structured query language expression discrimination result, and the structured query language statement discrimination result.
[0011] In this implementation manner, the implementation manner of identifying problematic data through different rule types is described. Identifying and judging the data to be cleaned from different rule types improves the identification accuracy of the data to be cleaned of different data types and improves the applicable data types of the data quality rules. It should be noted that the above rule types are only exemplary descriptions, and in actual scenarios, corresponding rule types can also be formulated according to the data types of the data.
[0012] In a possible implementation manner, the above-mentioned preset cleaning operators include at least one of a duplicate data operator, a normalization operator, and a data type operator.
[0013] In this implementation manner, examples of the preset cleaning operators are given. It should be noted that the content of this implementation manner does not constitute a limitation on the preset cleaning operators in the embodiments of the present application. The preset cleaning operators include multiple cleaning operators for cleaning different categories of data. One cleaning operator only processes one category of problematic data. In this way, each cleaning operator is easier to implement. By classifying the problematic data, the user can understand the problematic data under different categories and the category to which the current problematic data belongs, improving the user experience.
[0014] In a possible implementation, when the problem data is of the first type of problem data, the preset cleaning operator is used to clean the first type of problem data to obtain a first cleaning result, including: when the problem data is of the first type of problem data, determining whether the first type of problem data is duplicate data; when the first type of problem data is duplicate data, using the duplicate data operator to retain one of the multiple first type of problem data that are duplicates to obtain the cleaning result of the duplicate data operator; determining whether the first type of problem data meets the normative requirements; when the first type of problem data does not meet the normative requirements, using the normative operator to convert the format of the first type of problem data into a canonical format to obtain the cleaning result of the normative operator; determining whether the data type of the first type of problem data meets the requirements; when the data type of the first type of problem data does not meet the requirements, using the data type operator to convert the data type of the first type of problem data into a data type that meets the requirements to obtain the cleaning result of the data type operator; obtaining the first cleaning result based on at least one of the cleaning results of the duplicate data operator, the cleaning result of the normative operator, and the cleaning result of the data type operator.
[0015] In this implementation, the data cleaning processes of the duplicate data operator, the normative operator, and the data type operator are described. Different cleaning operators can handle different categories of problem data, thereby enabling targeted processing of the problem data and improving the accuracy of data cleaning.
[0016] In a possible implementation, the above-mentioned second type of problem data includes at least one of incomplete data and invalid data. When the problem data is of the second type of problem data, the data cleaning model is used to clean the second type of problem data to obtain a second cleaning result, including: inputting the incomplete data into the data cleaning model, and the data cleaning model modifies the incomplete data into complete data; inputting the invalid data into the data cleaning model, and the data cleaning model modifies the invalid data into valid data; obtaining the second cleaning result based on the modified complete data and valid data.
[0017] In this implementation, the above-mentioned second type of problem data includes incomplete data, invalid data, etc. The data cleaning model is a neural network model, which can well analyze problem data after training, expanding the cleaning scope of problem data, enabling more problem data to be cleaned, and improving the cleaning effect of problem data.
[0018] In a possible implementation, the above data cleaning model is a large language model. When the problem data is the second type of problem data, the data cleaning model is used to clean the second type of problem data to obtain a second cleaning result, including: inputting the second type of problem data into the large language model to output the second cleaning result.
[0019] In this implementation, the above data cleaning model can be a large language model. The large language model has high text understanding ability. The large language model can comprehensively analyze the second type of problem data and then modify the problem data into correct data.
[0020] In a possible implementation, the method further includes: when the problem data is the first type and the second type of problem data, using a preset cleaning operator and the data cleaning model for cleaning to obtain a third cleaning result.
[0021] In this implementation, the cleaning method for problem data that meets both the first type and the second type is described. For the first type and the second type of problem data, a preset cleaning operator and the data cleaning model can be used for cleaning. In this way, the problems of the problem data in the first type and the second type can be solved, and the data cleaning effect is improved.
[0022] In a possible implementation, the method further includes: when the problem data is the third type of problem data, outputting the third type of problem data so that the third type of problem data is manually cleaned to obtain a fourth cleaning result. The third type of problem data represents problem data that meets the manual cleaning conditions.
[0023] In this implementation, a method for manually cleaning data is provided. The third type can be considered a type of data that is more difficult to clean than the first type and the second type. Manually cleaning data can be used as a guarantee means in the data cleaning process, and thus it can ensure that all data is cleaned, guaranteeing the integrity and usability of the data. This is of extremely important significance in scenarios where data is highly important, such as in government affairs scenarios.
[0024] In a possible implementation, the method further includes: obtaining the fourth cleaning result; training the data cleaning model according to the fourth cleaning result.
[0025] In this implementation, a method for training the data cleaning model is provided. Training the data cleaning model according to the fourth cleaning result enables the data cleaning model to clean the third type of data, expanding the cleaning scope of the data cleaning model and improving the data cleaning ability of the data cleaning model. It also further reduces the amount of data to be manually cleaned and improves the user experience.
[0026] In a possible implementation, the method further includes: when the problem data is the problem data of the first type, the second type, and the third type, using a preset cleaning operator and a data cleaning model for cleaning to obtain a fifth cleaning result, and outputting the fifth cleaning result so that the data in the fifth cleaning result is manually cleaned to obtain a sixth cleaning result.
[0027] In this implementation, the cleaning method for the problem data that simultaneously meets the first type, the second type, and the third type is described. Combining the first cleaning with the preset cleaning operator and the second cleaning with the data cleaning model, the manual cleaning can be understood as the third cleaning, realizing the data cleaning in three stages of "preset cleaning operator + data cleaning model + manual cleaning". The manual cleaning ensures that all data is corrected, guaranteeing the integrity and availability of the data, which is of extremely important significance in scenarios with high data importance such as government affairs scenarios. Additionally, through the data cleaning operations in two stages of the preset cleaning operator and the data cleaning model, the amount of data that needs to be manually modified is lower, reducing the manual modification workload and improving the user experience.
[0028] In a second aspect, an embodiment of the present application provides a data cleaning device, including: an acquisition module for acquiring data to be cleaned; a processing module for identifying the data to be cleaned according to at least one data quality rule to obtain at least one problem data, where the data quality rule is used to represent the rule that conforms to the problem data; and, when the problem data is the problem data of the first type, using a preset cleaning operator to clean the problem data of the first type to obtain a first cleaning result, where the problem data of the first type represents the problem data that satisfies the cleaning using the preset cleaning operator; and, when the problem data is the problem data of the second type, using a data cleaning model to clean the problem data of the second type to obtain a second cleaning result, where the problem data of the second type represents the problem data that satisfies the cleaning using the data cleaning model.
[0029] In a possible implementation, the data quality rule represents the quality rules of at least one rule type in at least one dimension; the at least one dimension includes at least one of integrity, uniqueness, timeliness, validity, accuracy, and consistency, and the at least one rule type includes at least one of regular expression, enumerated value, range value, structured query language expression, and structured query language statement.
[0030] In a possible implementation manner, the above-mentioned processing module is specifically configured to: determine the integrity of the data to be cleaned according to the quality rules of the integrity dimension, and obtain an integrity recognition result; determine whether the data to be cleaned is unique according to the quality rules of the uniqueness dimension, and obtain a uniqueness recognition result; determine whether the update of the data to be cleaned meets the timeliness requirement according to the quality rules of the timeliness dimension, and obtain a timeliness recognition result; determine whether the data to be cleaned meets the validity requirement according to the quality rules of the validity dimension, and obtain a validity recognition result; determine whether the data to be cleaned is accurate according to the quality rules of the accuracy dimension, and obtain an accuracy recognition result; determine whether the data to be cleaned meets the consistency requirement according to the quality rules of the consistency dimension, and obtain a consistency recognition result; determine whether the data to be cleaned is problematic data according to at least one of the integrity recognition result, the uniqueness recognition result, the timeliness recognition result, the validity recognition result, the accuracy recognition result, and the consistency recognition result.
[0031] In a possible implementation manner, the above-mentioned processing module is specifically configured to: determine whether the data to be cleaned meets the requirements of the regular expression according to the regular expression, and obtain a regular expression discrimination result; determine whether the data to be cleaned is within the range of the enumerated value according to the enumerated value, and obtain an enumerated value discrimination result; determine whether the data to be cleaned is within the range represented by the range value according to the range value, and obtain a range value discrimination result; determine whether the data to be cleaned meets the requirements of the structured query language expression according to the structured query language expression, and obtain a structured query language expression discrimination result; determine whether the data to be cleaned meets the requirements of the structured query language statement according to the structured query language statement, and obtain a structured query language statement discrimination result; determine whether the data to be cleaned is problematic data according to at least one of the regular expression discrimination result, the enumerated value discrimination result, the range value discrimination result, the structured query language expression discrimination result, and the structured query language statement discrimination result.
[0032] In a possible implementation manner, the above-mentioned preset cleaning operators include at least one of a duplicate data operator, a normalization operator, and a data type operator.
[0033] In a possible implementation, the above-mentioned processing module is specifically configured to: when the problem data is the problem data of the first type, determine whether the problem data of the first type is duplicate data; when the problem data of the first type is duplicate data, use the duplicate data operator to retain one of the multiple duplicate problem data of the first type to obtain the cleaning result of the duplicate data operator; determine whether the problem data of the first type meets the normalization requirements; when the problem data of the first type does not meet the normalization requirements, use the normalization operator to convert the format of the problem data of the first type into the normalized format to obtain the cleaning result of the normalization operator; determine whether the data type of the problem data of the first type meets the requirements; when the data type of the problem data of the first type does not meet the requirements, use the data type operator to convert the data type of the problem data of the first type into a data type that meets the requirements to obtain the cleaning result of the data type operator; obtain the first cleaning result according to at least one of the cleaning results of the duplicate data operator, the cleaning result of the normalization operator, and the cleaning result of the data type operator.
[0034] In a possible implementation, the above-mentioned problem data of the second type includes at least one of incomplete data and invalid data. The above-mentioned processing module is specifically configured to: input the incomplete data into the data cleaning model, and the data cleaning model modifies the incomplete data into complete data; input the invalid data into the data cleaning model, and the data cleaning model modifies the invalid data into valid data; obtain the second cleaning result according to the modified complete data and valid data.
[0035] In a possible implementation, the above-mentioned data cleaning model is a large language model. The above-mentioned processing module is specifically configured to: input the problem data of the second type into the large language model to output the second cleaning result.
[0036] In a possible implementation, the above-mentioned processing module is further configured to: when the problem data is the problem data of the first type and the second type, perform cleaning using the preset cleaning operator and the data cleaning model to obtain the third cleaning result.
[0037] In a possible implementation, the above-mentioned processing module is further configured to: when the problem data is the problem data of the third type, output the problem data of the third type so that the problem data of the third type is manually cleaned to obtain the fourth cleaning result, and the problem data of the third type represents the problem data that meets the manual cleaning conditions.
[0038] In a possible implementation, the above-mentioned processing module is further configured to: obtain the fourth cleaning result; train the data cleaning model according to the fourth cleaning result.
[0039] In a possible implementation, the above-mentioned processing module is further configured to: when the problem data is of the first type, the second type, and the third type of problem data, perform cleaning using a preset cleaning operator and a data cleaning model to obtain a fifth cleaning result, and output the fifth cleaning result so that the data in the fifth cleaning result is manually cleaned to obtain a sixth cleaning result.
[0040] In a third aspect, an embodiment of the present application provides a server, which includes a processor and a memory. The processor is configured to execute instructions stored in the memory, so that the server is configured to execute the method described in the first aspect and any possible implementation thereof above.
[0041] In a fourth aspect, the present application provides a terminal device, which is communicatively connected to the server provided in the third aspect above. The terminal device is configured to obtain information about a user's data processing task and its parameter configuration. The data processing task includes a data cleaning task, so that the server provided in the third aspect above executes the method described in the first aspect and any possible implementation thereof above.
[0042] In a fifth aspect, an embodiment of the present application provides a data cleaning system, which includes the server provided in the third aspect above and the terminal device provided in the fourth aspect above.
[0043] In a sixth aspect, an embodiment of the present application provides a chip system, which includes a processor and a power supply circuit. The power supply circuit is configured to supply power to the processor, and the processor is configured to execute the method described in the first aspect and any possible implementation thereof above.
[0044] In a seventh aspect, an embodiment of the present application provides a computing device, which includes a processor and a memory. The processor is configured to execute instructions stored in the memory, so that the computing device executes the method described in the first aspect and any possible implementation thereof above.
[0045] In an eighth aspect, an embodiment of the present application provides a computing device cluster, which includes at least one computing device. Each computing device includes a processor and a memory; the processors of at least one computing device are configured to execute instructions stored in the memories of at least one computing device, so that the computing device cluster executes the method described in the first aspect and any possible implementation thereof above.
[0046] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes computer program instructions. When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the method described in the first aspect and any possible implementation thereof above, where the computing device cluster includes at least one computing device.
[0047] In a tenth aspect, an embodiment of the present application provides a computer program product including instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method described in the first aspect and any possible implementation thereof above, where the cluster of computing devices includes at least one computing device.
[0048] It can be understood that the beneficial effects of the second aspect to the tenth aspect above can be referred to the relevant descriptions in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The following briefly introduces the drawings required for the description of the embodiments or the prior art.
[0050] Figure 1 It is a schematic diagram of the architecture of a data cleaning system provided in an embodiment of the present application;
[0051] Figure 2 It is a schematic diagram of the process of a data cleaning method provided in an embodiment of the present application;
[0052] Figure 3a It is a schematic diagram of the training process of a data cleaning model provided in an embodiment of the present application;
[0053] Figure 3b It is a schematic diagram of the process of an example of a data cleaning method provided in an embodiment of the present application;
[0054] Figure 4 It is a schematic diagram of the architecture of another data cleaning system provided in an embodiment of the present application;
[0055] Figure 5 It is a schematic diagram of the user interface of another data cleaning system provided in an embodiment of the present application;
[0056] Figure 6 It is a schematic diagram of the composition of a data cleaning device provided in an embodiment of the present application;
[0057] Figure 7 It is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0058] Figure 8 It is a schematic diagram of the structure of a cluster of computing devices provided in an embodiment of the present application;
[0059] Figure 9 It is a schematic diagram of the structure of another cluster of computing devices provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In this text, the term "and / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In this text, the symbol " / " indicates that the related objects are in an "or" relationship. For example, A / B means A or B.
[0061] In the description of the specification and claims of this application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe the specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0062] In the embodiments of this application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0063] In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0064] In the description of the embodiments of this application, unless otherwise specified, the meaning of "several" refers to one or more. For example, several processing units refers to one or more processing units, etc.; several elements refers to one or more elements, etc.
[0065] To facilitate the understanding of the solution provided in the embodiments of this application, some terms related to this solution are briefly introduced first.
[0066] Data quality: refers to the degree to which data meets requirements such as integrity, uniqueness, consistency, accuracy, validity, and timeliness. The higher the degree to which data meets the requirements, the higher the data quality; the lower the degree to which data meets the requirements, the lower the data quality. Data quality is a key link in big data processing. It is the detection of data quality for different data sources, and problematic data information should be screened out through quality rules, and then the data should be cleaned, such as manually modifying and correcting problematic data.
[0067] Integrity: refers to the completeness of data. The tasks of data integrity usually include: (1) incomplete model design, such as incomplete uniqueness constraints and incomplete references; (2) incomplete data entries, such as missing or unavailable data records; (3) incomplete data attributes, such as null values of data attributes.
[0068] Uniqueness: It refers to the situation where there are no duplicate data values for a certain data item or a set of data. If there are duplicate and redundant data, it will lead to the inability to coordinate business and trace processes.
[0069] Consistency: It means that the types and meanings of data elements need to be consistent and clear. The tasks of data consistency usually include: inconsistent data models of multi-source data, such as inconsistent naming, data structures, and constraint rules; inconsistent data entities, such as inconsistent data encoding, naming and meanings, classification hierarchies, and lifecycles; data inconsistency and data content conflict tasks in the case of multiple copies of the same data.
[0070] Accuracy: It means that it is necessary to ensure that the data must reflect the real business content. That is to say, there should be no errors in the data. For example, the salary income of employees cannot be incorrect.
[0071] Validity: It means that the values and formats of data meet the requirements of data definitions or business definitions. For example, the formats of phone numbers and email addresses.
[0072] Timeliness: It refers to ensuring that the data is updated in a timely manner according to the timeliness requirements of users for information acquisition.
[0073] Data quality management: It refers to a series of activities such as identifying, measuring, monitoring, and warning about data quality problems that may occur in each stage of data from planning, acquisition, storage, sharing, maintenance, application, and its entire lifecycle. And by improving and enhancing the management level of the organization, the data quality can be improved to a certain extent. The ultimate goal of data management is to enhance the value of data in use through reliable data and ultimately win more economic benefits for the enterprise.
[0074] Data middle platform: It is the core in the new type of informatization application framework system. Exemplarily, the data middle platform is the precipitation of the business and data of each business unit during the digital transformation of government and enterprises, constructing a data construction, management, and usage system including data technology, data governance, data operation, etc., to realize data empowerment. These services are strongly related to the business of the enterprise, are unique and reusable to this enterprise. It is the precipitation of the enterprise's business and data, which can not only reduce duplicate construction and collaboration costs, but also be the source of differential competitive advantages. Broadly speaking, the data middle platform includes various data technologies, such as data models, algorithm services, data products, data management, etc.
[0075] Operator: It mainly refers to the processing unit used to process data, which can be interpreted as a mapping from a function space to a function space. When using an operator, input data and output data are required, and the operator then completes the corresponding data transformation.
[0076] Data cleaning: The process of reexamining and validating data, aiming to delete duplicate information, correct existing errors, and ensure data consistency. That is to say, data cleaning is equivalent to "washing away" the "dirty data", which is the last procedure to discover and correct identifiable errors in data files, including checking data consistency, handling invalid values and missing values, etc. Since the data in the data warehouse is a collection of data oriented to a certain theme, this data is extracted from multiple business systems and includes historical data, so it is inevitable that some data is incorrect and some data conflicts with each other. These incorrect or conflicting data are obviously what we don't want, called "dirty data". Washing away the "dirty data" according to certain rules is data cleaning. The task of data cleaning is to filter out the data that does not meet the requirements, and hand over the filtered results to the business management department to confirm whether to filter it out or let the business unit correct it before extraction. The data that does not meet the requirements mainly includes three categories: incomplete data, incorrect data, and duplicate data.
[0077] Calibration: A method of data correction or modification. For example, modifying data with incorrect information to data with correct information; modifying data with missing information to data with complete information.
[0078] Structured Query Language (SQL): A special-purpose programming language, a database query and programming language used to access data and query, update, and manage relational database systems. Structured Query Language is a high-level non-procedural programming language that allows users to work on high-level data structures. Structured Query Language does not require users to specify the storage method of data, nor does it require users to understand the specific data storage method. Therefore, different database systems with completely different underlying structures can use the same Structured Query Language as the interface for data input and management. Structured Query Language statements can be nested, with great flexibility and powerful functions.
[0079] Artificial Intelligence (AI): Refers to the discipline that, based on computer science, integrates knowledge such as information theory, psychology, physiology, linguistics, logic, and mathematics to create computer systems that can simulate human intelligent behavior. Currently, artificial intelligence has received extensive attention from academia and the industrial community, and the application of AI is becoming more and more extensive, and it exceeds the level of ordinary humans in many application fields. For example, the application of AI technology in the field of machine vision (human recognition, image classification, object detection, etc.) makes the accuracy of machine vision higher than that of humans, and AI technology also has good applications in the fields of natural language processing and recommendation systems.
[0080] Machine learning: It is a core means of implementing AI. For a technical problem to be solved, the computer constructs an AI model based on existing data and then uses the AI model to predict the results. This method enables the computer to solve technical problems by mimicking human learning abilities (such as cognitive ability, discrimination ability, classification ability), so this method is called machine learning.
[0081] AI model: It is a mathematical model used in various applications of implementing AI by machine learning, such as a neural network model. The essence of an AI model is an algorithm, which includes a large number of parameters and calculation formulas (or calculation rules). The AI model can learn the internal laws and representation levels of the input data to obtain a non-linear function for the mapping relationship between the input and output, and process and analyze new input data according to this non-linear function. The AI model can be used in multiple application scenarios such as biology, medicine, and transportation. For example, when the target event is to clean problem data, multiple types of data such as the problem data can be input into the AI model to use the AI model to clean the problem data, etc. There are various AI models, and one of the most widely used types of AI models is the neural network model, which is a mathematical algorithm model that mimics the structure and function of a biological neural network (the central nervous system of an animal).
[0082] Data cleaning model: It can be any AI model, such as a neural network model. Set the training data set according to the data cleaning task; train the AI model according to the training data set to obtain the data cleaning model.
[0083] AI operator: The calculation rules and the corresponding one or more parameters adopted by one or more neural network layers or nodes in the data cleaning model are called an AI operator. An AI operator can solve a specific problem or implement a specific function, etc. The data cleaning model can be divided into multiple layers or multiple nodes, and each layer or each node includes a type of calculation rule and one or more parameters (used to represent a certain mapping, relationship, or transformation). According to the functions implemented, a data cleaning model can be divided into one or more AI operators.
[0084] Large Language Model (LLM): It generally refers to a neural network model containing an extremely large number of parameters (usually more than one billion). The Large Language Model shows excellent capabilities in various natural language processing tasks, such as text classification, sentiment analysis, summary generation, translation, etc. The Large Language Model is also called a large model, a language large model, an AI large model, etc. The LLM is an example of a data cleaning model.
[0085] The embodiment of the present application provides a data cleaning method, which is implemented based on a data middle platform of a data cleaning model, solves the problem of data processing after data quality, and further replaces the technical means of manually processing problematic data that requires manual handling with the technical means of automatically processing through preset cleaning operators and a data cleaning model, saving labor costs and improving the efficiency of data cleaning.
[0086] Exemplarily, in a government affairs scenario, the problematic data scanned may be in the millions. The embodiment of the present application automatically processes this problematic data through preset cleaning operators and a data cleaning model, saving labor costs and improving the efficiency of data cleaning.
[0087] In a government affairs scenario, the problematic data scanned by data quality rules may involve some important information. In the process of cleaning the problematic data through preset cleaning operators and a data cleaning model in the embodiment of the present application, the scope of the problematic data to be cleaned can be increased, and the number of data that cannot be cleaned can be reduced, such as reducing the number of data that cannot be cleaned to a two-digit number or even a single-digit number. In this way, on the one hand, it makes it possible to manually modify the data that cannot be cleaned; on the other hand, it avoids the phenomena of being unable to perform manual cleaning and directly discarding the problematic data when the number of data that cannot be cleaned is too large. In these phenomena, for example, the high labor cost or even the inability to achieve due to the large number of manual cleaning; for example, directly discarding the problematic data has the risk of losing important data.
[0088] Exemplarily, Figure 1 shows a schematic architecture diagram of a data cleaning system provided by the embodiment of the present application. As Figure 1 shown, the embodiment of the present application provides a data cleaning system 100, which mainly includes: a terminal device 110 and a server 120. Exemplarily, the terminal device 110 obtains a user request and related parameters input by the user and sends them to the server 120; the server 120 determines a data processing task according to the user request and related parameters, and the data processing task includes data quality analysis, data cleaning, etc.; the server 120 performs data quality analysis to obtain problematic data; the server 120 is deployed with preset cleaning operators and a data cleaning model, and the server 120 combines the preset cleaning operators and the data cleaning model (also referred to as AI operators) to clean the problematic data and obtains the cleaned data, and optionally generates a data analysis report; the terminal device 110 can obtain the cleaned data and the data analysis report from the server 120, and the user can obtain the cleaned data and the data analysis report through the terminal device 110.
[0089] Optionally, during the process of the server 120 determining the cleaned data, the server 120 may also send the uncleaned data to the terminal device 110 for manual cleaning by the user; the server 120 may also obtain the result of the user's manual cleaning from the terminal device 110, and then determine the final cleaned data and generate a data analysis report.
[0090] Further, the server 120 may include one or more servers ( Figure 1 taking a single server as an example for illustration), and the server 120 may provide the methods and / or devices provided in the embodiments of the present application for one or more terminal devices.
[0091] Further, an application related to the method and / or device of the present application may be installed on the terminal device 110. The above application or web page may provide an interface. The terminal device 110 may receive the input information entered by the user on the interface and send the above input information to the server 120. The server 120 is deployed with a data cleaning model and may input the received input information into the data cleaning model to output the cleaned data, thereby realizing the information interaction between the terminal device 110 and the server 120. Optionally, the server 120 may also return the data analysis report to the terminal device 110 to feedback relevant information about the problematic data to the user, such as data quality information, data cleaning information, etc.
[0092] It should be understood that in some alternative implementations, the terminal device 110 is deployed with a data cleaning model, and the terminal device 110 may also complete the action of obtaining the processing result based on the received input information by itself without the cooperation of the server. The embodiments of the present application do not limit this.
[0093] Next, the product form of the terminal device 110 is described. The terminal device 110 in the embodiments of the present application may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of the present application do not impose any restrictions on this. Exemplary embodiments of the terminal device 110 involved in this solution include, but are not limited to, electronic devices equipped with iOS, Android, Windows, Harmony OS, or other operating systems. The embodiments of the present application do not specifically limit the type of the electronic device.
[0094] Next, the product form of the server 120 will be described. It can be further understood that the server 120 can be various servers, such as servers with an X86 architecture, specifically, it can be a whole rack server, a blade server, a high-density server, a rack server, or a high-performance server, etc. In other words, the embodiments of the present application do not specifically limit the specific categories of the servers. Further, it can be understood that Figure 1 The structure of the server shown does not constitute a limitation on the structure of the server. The server may include more or fewer components than those shown, or combine certain components, or have different component arrangements.
[0095] Further, the server 120 can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers. It can also be configured as a cloud server or a cloud server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The cloud server cluster is deployed in several cloud data centers; the software can be an application that implements the object control method, etc., but is not limited to the above forms.
[0096] Next, the communication connection method between the terminal device 110 and the server 120 will be described. Exemplarily, the terminal device 110 and the server 120 are connected through a network, so that the terminal device 110 can access the cloud management platform deployed by the cloud server cluster. Among them, the network can be a wired network or a wireless network. For example, the wired network can be a cable network, an optical fiber network, a Digital Data Network (DDN), etc., and the wireless network can be a telecommunications network, an internal network, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Network (WLAN), a Metropolitan Area Network (MAN), a Public Service Telephone Network (PSTN), a Bluetooth network, a ZigBee network, a Global System for Mobile Communications (GSM), a Code Division Multiple Access (CDMA) network, a General Packet Radio Service (GPRS) network, etc. or any combination thereof.
[0097] It can be understood that the network can use any known network communication protocol to implement communication between different terminal device layers and the gateway. The above network communication protocol can be various wired or wireless communication protocols, such as Ethernet, universal serial bus (USB), firewire, global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), new radio (NR), bluetooth, wireless fidelity (Wi-Fi), and other communication protocols.
[0098] In a possible scenario, the server 120 can be used as a cloud (a software platform that adopts application virtualization technology, integrating multiple functions such as software search, download, use, management, and backup). In specific use, the server 120 can deploy a cloud management platform and a data center, and the terminal device 110 and the cloud interact through the cloud management platform. Additionally, the data center can deploy nodes, where the nodes in the data center can be virtual machine instances, container instances, physical servers, etc.
[0099] In another possible scenario, the method provided by the embodiments of the present application can be implemented through software. The software has a terminal display module and a service running module. The terminal device 110 runs the terminal display module of the software, and the server 120 runs the service running module of the software. During the process of the terminal device 110 running the terminal display module of the software, it can call the service running module running on the server 120 to implement the method provided by the embodiments of the present application.
[0100] That is to say, the method provided by the embodiments of the present application can be applied to the above-mentioned terminal device 110 or the server 120. In specific implementation, it can run on the terminal device 110 or the server 120 in the form of software. For example, the software can be a service or an application program. The embodiments of the present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments of the present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0101] The above is the introduction to a data cleaning system 100 provided by the embodiments of the present application. Next, based on the above content, a data cleaning method provided by the embodiments of the present application will be introduced. It can be understood that the above method is proposed based on the data cleaning system 100 described above, and some or all of the content in the above method can refer to the description of the data cleaning system 100 above.
[0102] Exemplarily, Figure 2 shows a schematic flowchart of a data cleaning method provided by the embodiments of the present application. It can be understood that this method is executed by a computing device, and the computing device can be implemented by any device, equipment, platform, or device cluster with computing and processing capabilities, such as Figure 1 the server 120 shown. As Figure 2 shown, a data cleaning method mainly includes the following steps:
[0103] Step S210, obtain the data to be cleaned.
[0104] Step S220, identify the data to be cleaned according to at least one data quality rule to obtain at least one problematic data. The data quality rule is used to represent the rule that conforms to the problematic data.
[0105] Among them, the data quality rule is used to represent the rule that conforms to the problematic data. Problematic data refers to data whose data quality does not meet the requirements. For example, problematic data refers to data with quality problems such as missing data content, incorrect data type, and non-standard data form.
[0106] Step S230, in the case where the problematic data is the problematic data of the first type, use a preset cleaning operator to clean the problematic data of the first type to obtain a first cleaning result. Among them, the problematic data of the first type represents the problematic data that meets the requirement of being cleaned by the preset cleaning operator.
[0107] Step S240, when the problem data is the second type of problem data, use the data cleaning model to clean the second type of problem data to obtain a second cleaning result. Herein, the second type of problem data represents the problem data that meets the requirement of being cleaned by the data cleaning model.
[0108] In the embodiments of the present application, the problem data can be divided into multiple types, such as the first type and the second type. Optionally, the problem data can also be the third type. The first type of problem data represents the problem data that meets the requirement of being cleaned by a preset cleaning operator. The second type of problem data represents the problem data that meets the requirement of being cleaned by the data cleaning model. The third type of problem data represents the problem data that meets the condition of manual cleaning.
[0109] For the problem data, there may be one or more of the above-mentioned first type, second type, and third type.
[0110] For example, the first type of problem data indicates that the problem data can be cleaned by a preset cleaning operator to obtain the correct data. For example, the first type of problem data such as "075568849684", the standard format is: 0755 - 6620122, and the correct data of "0755 - 68849684" can be obtained by using the preset cleaning operator for cleaning.
[0111] For example, the second type of problem data indicates that the problem data can be cleaned by the data cleaning model to obtain the correct data. For example, the second type of problem data such as "Zhengda", which is not an error of the inherent rule and cannot be cleaned by the preset cleaning operator, can be cleaned by the data cleaning model. The data cleaning model has good text understanding ability, and the correct data of "Zhengzhou University" can be obtained after cleaning by the data cleaning model.
[0112] For example, the third type of problem data indicates that the problem data can be cleaned manually to obtain the correct data. The third type of problem data such as "West Steel Factory in the provincial capital", "West Steel Factory in the provincial capital" is a user-specific colloquial expression, which cannot be cleaned by the preset cleaning operator based on the inherent rule, nor can it be cleaned by the data cleaning model by understanding the meaning of the text. The third type of problem data needs to be manually cleaned to obtain the correct data of "West ** Steel Factory in Henan Province".
[0113] For example, the problem data of the first type and the second type indicate that the problem data needs to be cleaned by a preset cleaning operator and a data cleaning model to obtain correct data. The problem data of the first type and the second type are, for example, "The mobile phone number of the Zhengda Admissions Office is 075568849684". Among them, "075568849684" belongs to the problem data of the first type, and the correct data of "0755-68849684" can be obtained by using the preset cleaning operator; "Zhengda" belongs to the problem data of the second type, and the correct data of "Zhengzhou University" can be obtained after cleaning by the data cleaning model. Furthermore, after cleaning by the preset cleaning operator and the data cleaning model, the correct data of "The mobile phone number of the Zhengzhou University Admissions Office is 0755-68849684" can be obtained.
[0114] Optionally, in the case where the problem data is the problem data of the first type and the second type, use the preset cleaning operator and the data cleaning model for cleaning to obtain the third cleaning result. Among them, for the problem data that simultaneously meets the first type and the second type, the preset cleaning operator and the data cleaning model can be used for cleaning. In this way, the problems of the problem data in terms of the first type and the second type can be solved, and the effect of data cleaning is improved.
[0115] Optionally, the data quality rule represents a quality rule of at least one rule type in at least one dimension; the at least one dimension includes at least one of integrity, uniqueness, timeliness, validity, accuracy, and consistency, and the at least one rule type includes at least one of regular expression, enumerated value, range value, structured query language expression, and structured query language statement.
[0116] Optionally, according to the quality rule of the integrity dimension, judge the integrity of the data to be cleaned to obtain the integrity recognition result. For example, the loss or unavailability of data records results in incomplete data entries.
[0117] Optionally, according to the quality rule of the uniqueness dimension, judge whether the data to be cleaned is unique to obtain the uniqueness recognition result. Uniqueness means that for a certain data item or a group of data, there are no duplicate data values. If there are duplicate and redundant data situations, it will cause the business to be unable to coordinate and the process to be unable to be traced. For example, the occurrence of duplicate transaction records in bank data may lead to the failure of the accounting result.
[0118] Optionally, according to the quality rule of the timeliness dimension, judge whether the update of the data to be cleaned meets the timeliness requirement to obtain the timeliness recognition result. Timeliness means ensuring that the data is updated in a timely manner according to the user's time timeliness requirement for information acquisition. For example, judge whether the time stamp of the chat record data is the time of the latest chat record.
[0119] Optionally, according to the quality rules of the validity dimension, determine whether the data to be cleaned meets the validity requirements to obtain a validity recognition result. Validity means that the value and format of the data meet the requirements of the data definition or business definition, such as the formats of phone numbers and email addresses.
[0120] Optionally, according to the quality rules of the accuracy dimension, determine whether the data to be cleaned is accurate to obtain an accuracy recognition result. Accuracy means that it is necessary to ensure that the data must reflect the real business content. That is to say, there should be no errors in the data. For example, the salary income of employees should not be incorrect.
[0121] Optionally, according to the quality rules of the consistency dimension, determine whether the data to be cleaned meets the consistency requirements to obtain a consistency recognition result. Consistency means that the types and meanings of data elements need to be consistent and clear. For example, inconsistent data naming, inconsistent data structures, and inconsistent constraint rules.
[0122] Furthermore, based on at least one of the above-mentioned integrity recognition result, uniqueness recognition result, timeliness recognition result, validity recognition result, accuracy recognition result, and consistency recognition result, determine whether the data to be cleaned is problematic data.
[0123] Optionally, according to the regular expression, determine whether the data to be cleaned meets the requirements of the regular expression to obtain a regular expression discrimination result. A regular expression, also known as a rule expression, is a concept in computer science. Regular expressions are usually used to retrieve and replace text that conforms to a certain pattern (rule). For example, if the data to be cleaned is "name" data, the regular expression can be used to filter out the strings and numbers in the "name" data, and the filtered data is the problematic data.
[0124] Optionally, according to the enumerated values, determine whether the data to be cleaned is within the range of the enumerated values to obtain an enumerated value discrimination result. For example, if the data to be cleaned is "gender" data, the "gender" data can enumerate all results, that is, "male" and "female"; it is possible to determine the data outside the range of the enumerated values ("male", "female") as problematic data by judging whether "gender" is within the range of the enumerated values.
[0125] Optionally, according to the range values, determine whether the data to be cleaned is within the range represented by the range values to obtain a range value discrimination result. For example, if the data to be cleaned is "age" data, the value range of the "age" data can be determined based on experience. For example, the value range of the "age" data is greater than or equal to 0 and less than or equal to 150, that is, [0, 150]; it is possible to determine the data outside the value range [0, 150] as problematic data by judging whether "age" is within the data value range [0, 150].
[0126] Optionally, according to an expression and / or statement of the structured query language (SQL), it is determined whether the data to be cleaned meets the requirements of SQL, and a discrimination result corresponding to the SQL expression or statement is obtained. Exemplarily, the structured query language expression (SQL) is, for example, "Select [parameter A], From [parameter B]". The meaning of this SQL is that [parameter B] represents the source of the query data, and [parameter B] is, for example, a file name or a storage location; based on [parameter B], it is determined which data are the data to be cleaned; [parameter A] indicates the screening object; based on [parameter A], it is determined which data are screened out; according to the screening result, the problematic data is determined. For example, the data remaining after screening that does not meet the conditions set by [parameter A] is the problematic data.
[0127] Further, according to at least one of the discrimination result of the regular expression, the discrimination result of the enumerated value, the discrimination result of the range value, and the discrimination result corresponding to the SQL expression or statement, it is determined whether the data to be cleaned is problematic data. Among them, identifying and judging the data to be cleaned from different rule types improves the recognition accuracy of the data to be cleaned of different data types and improves the applicable data types of the data quality rules. It should be noted that the above rule types are only exemplary descriptions, and in actual scenarios, corresponding rule types can also be formulated according to the data types of the data.
[0128] Optionally, the above-mentioned preset cleaning operators include at least one of a duplicate data operator, a normalization operator, and a data type operator.
[0129] It should be noted that the content of this implementation manner does not constitute a limitation on the preset cleaning operators in the embodiments of the present application. The preset cleaning operators include multiple cleaning operators for cleaning different categories of data. One cleaning operator only processes the problematic data of one category. In this way, each cleaning operator is easier to implement. By classifying the problematic data, the user can understand the problematic data under different categories and the category to which the current problematic data belongs, improving the user experience.
[0130] Optionally, in the case where the problematic data is the problematic data of the first type, it is determined whether the problematic data of the first type is duplicate data; when the problematic data of the first type is duplicate data, the duplicate data operator is used to retain one of the multiple duplicate problematic data of the first type, and the cleaning result of the duplicate data operator is obtained. For example, a set of data of the problematic data is [data 1, data 2, data 3, data 2], and the duplicate data is [data 2]. Using the duplicate data operator to retain one of the multiple [data 2], the cleaning result [data 1, data 2, data 3] of the duplicate data operator is obtained.
[0131] Optionally, determine whether the problem data of the first type meets the normative requirements; when the problem data of the first type does not meet the normative requirements, use a normative operator to convert the format of the problem data of the first type into a normative format to obtain the cleaning result of the normative operator. A data format that meets the normative requirements is, for example: 0755-6620122; the problem data is, for example, (0755)68849686, 075568849684, etc.; use a normative operator to convert the format of the problem data of the first type into a normative format to obtain the cleaning result of the normative operator, such as 0755-68849686, 0755-68849684, etc.
[0132] Optionally, determine whether the data type of the problem data of the first type meets the requirements; when the data type of the problem data of the first type does not meet the requirements, use a data type operator to convert the data type of the problem data of the first type into a data type that meets the requirements to obtain the cleaning result of the data type operator. For example, in a government affairs scenario, "male" and "female" in the "gender" data are usually represented by 0 and 1 respectively, and the problem data is the characters "male" and "female"; determine whether the data type of the "gender" data meets the requirements of the "0, 1" data type; when the data type of the problem data of the first type is the non-conforming characters "male" and "female", use a data type operator to convert the data type of the problem data "male" into the conforming data type "0", and convert the data type of the problem data "female" into the conforming data type "1".
[0133] Furthermore, obtain the first cleaning result according to at least one of the cleaning result of the duplicate data operator, the cleaning result of the normative operator, and the cleaning result of the data type operator. Among them, the data cleaning processes of the duplicate data operator, the normative operator, and the data type operator are described. Different preset cleaning operators can process different categories of problem data, so as to perform targeted processing on the problem data and improve the accuracy of data cleaning.
[0134] Optionally, the above-mentioned problem data of the second type includes at least one of incomplete data and invalid data.
[0135] Optionally, input the incomplete data into the data cleaning model, and the data cleaning model modifies the incomplete data into complete data. For example, the incomplete data is "Zhengda", input "Zhengda" into the data cleaning model, and the data cleaning model modifies "Zhengda" into "Zhengzhou University".
[0136] Optionally, the invalid data is input into the data cleaning model, which modifies the invalid data into valid data. For example, if the invalid data is "Zhengda is a university in Hebei Province", the data cleaning model can deeply understand the meaning of the text and answer based on the existing knowledge. Therefore, when the problem data "Zhengda is a university in Hebei Province" is input into the data cleaning model, the data cleaning model modifies the data "Zhengda is a university in Hebei Province" into "Zhengzhou University is a university in Henan Province" according to the existing knowledge.
[0137] The prompt words of the data cleaning model are, for example, "Please, based on your existing knowledge, determine whether the following sentence is correct. If it is incorrect, please correct it and output the data after your correction; [problem data]".
[0138] Furthermore, based on the modified complete data and valid data, a second cleaning result is obtained. The data cleaning model is a neural network model that can well analyze problem data after training, expanding the cleaning scope of problem data, enabling more problem data to be cleaned, and improving the cleaning effect of problem data.
[0139] Next, an exemplary introduction to the training method of the data cleaning model is provided.
[0140] Exemplarily, Figure 3a is a schematic diagram of the training process of the data cleaning model provided in the embodiments of the present application. As Figure 3a shown, using the problem data as the training sample and the true value of the correct data corresponding to the problem data as the label; to distinguish the correct data corresponding to the label from the correct data output by the data cleaning model, the correct data corresponding to the label is also called the true value of the correct data, and the correct data output by the data cleaning model is also called the predicted value of the correct data.
[0141] During the training process, the problem data is input into the data cleaning model, and the data cleaning model outputs the predicted value of the correct data. Calculate the loss value between the true value of the correct data and the predicted value of the correct data. When the loss value meets the threshold requirement (for example, the loss value is less than the loss threshold), the training of the current training sample ends, and the iterative training of the next training sample continues until the training process of all training samples in the training dataset is completed. When the loss value does not meet the threshold requirement (for example, the loss value is greater than or equal to the loss threshold), the parameters of the data cleaning model are adjusted.
[0142] Optionally, the training sample for each training can be a single one or multiple ones, realizing the model training of batch training samples. That is to say, this training process can be a batch training process (after calculating the loss value of the batch samples, then adjusting the parameters of the data cleaning model); it can also be a training process of a single training sample (after calculating the loss value of a single training sample, adjusting the parameters of the data cleaning model).
[0143] It should be noted that during the training process, the above parameters for adjusting the data cleaning model can be all or most of the parameters of the data cleaning model, or some parameters of the data cleaning model (such as the parameters of the output layer). The corresponding training process for the latter is also called fine-tuning training or fine-tuning. After training, the neural network model already has basic analysis capabilities. Fine-tuning is a re-training process for optimization by adjusting a small number of parameters.
[0144] Optionally, the above data cleaning model is a large language model. The second type of problem data can be input into the large language model to output the second cleaning result. The large language model has high text understanding capabilities. The large language model can comprehensively analyze the second type of problem data and then modify the problem data into correct data.
[0145] Optionally, in the case where the problem data is the third type of problem data, the third type of problem data is output so that the third type of problem data can be manually cleaned to obtain the fourth cleaning result. The third type of problem data represents problem data that meets the conditions for manual cleaning.
[0146] Among them, the third type can be considered a type of data that is more difficult to clean than the first type and the second type. Manually cleaning data can be used as a guarantee means in the data cleaning process, which can ensure that all data is cleaned, guarantee the integrity and availability of the data. This is of extremely important significance in scenarios with high data importance such as government affairs scenarios.
[0147] Optionally, obtain the fourth cleaning result; according to the fourth cleaning result, train the data cleaning model. Optionally, the "training" here can be "fine-tuning". According to the fourth cleaning result, fine-tune the data cleaning model to optimize the data cleaning ability of the data cleaning model, so that the data cleaning model can clean the third type of data, expand the cleaning scope of the data cleaning model, and improve the data cleaning ability of the data cleaning model. It also further reduces the amount of data to be manually cleaned and improves the user experience.
[0148] Optionally, in the case where the problem data is the first type, the second type, and the third type of problem data, use the preset cleaning operator and the data cleaning model for cleaning to obtain the fifth cleaning result, and output the fifth cleaning result so that the data in the fifth cleaning result can be manually cleaned to obtain the sixth cleaning result.
[0149] In this implementation manner, a method for cleaning problem data that simultaneously meets the first type, the second type, and the third type is described. Combining the first cleaning by the preset cleaning operator and the second cleaning by the data cleaning model, manual cleaning can be understood as the third cleaning, realizing data cleaning in three stages of "preset cleaning operator + data cleaning model + manual cleaning". Manual cleaning ensures that all data is corrected, guaranteeing the integrity and availability of the data, which is of extremely important significance in scenarios with high data importance such as government affairs scenarios. Additionally, through the data cleaning operations in two stages of the preset cleaning operator and the data cleaning model, the amount of data that needs to be manually modified is lower, reducing the workload of manual modification and improving the user experience.
[0150] Optionally, the method further includes: generating a data analysis report based on the problem data and the cleaned data, where the data analysis report is used to represent the data quality information and data cleaning information corresponding to the problem data. The data analysis report is a visualization technical means, facilitating users to understand the data quality information, information related to the data cleaning work, etc., and improving the user experience. The embodiments of the present application do not limit the specific form and specific content of the data analysis report.
[0151] The above is the introduction to a data cleaning system 100 and a data cleaning method provided by the embodiments of the present application. Next, based on the above content, an example of the data cleaning method is introduced to facilitate further understanding of the above data cleaning system and the above data cleaning method. It can be understood that part or all of the content of an example of the above data cleaning method can refer to the description of the data cleaning system 100 and a data cleaning method above.
[0152] Exemplarily, Figure 3b shows a schematic flowchart of an example of the data cleaning method provided by the embodiments of the present application. As Figure 3b shown, the embodiments of the present application provide an example of the data cleaning method, mainly including the following steps:
[0153] Step S301, the user formulates data quality rules and creates a data cleaning job. Exemplarily, the user formulates the scanning rules for data quality, and can add new data quality rules or select from existing data quality rules. An example of the above preset data quality rules is the formulated data quality rules. For example, the user formulates data quality rules on the terminal and creates a data cleaning job, and the server executes the data cleaning job (i.e., the data cleaning method of the present application). Exemplarily, after formulating the data quality rules, the user can create a new data cleaning job, configure basic parameters, and bind multiple data quality rules.
[0154] Exemplarily, Figure 4A schematic diagram of a user interface provided by an embodiment of the present application is shown. As Figure 4 shown, taking the terminal device 110 as an example, an application program related to the method and / or device of the present application can be installed on the terminal device 110. The above application program or web page can provide an interface. The terminal device 110 can receive input information input by the user on the interface and send the above input information to the server 120. The server 120 is deployed with a preset cleaning operator and a data cleaning model, and can input the received input information into the data cleaning model to output the cleaned data, realizing the information interaction between the terminal device 110 and the server 120. Optionally, the server 120 can also return the data analysis report to the terminal device 110 to feedback relevant information about the problem data to the user, such as data quality information, data cleaning information, etc.
[0155] Exemplarily, in Figure 4 it, in the first interface, the software name is "a certain data middle platform", and the user can edit data quality rules. Specifically, the user clicks the "Add Rule" button to add a custom data quality rule and jumps to the second interface. In the second interface, the detailed content of the newly added custom rule can be set, such as parameters like rule name, belonging directory, dimension, rule type, scenario classification, rule description, and rule definition, and then click Save.
[0156] Exemplarily, in Figure 4 it, in the third interface, the user clicks the "New" button to create a new data cleaning job; open the third interface, configure parameters such as job name, problem handler, subscription configuration, scheduling configuration, data object, and bind the newly created data quality rule, and then click Submit to generate a data cleaning job. Taking the fourth interface of "binding the newly created data quality rule" as an example, in the fourth interface, several data quality rules can be associated. After completing the creation of the new data cleaning job, the user can start the job.
[0157] Exemplarily, taking the addition of data quality rules as an example, the newly added data quality rules usually have six different dimensions, namely integrity, uniqueness, timeliness, validity, accuracy, and consistency; at the same time, each dimension corresponds to five different rule types, namely: regular expression, enumerated value, range value, SQL (structured query language) expression, and SQL statement. Furthermore, quality rule formulation can be carried out through the above dimensions and rule types. For example, if there are strings and numbers in "Name", it is dirty data, and the "Name" can be corrected through a regular expression; taking gender as an example, the enumerated values can be male and female; taking gender as an example, the range value can be from zero to one hundred years old, etc.; SQL expressions and SQL statements are SQL-format expressions.
[0158] Step S302: The server performs a quality scan on the data to be cleaned. The server performs a quality scan on the data to be cleaned according to the relevant information of the data cleaning job, scans to obtain problematic data, and performs data cleaning on the problematic data.
[0159] Step S303: The server classifies the problematic data to obtain different types of problematic data, such as the problematic data of the first type, the problematic data of the second type, the problematic data of the first type and the second type, etc.
[0160] Exemplarily, when the server executes the data cleaning job, the user can view the corresponding data cleaning results in different categories in the job monitoring, such as viewing the problematic data of different types. This step can be executed periodically to achieve the purpose of regular data cleaning, and the user can regularly view the data cleaning results in the job monitoring.
[0161] Exemplarily, in Figure 4 , in the fifth interface, the user can perform task supervision and view the data cleaning results through job monitoring. For example, the user can view the running status and running details of the data cleaning job, and perform operations such as rerunning, canceling, manually retrying, and viewing details on the data cleaning job. For example, after the data cleaning task fails, the user can click on problem handling, select a preset cleaning operator and a data cleaning model to perform data cleaning again. For data that cannot be processed by the preset cleaning operator and the data cleaning model, manual cleaning can also be performed, and the data marked after manual cleaning is used for training and fine-tuning of the data cleaning model.
[0162] Step S304: The server uses a preset cleaning operator to perform data cleaning on the problematic data that at least includes the problematic data of the first type. Exemplarily, the preset cleaning operator can be obtained by manual writing and is also called a written operator. For different types of problematic data, different operators need to be written to clean and modify the problematic data to ensure the authenticity and usability of the data. Among them, the problematic data that at least includes the problematic data of the first type means that the problematic data can belong to the first type or can belong to the first type and other types.
[0163] Exemplarily, common operators include: repetitive data operator, normalization operator, data type operator, etc. Repetitive data operator: mainly solves the problem of some repetitive data in the data. For multiple identical repetitive data, only one of them is retained. Normalization operator: mainly targets some non-standard data; for example, there may be multiple non-compliant filling methods for contact information. The compliant format is: 1234-1234567, and the actual filled data are (1234)1234567, 12341234567, etc. The normalization operator needs to convert the non-compliant operators into a compliant format. Data type operator: mainly targets data with inconsistent data types; for example, in the government affairs scenario, gender male and female are usually represented by 0 and 1, but characters such as "male" and "female" often appear during the filling process. The data type operator needs to convert these characters into compliant integer numbers.
[0164] Step S305, the server uses the data cleaning model to clean at least the problem data of the second type. The manually written preset cleaning operator is obtained by manual writing and has certain limitations. It may be ineffective for some data. The data cleaning model can further clean the data that cannot be cleaned by the preset cleaning operator, improving the usability of the data. Among them, the problem data that at least contains the second type means that the problem data can belong to the second type, or can belong to the second type and other types.
[0165] Exemplarily, the manually written preset cleaning operator can handle problem data as much as possible and performs well in dealing with some repetitive, normalized, consistent, and typed data. However, it cannot handle the data integrity and validity scenarios. For example, in terms of integrity, the school information filled in by users may be some abbreviations, with integrity problems, such as: USTC, SZU, ZZU. These data cannot be processed by manual operators, but the data cleaning model can well solve this problem. Through the data cleaning model, "ZZU" can be quickly converted into "Zhengzhou University", and the data that cannot be cleaned in step S304 is assisted to be processed by the data cleaning model, that is, the problem data is cleaned and modified by the data cleaning model. Then, based on the above two different operators, almost all problem data can be solved, avoiding the risk of important data loss caused by discarding problem data and reducing or avoiding the workload of manual cleaning by users.
[0166] Step S306, the server determines whether all the problem data has been completely cleaned. If the determination result is "yes", then step S309 is executed; if the determination result is "no", then step S307 is executed to output the problem data that cannot be cleaned through the user interface.
[0167] Step S307: For the problem data that cannot be cleaned, the user manually cleans and marks the data. In this step, the way to clean the data is manual operation and marking, which is a further process for the data that cannot be cleaned by the preset cleaning operator and the data cleaning model, so as to ensure the integrity and availability of the data. The problem data cleaned manually is also called the third type of problem data.
[0168] Exemplarily, for the data that cannot be cleaned by the data cleaning model, manual marking processing can be carried out. In the government affairs scenario, there are some data whose integrity cannot be directly cleaned by the data cleaning model. For example, for the work unit data such as "West Steel Factory in the provincial capital", the data cleaning model cannot solve the data integrity problem. At this time, it is necessary to manually clean the data by supplementing the correct work unit information, and complete the final data cleaning through manual processing, and record the problem data processed manually and the corrected data.
[0169] Step S308: The server fine-tunes the trained data cleaning model based on the data cleaned manually. The server can put the data marked in Step S307 into the data cleaning model for fine-tuning to make the data cleaning model more intelligent; and through iterative training of multiple fine-tuning, the data cleaning model becomes more intelligent and efficient.
[0170] Step S309: The server generates a data analysis report. The quality report generated by the server, for example, includes the number of problem data with data quality problems in the data source that can be shown, the proportion of problem data in the total amount of the overall data, the number of problem data cleaned by the preset cleaning operator and the data cleaning model, etc.; so that the user can see the data quality and the data cleaning situation in the first time.
[0171] Exemplarily, in Figure 4 In the sixth interface, the user can view the data cleaning job results. Taking the data analysis report as an example, the user can select the data cleaning job, click on several controls under the data quality information, and then view the relevant index situations of the data quality task, such as the comprehensive quality index, the historical trend of the quality index, the proportion of abnormal data, etc.; click on several controls under the data cleaning information, and then view the relevant index situations of the data cleaning task, such as the data results of data cleaning by the preset cleaning operator, the data cleaning model, and manual cleaning.
[0172] The above is the introduction of a data cleaning system 100, a data cleaning method, and their examples provided by the embodiments of the present application. Next, based on the above content, from the perspective of software implementation, another data cleaning system will be introduced to further exemplarily illustrate the specific implementation of the above data cleaning system, the above data cleaning method, and their examples in terms of software modules. It can be understood that part or all of the content of the above another data cleaning system can refer to the description of the data cleaning system 100, the above data cleaning method, and their examples in the above text.
[0173] Exemplarily, Figure 5 shows a schematic architecture diagram of another data cleaning system provided by the embodiments of the present application. As Figure 5 shown, the embodiments of the present application provide another data cleaning system, mainly including:
[0174] A data quality rule formulation module 510, which is used to formulate scanning rules for data quality. Exemplarily, newly added data quality rules usually have six different dimensions, namely integrity, uniqueness, timeliness, validity, accuracy, and consistency.
[0175] A data quality scanning module 520, which is used to perform quality scanning on data through custom quality rules, filter out problematic data, and classify the problematic data of different error types to obtain different types of problematic data.
[0176] A data cleaning module 530 includes a preset cleaning operator sub-module 531, a data cleaning model sub-module 532, a manual cleaning sub-module 533, and a data cleaning model fine-tuning sub-module 534. The preset cleaning operator sub-module 531 is used to correct at least the problematic data of the first type scanned out through preset cleaning operators written manually. The data cleaning model sub-module 532 is used to call a data cleaning model to perform data cleaning on at least the problematic data of the second type. The manual cleaning sub-module 533 is used to perform manual correction on at least the problematic data of the third type, such as problematic data that cannot be cleaned by both the preset cleaning operator and the data cleaning model, to clean at least the problematic data of the second type. The data cleaning model fine-tuning sub-module 534 is used to fine-tune and train the data cleaning model according to the data that cannot be cleaned by the data cleaning model, that is, the data cleaned manually, so as to improve the data cleaning ability of the data cleaning model and reduce the workload of manual cleaning.
[0177] A data analysis report generation module 540, which is used to generate a quality report for the data after data cleaning and repair, and is used to display the quality information, cleaning results, and other indicators of the problematic data.
[0178] Next, based on the same inventive concept, in combination with Figure 6 , a detailed example description will be given to the data cleaning device provided in the embodiments of the present application respectively.
[0179] Exemplarily, Figure 6 shows a schematic composition diagram of a data cleaning device provided in an embodiment of the present application. As Figure 6 shown, an embodiment of the present application provides a data cleaning device 600, which is applied to a server and mainly includes:
[0180] An acquisition module 610, configured to acquire data to be cleaned.
[0181] A processing module 620, configured to identify the data to be cleaned according to at least one data quality rule to obtain at least one problem data, where the data quality rule is used to represent a rule that conforms to the problem data; and, in the case that the problem data is the first type of problem data, use a preset cleaning operator to clean the first type of problem data to obtain a first cleaning result, where the first type of problem data represents problem data that satisfies cleaning using a preset cleaning operator; and, in the case that the problem data is the second type of problem data, use a data cleaning model to clean the second type of problem data to obtain a second cleaning result, where the second type of problem data represents problem data that satisfies cleaning using a data cleaning model.
[0182] In a possible implementation manner, the data quality rule represents a quality rule of at least one rule type in at least one dimension; the at least one dimension includes at least one of integrity, uniqueness, timeliness, validity, accuracy, and consistency, and the at least one rule type includes at least one of a regular expression, an enumerated value, a range value, a structured query language expression, and a structured query language statement.
[0183] In a possible implementation manner, the above-mentioned processing module 620 is specifically configured to: judge the integrity of the data to be cleaned according to the quality rule of the integrity dimension to obtain an integrity identification result; judge whether the data to be cleaned is unique according to the quality rule of the uniqueness dimension to obtain a uniqueness identification result; judge whether the update of the data to be cleaned meets the timeliness requirement according to the quality rule of the timeliness dimension to obtain a timeliness identification result; judge whether the data to be cleaned meets the validity requirement according to the quality rule of the validity dimension to obtain a validity identification result; judge whether the data to be cleaned is accurate according to the quality rule of the accuracy dimension to obtain an accuracy identification result; judge whether the data to be cleaned meets the consistency requirement according to the quality rule of the consistency dimension to obtain a consistency identification result; and judge whether the data to be cleaned is problem data according to at least one of the integrity identification result, the uniqueness identification result, the timeliness identification result, the validity identification result, the accuracy identification result, and the consistency identification result.
[0184] In a possible implementation, the above-mentioned processing module 620 is specifically configured to: determine whether the data to be cleaned meets the requirements of the regular expression according to the regular expression, and obtain a regular expression discrimination result; determine whether the data to be cleaned is within the range of the enumeration value according to the enumeration value, and obtain an enumeration value discrimination result; determine whether the data to be cleaned is within the range represented by the range value according to the range value, and obtain a range value discrimination result; determine whether the data to be cleaned meets the requirements of the structured query language expression according to the structured query language expression, and obtain a structured query language expression discrimination result; determine whether the data to be cleaned meets the requirements of the structured query language statement according to the structured query language statement, and obtain a structured query language statement discrimination result; determine whether the data to be cleaned is problematic data according to at least one of the regular expression discrimination result, the enumeration value discrimination result, the range value discrimination result, the structured query language expression discrimination result, and the structured query language statement discrimination result.
[0185] In a possible implementation, the above-mentioned preset cleaning operators include at least one of a duplicate data operator, a normalization operator, and a data type operator.
[0186] In a possible implementation, the above-mentioned processing module 620 is specifically configured to: when the problematic data is the problematic data of the first type, determine whether the problematic data of the first type is duplicate data; when the problematic data of the first type is duplicate data, use the duplicate data operator to retain one of the multiple duplicate problematic data of the first type, and obtain a cleaning result of the duplicate data operator; determine whether the problematic data of the first type meets the normalization requirements; when the problematic data of the first type does not meet the normalization requirements, use the normalization operator to convert the format of the problematic data of the first type into a normalized format, and obtain a cleaning result of the normalization operator; determine whether the data type of the problematic data of the first type meets the requirements; when the data type of the problematic data of the first type does not meet the requirements, use the data type operator to convert the data type of the problematic data of the first type into a data type that meets the requirements, and obtain a cleaning result of the data type operator; obtain a first cleaning result according to at least one of the cleaning result of the duplicate data operator, the cleaning result of the normalization operator, and the cleaning result of the data type operator.
[0187] In a possible implementation, the above-mentioned problematic data of the second type includes at least one of incomplete data and invalid data. The above-mentioned processing module 620 is specifically configured to: input the incomplete data into the data cleaning model, and the data cleaning model modifies the incomplete data into complete data; input the invalid data into the data cleaning model, and the data cleaning model modifies the invalid data into valid data; obtain a second cleaning result according to the modified complete data and valid data.
[0188] In a possible implementation, the above data cleaning model is a large language model. The above processing module 620 is specifically configured to: input the problem data of the second type into the large language model to output a second cleaning result.
[0189] In a possible implementation, the above processing module 620 is further configured to: when the problem data is the problem data of the first type and the second type, perform cleaning using a preset cleaning operator and the data cleaning model to obtain a third cleaning result.
[0190] In a possible implementation, the above processing module 620 is further configured to: when the problem data is the problem data of the third type, output the problem data of the third type so that the problem data of the third type is manually cleaned to obtain a fourth cleaning result, where the problem data of the third type represents the problem data that meets the manual cleaning condition.
[0191] In a possible implementation, the above processing module 620 is further configured to: obtain the fourth cleaning result; and train the data cleaning model according to the fourth cleaning result.
[0192] In a possible implementation, the above processing module 620 is further configured to: when the problem data is the problem data of the first type, the second type, and the third type, perform cleaning using a preset cleaning operator and the data cleaning model to obtain a fifth cleaning result, and output the fifth cleaning result so that the data in the fifth cleaning result is manually cleaned to obtain a sixth cleaning result.
[0193] In a possible implementation, the above processing module 620 is further configured to: generate a data analysis report according to the problem data and the cleaned data, where the data analysis report is used to represent the data quality information and data cleaning information corresponding to the problem data.
[0194] It should be noted that Figure 6 the multiple functional modules shown in can be implemented by software or can be implemented by hardware. Exemplarily, next, taking the processing module 620 as an example, the implementation manner of the processing module 620 is introduced. Similarly, the implementation manner of the acquisition module 910 can refer to the implementation manner of the processing module 620.
[0195] As an example of a software functional unit, the processing module 620 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the processing module 620 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.
[0196] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.
[0197] As an example of a hardware functional unit, the processing module 620 may include at least one computing device, such as a server, etc. Alternatively, the processing module 620 may also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0198] The multiple computing devices included in the processing module 620 may be distributed in the same region or in different regions. The multiple computing devices included in the processing module 620 may be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the processing module 620 may be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0199] It should be noted that in other embodiments, the above device (i.e., the data cleaning device 600) may additionally be provided with one or more modules for performing any of the steps included in the above implementation manner. The above device is also referred to as the related device of the embodiments of the present application. The steps to be implemented by one or more modules in the above device may be specified as needed. The functions of the above device may also be implemented using more or fewer modules than those in the embodiments of the present application. One or more modules in the above device are used to implement different steps in the above method respectively, so as to implement all the functions of the above device.
[0200] The present application also provides a computing device 700. As Figure 7 shown, the computing device 700 includes: a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other through the bus 702. The computing device 700 may be a server, such as a central server, an edge server, or a local server in a local data center, or may be an electronic device such as a desktop computer, a laptop computer, or a smart phone. It should be understood that the present application does not limit the number of processors and memories in the computing device 700.
[0201] The bus 702 can be a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. Among them, the unified bus is also known as the Lingqu bus. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only one line is used in Figure 7 , but it does not mean that there is only one bus or one type of bus. The bus 704 can include a path for transmitting information between various components of the computing device 700 (e.g., the memory 706, the processor 704, the communication interface 708).
[0202] The processor 704 can include any one or more of computing devices such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP), an ASIC, an FPGA, a CPLD, an NPU, a SoC, an offload card, an acceleration card, etc.
[0203] The memory 706 can include volatile memory, such as Random Access Memory (RAM). The processor 704 can also include non-volatile memory, such as Read-Only Memory (ROM), flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD). In addition, the memory 706 can also be implemented through Storage Class Memory (SCM), Phase Change Memory (PCM), or other types of storage media.
[0204] It should be noted that in the same computing device, the same type of storage medium can be configured to implement the function of the memory 706, or two or more types of storage media can be configured to implement the function of the memory 706. This application does not make any limitations in this regard.
[0205] The memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement the functions of the foregoing one or more modules, thereby implementing the method described in the above embodiments. That is to say, the memory 706 stores instructions for executing the method described in the above embodiments.
[0206] Alternatively, the memory 706 stores executable codes, and the processor 704 executes the executable codes to respectively implement the functions of the above device, thereby implementing the method described in the above embodiments. That is to say, the memory 706 stores instructions for executing the method described in the above embodiments.
[0207] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 700 and other devices or communication networks.
[0208] As a possible implementation, the computing device 700 may also include a chip system. The chip system includes a processor and a power supply circuit. The power supply circuit is used to supply power to the processor, and the processor is used to execute the operation steps corresponding to the method of the embodiments of the present application. For the sake of brevity, it will not be elaborated here. Among them, the processor can be implemented by a GPU, or can be implemented by a computing device or an AI chip such as a DPU, an NPU, an XPU, a SoC, an offload card, or an acceleration card.
[0209] As a possible implementation, the computing device 700 may include multiple types of processors 704, that is, the computing device 700 is a heterogeneous device. For example, the computing device 700 includes a CPU and a GPU, and at least one of the processors 704 can execute the operation steps corresponding to the method of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0210] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be an electronic device such as a desktop computer, a laptop computer, or a smart phone.
[0211] Such as Figure 8As shown, the computing device cluster includes at least one computing device 700. Instructions for executing the methods described in the above embodiments may be stored in the memory 706 of one or more of the computing devices 700 in the computing device cluster.
[0212] In some possible implementation manners, partial instructions for executing the methods described in the above embodiments may also be separately stored in the memory 706 of one or more of the computing devices 700 in the computing device cluster. In other words, a combination of one or more computing devices 700 may jointly execute the instructions for executing the methods described in the above embodiments.
[0213] It should be noted that the memories 706 in different computing devices 700 in the computing device cluster may store different instructions, respectively for executing partial functions of the above device. That is, the instructions stored in the memories 706 of different computing devices 700 may implement the functions of one or more modules in the above device.
[0214] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Wherein, the network may be a wide area network or a local area network, etc. Figure 9 A possible implementation manner is shown. As Figure 9 shown, two computing devices 700A and 700B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manners, instructions for the functions of one or more modules in the above device are stored in the memory 706 of the computing device 700A. At the same time, instructions for the functions of one or more other modules in the above device are stored in the memory 706 of the computing device 700B.
[0215] It should be understood that Figure 9 the functions of the computing device 700A shown in
[0216] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in this computing device cluster may be similarly referred to the Figure 8 and Figure 9 connection manner of the described computing device cluster. The difference is that instructions for executing the methods in the above embodiments may be stored in the memory 706 of one or more of the computing devices 700 in this computing device cluster.
[0217] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store some instructions for executing the foregoing method respectively. In other words, the combination of one or more computing devices 700 can jointly execute the instructions for executing the foregoing method.
[0218] In addition to the foregoing method and electronic device, an embodiment of the present application may further provide a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the method according to various embodiments of the present application described in the foregoing "method" section of this specification. Among them, the computer program product can be written in any combination of one or more programming languages for computer program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. Among them, the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer program code can be executed completely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or completely executed on a remote computing device or server. In some examples, when the computer program instructions are executed by a computing device cluster including at least one computing device, at least one computing device in the computing device cluster is caused to execute the method in the foregoing embodiment.
[0219] In addition, embodiments of the present application may also provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present application described in the "Methods" section above of this specification. The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but not be limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. In some examples, when the instructions are run by a cluster of computing devices including at least one computing device, at least one computing device in the cluster of computing devices is caused to execute the methods in the above embodiments.
[0220] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0221] The method steps in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0222] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0223] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0224] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not described or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0225] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the specific details disclosed above are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present application to necessarily adopt the above specific details for implementation.
[0226] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0227] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.
[0228] The above description has been given for purposes of illustration and description. In addition, this description does not intend to limit the embodiments of the present application to the form disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
[0229] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data cleaning method, characterized in that: include: Get the data to be cleaned; According to at least one data quality rule, the data to be cleaned is identified to obtain at least one problematic data, wherein the data quality rule is used to represent a rule that satisfies the problematic data; In the case where the problem data is a first type of problem data, using a preset cleaning operator to clean the first type of problem data to obtain a first cleaning result, wherein the first type of problem data indicates problem data that satisfies the requirement of using the preset cleaning operator for cleaning; In the case where the problem data is the second type of problem data, the second type of problem data is cleaned using a data cleaning model to obtain a second cleaning result, where the second type of problem data represents problem data that satisfies the need for cleaning using the data cleaning model.
2. The method according to claim 1, characterized in that The data quality rules represent quality rules of at least one rule type of at least one dimension; the at least one dimension includes at least one of completeness, uniqueness, timeliness, validity, accuracy and consistency, and the at least one rule type includes at least one of a regular expression, an enumeration value, a range value, a structured query language expression and a structured query language statement.
3. The method according to claim 2, characterized in that The step of identifying the data to be cleaned according to at least one data quality rule to obtain at least one problematic data includes: According to the quality rules of the integrity dimension, the integrity of the data to be cleaned is judged to obtain an integrity identification result; According to the quality rules of the uniqueness dimension, determine whether the data to be cleaned is unique, and obtain a unique identification result; According to the quality rules of the timeliness dimension, it is judged whether the update of the data to be cleaned meets the timeliness requirements, and a timeliness identification result is obtained; According to the quality rules of the validity dimension, determine whether the data to be cleaned meets the validity requirements and obtain the validity identification result; According to the quality rules of the accuracy dimension, determine whether the data to be cleaned is accurate, and obtain an accuracy recognition result; According to the quality rules of the consistency dimension, determine whether the data to be cleaned meets the consistency requirements and obtain a consistency identification result; Whether the data to be cleaned is problematic data is determined based on at least one of the integrity identification result, the uniqueness identification result, the timeliness identification result, the validity identification result, the accuracy identification result, and the consistency identification result.
4. The method according to claim 2 or 3, characterized in that: The step of identifying the data to be cleaned according to at least one data quality rule to obtain at least one problematic data includes: According to the regular expression, determine whether the data to be cleaned meets the requirements of the regular expression, and obtain a regular expression judgment result; According to the enumeration value, judging whether the data to be cleaned is within the range of the enumeration value, and obtaining the enumeration value judgment result; According to the range value, determining whether the data to be cleaned is within the range represented by the range value, and obtaining a range value determination result; According to the structured query language expression, determining whether the data to be cleaned meets the requirements of the structured query language expression, and obtaining a structured query language expression determination result; According to the structured query language statement, determining whether the data to be cleaned meets the requirements of the structured query language statement, and obtaining a structured query language statement determination result; Whether the data to be cleaned is problematic data is determined according to at least one of the regular expression determination result, the enumeration value determination result, the range value determination result, the structured query language expression determination result, and the structured query language statement determination result.
5. The method according to any one of claims 1 to 4, characterized in that: The preset cleaning operator includes at least one of a repetitive data operator, a standardization operator, and a data type operator.
6. The method according to claim 5, characterized in that When the problem data is the first type of problem data, using a preset cleaning operator to clean the first type of problem data to obtain a first cleaning result includes: When the first type of problematic data is repeated data, a repetitive data operator is used to retain one of the repeated problematic data of the first type, and a cleaning result of the repetitive data operator is obtained; When the first type of problem data does not meet the normative requirements, a normative operator is used to convert the format of the first type of problem data into a normative format, and a cleaning result of the normative operator is obtained; When the data type of the first type of problem data does not meet the requirements, a data type operator is used to convert the data type of the first type of problem data into a data type that meets the requirements, and a cleaning result of the data type operator is obtained; The first cleaning result is obtained according to at least one of the cleaning result of the repetitive data operator, the cleaning result of the standardization operator, and the cleaning result of the data type operator.
7. The method according to any one of claims 1 to 6, characterized in that: The second type of problematic data includes at least one of incomplete data and invalid data; When the problem data is the second type of problem data, using the data cleaning model to clean the second type of problem data to obtain a second cleaning result includes: Inputting the incomplete data into the data cleaning model, and the data cleaning model modifies the incomplete data into complete data; and / or, inputting the invalid data into the data cleaning model, and the data cleaning model modifies the invalid data into valid data; The second cleaning result is obtained according to the modified complete data and / or the valid data.
8. The method according to any one of claims 1 to 7, characterized in that: The data cleaning model is a large language model; When the problem data is the second type of problem data, using the data cleaning model to clean the second type of problem data to obtain a second cleaning result includes: The second type of question data is input into the large language model to output the second cleaning result.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: In the case where the problem data is the first type and the second type of problem data, the preset cleaning operator and the data cleaning model are used to perform cleaning to obtain a third cleaning result.
10. A server, characterized in that: The server includes a processor and a memory, and the processor is used to execute instructions stored in the memory, so that the server executes the method according to any one of claims 1 to 9.
Citation Information
Cited By
Data quality evaluation method and system based on rule engine and machine learning
CN121030265A