Data cleaning method and system based on large language model
Through the data cleaning method based on the large language model, the efficiency and accuracy problems of traditional methods when dealing with complex and dynamic data are solved, and more efficient and accurate data cleaning is achieved, and the degree of adaptability and automation is significantly improved.
Patent Information
- Application Number
- CN202411939526.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional data cleaning methods are difficult to accurately identify semantic errors in text, and require cumbersome rule settings when processing multiple data types, which cannot adapt to the dynamic changes of the data, resulting in reduced efficiency and accuracy.
The data cleaning method based on the large language model is adopted, and the data set to be cleaned is collected and sorted, and formatted into a text format suitable for model processing is carried out, text error correction, exception processing, entity recognition, semantic consistency check, missing data generation and duplicate data processing are carried out, and problems and processing measures in the cleaning process are recorded to generate structured reports.
Improve the accuracy and efficiency of data cleaning, enhance the adaptability to complex and dynamic data, reduce manual intervention, significantly reduce error rates, and provide a more reliable basis for data analysis and decision-making.
Smart Images

Figure CN119988832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data cleaning, and more specifically, to a data cleaning method and system based on a large language model. Background Art
[0002] Data cleaning is an important step in data analysis and machine learning. Traditional data cleaning methods usually rely on rule-driven processing processes, including missing value filling, outlier detection, duplicate record removal, and format standardization. Although these methods are effective in dealing with simple data problems, they often fail to cope with complex and heterogeneous data.
[0003] For example, traditional methods may not be able to accurately identify semantic errors in text, or require cumbersome rule settings when processing multiple data types. In addition, manually formulated data cleaning rules may not be able to adapt to the dynamic changes of data, thereby reducing the efficiency and accuracy of data cleaning. With the increase in data volume and the diversification of data types, traditional methods are becoming less flexible and inefficient. Summary of the invention
[0004] According to the present invention, a data cleaning method and system based on a large language model are provided to solve the problem that traditional methods may not be able to accurately identify semantic errors in text, or require cumbersome rule settings when processing multiple data types. In addition, manually formulated data cleaning rules may not be able to adapt to the dynamic changes of data, thereby reducing the efficiency and accuracy of data cleaning. With the increase in data volume and the diversification of data types, traditional methods appear to be less flexible and inefficient.
[0005] According to a first aspect of the present invention, a data cleaning method based on a large language model is provided, comprising:
[0006] Collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned;
[0007] Based on the large language model, the cleaned text data is corrected and normalized, the cleaned text data is handled abnormally, and the cleaned text data is recognized and normalized based on the large language model.
[0008] Perform semantic consistency check on the cleaned text data based on the large language model, generate missing data on the cleaned text data based on the large language model, and process duplicate data on the cleaned text data based on the large language model;
[0009] All problems, implemented processing measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model are recorded, and a structured report is generated based on the records.
[0010] Optionally, the data set to be cleaned is collected and sorted, and formatted into a text format suitable for model processing to obtain text data to be cleaned, including:
[0011] Collect raw data from multiple data sources, and extract required data fields and records according to business needs. The multiple data sources include databases, file systems, and network APIs. The raw data includes structured data or unstructured data.
[0012] The collected data is converted into a text format that can be processed by the large language model, that is, JSON format, to obtain text data.
[0013] Optionally, text error correction and normalization processing are performed on the text data to be cleaned based on the large language model, including:
[0014] The text data to be cleaned and the prompt template are constructed into error correction standard instructions, the error correction standard instructions are sent to the large language model, and the text data is corrected and standardized based on the large language model;
[0015] Correct spelling errors and grammatical errors in the text, and unify capitalization and punctuation;
[0016] If there are abbreviations in the text, expand them to their full forms, ensuring that the processed text remains semantically accurate and coherent.
[0017] Optionally, exception processing is performed on the text data to be cleaned based on the large language model, including:
[0018] The text data to be cleaned and the prompt template are constructed into exception handling instructions, the exception handling instructions are sent to the large language model, and the text data is processed based on the large language model;
[0019] Analyze the semantic context of the text data to be cleaned and mark outliers that do not conform to the context or convention;
[0020] For detected outliers, correction suggestions are provided to ensure the accuracy and consistency of the content.
[0021] Optionally, entity recognition and standardization processing are performed on the cleaned text data based on the large language model, including:
[0022] Develop a set of entity standardization rules based on the characteristics of the data set and business needs;
[0023] Construct entity standardization instructions including text data to be cleaned, entity standardization rules and prompt templates, send the entity standardization instructions to the large language model, and perform entity recognition and standardization processing on the text data based on the large language model.
[0024] Optionally, a semantic consistency check is performed on the text data to be cleaned based on a large language model, including:
[0025] Construct a semantic check instruction containing the text data to be cleaned and a prompt template, send the semantic check instruction to the large language model, and perform a semantic consistency check on the text data based on the large language model;
[0026] Check whether the text content is semantically consistent and whether there are any logical contradictions;
[0027] If any logical contradictions or semantic inconsistencies are found, correction suggestions are provided to ensure the accuracy and coherence of the data.
[0028] Optionally, missing data is generated for the text data to be cleaned based on the large language model, including:
[0029] Construct a missing data generation instruction including the text data to be cleaned and a prompt template, send the missing data generation instruction to the large language model, and generate missing data for the text data based on the large language model;
[0030] Identify missing parts in text content and automatically fill in the missing parts. By analyzing the context, structure and semantics of existing data, generate supplementary content that is consistent with the overall style of the dataset and logically reasonable, ensuring the integrity and availability of the data.
[0031] Optionally, duplicate data processing is performed on the text data to be cleaned based on the large language model, including:
[0032] Construct a duplicate data processing instruction that includes the text data to be cleaned and a prompt template, send the duplicate data processing instruction to the large language model, process the duplicate data in the text based on the large language model, and eliminate redundant information in the data set by identifying and merging duplicate data.
[0033] Optionally, all problems, implemented treatment measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model are recorded, and a structured report is generated based on the records, including:
[0034] At each step of data cleaning, identify and record problems in the dataset, including spelling errors, grammatical errors, semantic inconsistencies, missing data, duplicate records, and outliers;
[0035] For each identified problem, record the corresponding treatment measures and correction methods;
[0036] List the data before and after cleaning for comparison, and evaluate the effectiveness of the cleaning operation based on the results of data cleaning and the preset goals;
[0037] All records are combined to generate a structured report containing charts and statistical information to demonstrate the process and effect of data cleaning.
[0038] According to another aspect of the present invention, a data cleaning system based on a large language model is also provided, comprising:
[0039] A module for obtaining text data to be cleaned is used to collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned;
[0040] The first module for processing text data is used to perform text error correction and normalization processing on the cleaned text data based on the large language model, perform exception processing on the cleaned text data based on the large language model, and perform entity recognition and normalization processing on the cleaned text data based on the large language model;
[0041] The second module for processing text data is used to perform semantic consistency check on the text data to be cleaned based on the large language model, generate missing data on the text data to be cleaned based on the large language model, and process duplicate data on the text data to be cleaned based on the large language model;
[0042] A structured report generation module is used to record all problems, implemented processing measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model, and generate a structured report based on the records.
[0043] Therefore, through the powerful text understanding and generation capabilities of the large language model, human intervention is reduced, the accuracy and efficiency of data cleaning are improved, and the adaptability of the data cleaning system to complex and dynamic data is enhanced. This not only improves the automation level of data cleaning, but also significantly reduces the error rate, providing a more reliable foundation for data analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:
[0045] Figure 1 A schematic diagram of a process of a data cleaning method based on a large language model described in this embodiment;
[0046] Figure 2 This is a schematic diagram of data cleaning based on a large language model according to this embodiment;
[0047] Figure 3 This is a schematic diagram of a data cleaning system based on a large language model described in this embodiment. DETAILED DESCRIPTION
[0048] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.
[0049] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.
[0050] According to a first aspect of the present invention, a data cleaning method 100 based on a large language model is provided. Figure 1 As shown, the method 100 includes:
[0051] S101: Collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned;
[0052] S102: performing text error correction and normalization processing on the cleaned text data based on the large language model, performing exception processing on the cleaned text data based on the large language model, and performing entity recognition and normalization processing on the cleaned text data based on the large language model;
[0053] S103: performing semantic consistency check on the text data to be cleaned based on the large language model, generating missing data on the text data to be cleaned based on the large language model, and processing duplicate data on the text data to be cleaned based on the large language model;
[0054] S104: Record all problems, implemented processing measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model, and generate a structured report based on the records.
[0055] Specifically, Large Language Models (LLMs) are artificial intelligence models trained with large amounts of data, with strong text understanding and generation capabilities, designed to understand and generate natural language text. These models are usually based on deep learning technology, especially the Transformer architecture, which can capture the complexity and diversity of language, thus showing strong capabilities in the field of natural language processing (NLP).
[0056] Large language models can understand the context and semantic relationships of text, thus providing more intelligent data cleaning solutions. They can automatically identify potential problems in the data, such as grammatical errors, inconsistent data formats, or ambiguous records, and correct them through contextual reasoning. In addition, large language models can handle a variety of data types, including text, numbers, and dates, and automatically perform missing value filling, outlier detection, and data standardization tasks by learning patterns and structures in the data.
[0057] refer to Figure 2 As shown in the figure, the data cleaning method based on the large language model is mainly divided into eight modules: 1) data preprocessing; 2) text error correction and normalization; 3) outlier processing; 4) entity recognition and standardization; 5) semantic consistency check; 6) missing data generation; 7) duplicate data processing; 8) data cleaning report generation
[0058] 1. Data preprocessing module
[0059] 1) Collect and organize the data set to be cleaned and format it into a text format suitable for model processing.
[0060] 2) Collect raw data from multiple data sources (such as databases, file systems, network APIs. Raw data may include structured data (such as CSV, SQL databases) or unstructured data (such as text files, HTML pages). Extract the required data fields and records based on business needs.
[0061] Convert the collected data into a text format that can be processed by the large language model, that is, the json format. Perform type conversion on numerical data, date data, etc. For example, convert the date format to the "YYYY-MM-DD" format and convert numerical data into a text description containing units. Perform preliminary processing on unstructured text, such as removing special characters (such as HTML tags, non-printing characters), unifying the encoding format to UTF-8, and deleting unnecessary spaces and line breaks.
[0062] 2. Text Correction and Standardization
[0063] This module uses the natural language processing capabilities of a large language model to automatically identify and correct spelling errors and grammatical errors in text data, and normalizes the text (including unifying capitalization, punctuation, abbreviations, etc.) to improve data accuracy, consistency, and quality.
[0064] First, the text data to be cleaned and the prompt template are constructed into instructions, and then the instructions are sent to the large language model to guide it to perform text error correction and normalization on the text data. The prompt template is as follows:
[0065] Please correct and normalize the following text data. The requirements are as follows:
[0066] -Correct spelling errors, grammatical errors and other issues in the text, and unify capitalization and punctuation;
[0067] - If there are abbreviations in the text, try to expand them into their full form;
[0068] - The processed text should maintain semantic accuracy and coherence.
[0069] The text data is as follows:
[0070] {Text data to be cleaned}
[0071] 3. Outlier processing
[0072] This module uses the semantic understanding and reasoning capabilities of large language models to automatically detect and process data points in the dataset that do not conform to the semantic context or conventions (i.e., outliers). By analyzing the semantic and statistical characteristics of the data, it identifies potential outliers and provides correction or processing suggestions, thereby improving the quality and consistency of the data.
[0073] First, you need to build an instruction that contains the text data to be cleaned and a prompt template, and then send the instruction to the large language model to guide it to process the outliers in the text data. The prompt template is as follows:
[0074] Please perform outlier detection and processing on the following text content. The requirements are as follows:
[0075] -Pay attention to analyzing the semantic context of the text and marking outliers that do not fit the context or convention;
[0076] - For detected outliers, provide correction suggestions, but ensure the accuracy and coherence of the content.
[0077] The text content is as follows:
[0078] {Text data to be cleaned}
[0079] 4. Entity Recognition and Standardization
[0080] This module uses the entity recognition function of the large language model to identify entities in the text, including but not limited to names of people, places, organizations, dates, etc., and standardizes the identified entities according to predefined entity standardization rules, including unified format, disambiguation, normalized names, etc., to ensure the availability and structure of the data.
[0081] 1) According to the characteristics of the data set and business needs, formulate a set of entity standardization rules. As follows (for reference only):
[0082] a. Naming conventions
[0083] For example, for names of people, you can formulate rules to unify all names into the format of "surname, first name".
[0084] b.Abbreviation expansion
[0085] Abbreviations in institution or place names may require expansion to their full form.
[0086] c. Unified unit
[0087] When processing entities containing numeric values, rules may require the unification of measurement units. For example, converting all length units from "feet" to "meters" or converting currency units from "dollars" to "RMB" (the specific conversion depends on the target application scenario).
[0088] d. Remove irrelevant characters
[0089] Remove irrelevant characters in the entity name, such as special symbols, extra spaces, etc. For example, remove the brackets and spaces in "Apple Inc." and standardize it to "Apple Inc.".
[0090] e. Standardized format
[0091] For entities such as date and time, the rules may require a uniform format, for example, all dates should be in the "YYYY-MM-DD" format and all times should be in the "HH:MM:SS" format.
[0092] f.Multi-language support
[0093] In a multilingual environment, rules may need to handle entity names in different languages. For example, for Chinese names, their original writing order and format may be maintained; while for English names, they may be standardized according to the above "surname, first name" format (although this is not common in English, it is only used as an example).
[0094] 2) Construct an instruction containing the text data to be cleaned, entity standardization rules, and prompt template, and then send the instruction to the large language model to guide it to perform entity recognition and standardization on the text data. The prompt template is as follows:
[0095] **Task**: Entity Recognition and Normalization
[0096] **Input text**: {text data to be cleaned}
[0097] **Entity recognition part**:
[0098] -Please identify specific types of entities such as names of people, places, and organizations from the above text.
[0099] -The context of the text should be considered during the recognition process to ensure the accuracy of the entity.
[0100] - Output the recognized entities in the form of a list, each entity contains its type and corresponding text.
[0101] **Entity Standardization Section**:
[0102] Please process the identified entity list according to the following normalization rules: {Predefined entity normalization rules}
[0103] 5. Semantic consistency check
[0104] This module uses the deep semantic understanding capabilities of large language models to perform semantic consistency checks on the text in the dataset and provide repair suggestions to ensure that the information in the data is logically reasonable, coherent, and free of logical contradictions. This module is designed to improve the accuracy and reliability of data, especially when dealing with text containing complex information.
[0105] First, you need to build an instruction that contains the text data to be cleaned and a prompt template, and then send the instruction to the large language model to guide it to check the semantic consistency of the text data. The prompt template is as follows:
[0106] Please carefully check the semantics of the following text content to see if there are any logical contradictions. If you find any logical contradictions or semantic inconsistencies, please provide correction suggestions to ensure the accuracy and coherence of the data.
[0107] The text content is as follows:
[0108] {Data to be cleaned}
[0109] Missing data generation
[0110] This module uses the contextual understanding and generation capabilities of large language models to automatically fill in the missing parts of text data. By analyzing the context, structure, and semantics of existing data, it generates supplementary content that is consistent with the overall style of the dataset and logically reasonable, ensuring the integrity and availability of the data.
[0111] First, you need to build an instruction that contains the text data to be cleaned and a prompt template, and then send the instruction to the large language model to guide it to identify and generate the missing data in the text. The prompt template is as follows:
[0112] **Task**: Identify and generate missing parts of text
[0113] **Input text**: {data to be cleaned}
[0114] **Identify the missing parts of the text**: Identify the missing parts of the above text and mark them out
[0115] **Generate the missing content in the text**: Generate the missing content based on the context of the above text. The generated content must be consistent with the context logic and the sentences must be fluent.
[0116] 7. Duplicate data processing
[0117] This module uses the contextual understanding and semantic analysis capabilities of the large language model to automatically identify and process duplicate records in the data set. This module aims to improve the neatness and quality of data by identifying and merging duplicate data and eliminating redundant information in the data set, thereby improving the accuracy and efficiency of data analysis.
[0118] First, you need to build an instruction that contains the text data to be cleaned and a prompt template, and then send the instruction to the large language model to guide it to process repeated data in the text. The prompt template is as follows:
[0119] Please use your powerful natural language processing skills to help me process the duplicate records in the following text. The specific tasks are as follows:
[0120] 1. **Mission Description**:
[0121] -Text to be processed: {data to be cleaned};
[0122] - I need you to identify all duplicate or highly similar records in the text;
[0123] 2. **Similarity evaluation**:
[0124] - When evaluating similarity between records, consider semantic proximity rather than just superficial textual matching;
[0125] -You can use your built-in text similarity calculation function, or adopt other suitable technical means;
[0126] 3. **Duplicate record processing**:
[0127] - For the duplicate records identified, please provide merge suggestions. The suggestions should include which records should be merged and which field values should be retained after the merge;
[0128] - If possible, automatically merge these records according to preset rules and update the text;
[0129] 4.**Output requirements**:
[0130] - Output a report containing duplicate record identification results and merging suggestions;
[0131] - If an automatic merge operation was performed, please update the text content as well.
[0132] 8. Data cleaning report generation
[0133] The main function of this module is to generate a detailed data cleaning report, recording all the problems identified during the data cleaning process, the treatment measures implemented, the results generated, and the various parameters in the process. This report not only helps users understand the overall process and effect of data cleaning, but also provides a basis for subsequent data analysis, decision-making and auditing.
[0134] 1) Problem identification records
[0135] At each step of data cleaning, identify and record problems in the data set. These problems include spelling errors, grammatical errors, semantic inconsistencies, missing data, duplicate records, outliers, etc. a. Each data cleaning module marks and records all problems found when processing data;
[0136] b. Classify the identified issues by type to facilitate further analysis and processing. The classification criteria may include the type of issue (such as grammatical errors, missing data, etc.), severity (such as high, medium, low), scope of impact, etc.
[0137] 2) Processing measures and correction records
[0138] For each problem identified, the corresponding treatment measures and correction methods are recorded. These records describe in detail the actions taken by the model, the parameters used, etc.
[0139] a. Record of treatment measures: Record the treatment methods for each problem, such as the generation method of filling missing data, the replacement strategy for outliers, the merging operation of duplicate data, etc.
[0140] b. Parameter and confidence records: Record the parameters used by the model when processing problems (such as threshold settings, etc.) to ensure transparency and traceability of all operations.
[0141] 3) Record of cleaning results
[0142] a. Comparison before and after cleaning: List the data comparison before and after cleaning, highlighting the improved parts in the cleaning process and the improvement of data quality.
[0143] b. Result effectiveness evaluation: Based on the results of data cleaning and preset goals, evaluate the effectiveness of the cleaning operation, such as reduction in error rate, improvement in data consistency, etc.
[0144] 4) Report Generation
[0145] All records are combined to generate a structured report. The report is easy to read and contains charts and statistical information to clearly show the process and effect of data cleaning.
[0146] Therefore, compared with traditional data cleaning methods, it has many beneficial effects. Improve the degree of automation: The present invention uses a large language model to reduce the reliance on manual intervention in the data cleaning process, significantly improving the level of automation. Improve accuracy and efficiency: Through intelligent semantic analysis and context understanding, the present invention can more accurately identify and correct data problems, and improve the accuracy and efficiency of data cleaning. Enhance adaptability and flexibility: The large language model has adaptive capabilities and can dynamically adjust the cleaning strategy to adapt to changes in different data types and data sets, ensuring continuous and effective data cleaning. Improve transparency and traceability: Generate a detailed data cleaning report, record each step and result of the cleaning process, ensure the transparency and traceability of the cleaning process, and facilitate auditing and compliance inspections. Through the above beneficial effects, the present invention not only improves the efficiency and accuracy of data cleaning, but also enhances the flexibility and adaptability of data cleaning, providing strong support for various data processing application scenarios.
[0147] Optionally, the data set to be cleaned is collected and sorted, and formatted into a text format suitable for model processing to obtain text data to be cleaned, including:
[0148] Collect raw data from multiple data sources, and extract required data fields and records according to business needs. The multiple data sources include databases, file systems, and network APIs. The raw data includes structured data or unstructured data.
[0149] The collected data is converted into a text format that can be processed by the large language model, that is, JSON format, to obtain text data.
[0150] Optionally, text error correction and normalization processing are performed on the text data to be cleaned based on the large language model, including:
[0151] The text data to be cleaned and the prompt template are constructed into error correction standard instructions, the error correction standard instructions are sent to the large language model, and the text data is corrected and standardized based on the large language model;
[0152] Correct spelling errors and grammatical errors in the text, and unify capitalization and punctuation;
[0153] If there are abbreviations in the text, expand them to their full forms, ensuring that the processed text remains semantically accurate and coherent.
[0154] Optionally, exception processing is performed on the text data to be cleaned based on the large language model, including:
[0155] The text data to be cleaned and the prompt template are constructed into exception handling instructions, the exception handling instructions are sent to the large language model, and the text data is processed based on the large language model;
[0156] Analyze the semantic context of the text data to be cleaned and mark outliers that do not conform to the context or convention;
[0157] For detected outliers, correction suggestions are provided to ensure the accuracy and consistency of the content.
[0158] Optionally, entity recognition and standardization processing are performed on the cleaned text data based on the large language model, including:
[0159] Develop a set of entity standardization rules based on the characteristics of the data set and business needs;
[0160] Construct entity standardization instructions including text data to be cleaned, entity standardization rules and prompt templates, send the entity standardization instructions to the large language model, and perform entity recognition and standardization processing on the text data based on the large language model.
[0161] Optionally, a semantic consistency check is performed on the text data to be cleaned based on a large language model, including:
[0162] Construct a semantic check instruction containing the text data to be cleaned and a prompt template, send the semantic check instruction to the large language model, and perform a semantic consistency check on the text data based on the large language model;
[0163] Check whether the text content is semantically consistent and whether there are any logical contradictions;
[0164] If any logical contradictions or semantic inconsistencies are found, correction suggestions are provided to ensure the accuracy and coherence of the data.
[0165] Optionally, missing data is generated for the text data to be cleaned based on the large language model, including:
[0166] Construct a missing data generation instruction including the text data to be cleaned and a prompt template, send the missing data generation instruction to the large language model, and generate missing data for the text data based on the large language model;
[0167] Identify missing parts in text content and automatically fill in the missing parts. By analyzing the context, structure and semantics of existing data, generate supplementary content that is consistent with the overall style of the dataset and logically reasonable, ensuring the integrity and availability of the data.
[0168] Optionally, duplicate data processing is performed on the text data to be cleaned based on the large language model, including:
[0169] Construct a duplicate data processing instruction that includes the text data to be cleaned and a prompt template, send the duplicate data processing instruction to the large language model, process the duplicate data in the text based on the large language model, and eliminate redundant information in the data set by identifying and merging duplicate data.
[0170] Optionally, all problems, implemented treatment measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model are recorded, and a structured report is generated based on the records, including:
[0171] At each step of data cleaning, identify and record problems in the dataset, including spelling errors, grammatical errors, semantic inconsistencies, missing data, duplicate records, and outliers;
[0172] For each identified problem, record the corresponding treatment measures and correction methods;
[0173] List the data before and after cleaning for comparison, and evaluate the effectiveness of the cleaning operation based on the results of data cleaning and the preset goals;
[0174] All records are combined to generate a structured report containing charts and statistical information to demonstrate the process and effect of data cleaning.
[0175] Thus, by using the powerful text understanding and generation capabilities of the advanced large language model, problems in the data can be automatically identified and corrected, improving the automation and accuracy of data cleaning. At the same time, the invention automatically generates a cleaning report, records the cleaning process and results, improves the transparency and traceability of data cleaning, and facilitates audits and compliance checks.
[0176] According to another aspect of the present invention, a data cleaning system 300 based on a large language model is also provided. Figure 3 As shown, the system 300 includes:
[0177] A module 310 for obtaining text data to be cleaned is used to collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned;
[0178] A first module 320 for processing text data is used to perform text error correction and normalization processing on the cleaned text data based on the large language model, perform exception processing on the cleaned text data based on the large language model, and perform entity recognition and normalization processing on the cleaned text data based on the large language model;
[0179] A second text data processing module 330 is used to perform semantic consistency check on the cleaned text data based on the large language model, generate missing data on the cleaned text data based on the large language model, and process duplicate data on the cleaned text data based on the large language model;
[0180] The structured report generation module 340 is used to record all problems, implemented processing measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model, and generate a structured report based on the records.
[0181] A data cleaning system 300 based on a large language model in an embodiment of the present invention corresponds to a data cleaning method 100 based on a large language model in another embodiment of the present invention, and will not be described in detail here.
[0182] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0183] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0184] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0186] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0187] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A data cleaning method based on a large language model, characterized in that: include: Collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned; Based on the large language model, the cleaned text data is corrected and normalized, the cleaned text data is handled abnormally, and the cleaned text data is recognized and normalized based on the large language model. Perform semantic consistency check on the cleaned text data based on the large language model, generate missing data on the cleaned text data based on the large language model, and process duplicate data on the cleaned text data based on the large language model; All problems, implemented processing measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model are recorded, and a structured report is generated based on the records.
2. The method according to claim 1, characterized in that Collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned, including: Collect raw data from multiple data sources, and extract required data fields and records according to business needs. The multiple data sources include databases, file systems, and network APIs. The raw data includes structured data or unstructured data. The collected data is converted into a text format that can be processed by the large language model, that is, JSON format, to obtain text data.
3. The method according to claim 1, characterized in that: Based on the large language model, the text data to be cleaned is corrected and normalized, including: The text data to be cleaned and the prompt template are constructed into error correction standard instructions, the error correction standard instructions are sent to the large language model, and the text data is corrected and standardized based on the large language model; Correct spelling errors and grammatical errors in the text, and unify capitalization and punctuation; If there are abbreviations in the text, expand them to their full forms, ensuring that the processed text remains semantically accurate and coherent.
4. The method according to claim 1, characterized in that Perform exception processing on the cleaned text data based on the large language model, including: The text data to be cleaned and the prompt template are constructed into exception handling instructions, the exception handling instructions are sent to the large language model, and the text data is processed based on the large language model; Analyze the semantic context of the text data to be cleaned and mark outliers that do not conform to the context or convention; For detected outliers, correction suggestions are provided to ensure the accuracy and consistency of the content.
5. The method according to claim 1, characterized in that Perform entity recognition and standardization on the cleaned text data based on the large language model, including: Develop a set of entity standardization rules based on the characteristics of the data set and business needs; Construct entity standardization instructions including text data to be cleaned, entity standardization rules and prompt templates, send the entity standardization instructions to the large language model, and perform entity recognition and standardization processing on the text data based on the large language model.
6. The method according to claim 1, characterized in that Perform semantic consistency checks on the text data to be cleaned based on a large language model, including: Construct a semantic check instruction containing the text data to be cleaned and a prompt template, send the semantic check instruction to the large language model, and perform a semantic consistency check on the text data based on the large language model; Check whether the text content is semantically consistent and whether there are any logical contradictions; If any logical contradictions or semantic inconsistencies are found, correction suggestions are provided to ensure the accuracy and coherence of the data.
7. The method according to claim 1, characterized in that Generate missing data for the cleaned text data based on the large language model, including: Construct a missing data generation instruction including the text data to be cleaned and a prompt template, send the missing data generation instruction to the large language model, and generate missing data for the text data based on the large language model; Identify missing parts in text content and automatically fill in the missing parts. By analyzing the context, structure and semantics of existing data, generate supplementary content that is consistent with the overall style of the dataset and logically reasonable, ensuring the integrity and availability of the data.
8. The method according to claim 1, characterized in that Based on the large language model, duplicate data processing is performed on the cleaned text data, including: Construct a duplicate data processing instruction that includes the text data to be cleaned and a prompt template, send the duplicate data processing instruction to the large language model, process the duplicate data in the text based on the large language model, and eliminate redundant information in the data set by identifying and merging duplicate data.
9. The method according to claim 1, characterized in that: Record all problems, implemented treatment measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model, and generate a structured report based on the records, including: At each step of data cleaning, identify and record problems in the dataset, including spelling errors, grammatical errors, semantic inconsistencies, missing data, duplicate records, and outliers; For each identified problem, record the corresponding treatment measures and correction methods; List the data before and after cleaning for comparison, and evaluate the effectiveness of the cleaning operation based on the results of data cleaning and the preset goals; All records are combined to generate a structured report containing charts and statistical information to demonstrate the process and effect of data cleaning.
10. A data cleaning system based on a large language model, characterized in that: include: A module for obtaining text data to be cleaned is used to collect and organize the data set to be cleaned, format it into a text format suitable for model processing, and obtain the text data to be cleaned; The first module for processing text data is used to perform text error correction and normalization processing on the cleaned text data based on the large language model, perform exception processing on the cleaned text data based on the large language model, and perform entity recognition and normalization processing on the cleaned text data based on the large language model; The second module for processing text data is used to perform semantic consistency check on the text data to be cleaned based on the large language model, generate missing data on the text data to be cleaned based on the large language model, and process duplicate data on the text data to be cleaned based on the large language model; A structured report generation module is used to record all problems, implemented processing measures, generated results and various parameters in the process of cleaning the text data to be cleaned based on the large language model, and generate a structured report based on the records.
Citation Information
Patent Citations
Text data cleaning system based on large language model
CN117910458A
Named entity recognition method and equipment based on large language model
CN119129596A
Cited By
Data reliability improvement method and device based on large model, medium and equipment
CN120653704A
Data cleaning method and system based on AI identification
CN120804519A
LLM-based industry map multi-modal report generation system
CN120996158A
Bid inviting and tendering template generation method based on large language model
CN121031563A
Statistically interpretable large-model active learning mass data screening method and device and software system
CN121188477A