Large model-based data cleaning method and related device
Through the data cleaning method based on the big model, using prompt words to generate regular expressions and executable code, the problem of cost-effectiveness caused by relying on manual code writing in the prior art is solved, and a more efficient data cleaning process is achieved.
Patent Information
- Application Number
- CN202411898342.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, data cleaning relies on manual code writing, resulting in high cost and low efficiency.
The data to be processed is processed by acquiring communication operators, obtaining regular expressions and executable code based on prompt words.
Reduces data cleaning costs, improves efficiency, and simplifies the complex process of manually writing regular expressions and code by users.
Smart Images

Figure CN120045837A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of large model technology, and more particularly to a data cleaning method and device based on a large model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the outbreak of big model technology, the competition for big models has focused on computing power and data, and high-quality data is the goal pursued by all relevant research parties. During the pre-training and fine-tuning training stages of the big model, the initial data collected at the beginning has various sources and formats, and the quality is poor, and needs further cleaning. The data cleaning work in the existing technology mostly relies on data R&D personnel to write code for different data. Whenever new data arrives, a lot of development work needs to be carried out, which is costly and inefficient. In view of this, the present invention aims to use big models to build a data cleaning framework for such data cleaning scenarios to reduce data cleaning costs and improve its efficiency. Summary of the Invention
[0003] The present disclosure provides a data cleaning method and device based on a large model, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] According to the first aspect, a data cleaning method based on a large model is provided, characterized in that the method includes: obtaining a communication operator in the distributed computing framework; obtaining a regular expression and an executable code for the data to be processed based on the prompt word; and processing the data to be processed based on the regular expression and the executable code to obtain processed data.
[0005] According to the second aspect, a data cleaning device based on a large model is provided, characterized in that the device includes: a prompt word template acquisition unit, which acquires the user's instructions for the data to be processed; acquires a prompt word template and a prompt word that match the instruction, and the prompt word is obtained by filling the prompt word template based on the instruction; a regular code acquisition unit, which acquires a regular expression and executable code for the data to be processed based on the prompt word; and a framework upgrade verification unit, which processes the data to be processed based on the regular expression and the executable code to obtain processed data.
[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.
[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as described in any implementation of the first aspect.
[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program, which implements the method described in any implementation manner of the first aspect when executed by a processor.
[0009] The technical solution disclosed in the present invention specifically solves the key problem of high cost and low efficiency caused by the reliance on manual code writing in the existing technology for data cleaning, helps to improve the flexible configuration of data cleaning, and lowers the threshold for using data cleaning technology.
[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 This is a schematic diagram of a scenario of data cleaning based on a large model provided by an embodiment of the present disclosure;
[0013] Figure 2 This is a flow chart of a data cleaning method based on a large model provided by an embodiment of the present disclosure;
[0014] Figure 3 This is a structural diagram of a data cleaning device based on a large model provided by an embodiment of the present disclosure;
[0015] Figure 4 It is a block diagram of an electronic device used to implement the large model-based data cleaning method of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0017] The present disclosure proposes a data cleaning method and related devices based on a large model.
[0018] It is understandable that the method can be integrated into various types of computing devices, including but not limited to terminal devices with text processing capabilities (such as smart phones, tablets, personal computers, etc.) and servers (such as application servers, Web servers, whether they are single servers or cluster-deployed server systems). The method does not depend on a specific hardware platform or software architecture. During the execution of the method, both independently running terminal devices and terminal-server architectures that work together through a network can effectively utilize the method disclosed herein. When the terminal device is executed independently, the method can be independent of the external network. In scenarios where higher processing performance or wider data resource support is required, the method of the present invention also supports communication between the terminal device and the server, utilizing the powerful computing power and rich data resources of the server to jointly complete the method. The method can adapt to different operating systems and platform environments, including mobile operating systems such as iOS, Android, and Hongmeng, desktop operating systems such as Windows and macOS, and server operating systems such as Linux and Unix.
[0019] It can be understood that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution of the present disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0020] It is understandable that the various embodiments covered by this disclosure have been deeply described and explained from the unique perspectives and different dimensions of the overall technical solution. In actual application scenarios, based on diverse needs, the technical solutions involved in different embodiments can be flexibly and freely combined according to specific circumstances. This combination is not a simple patchwork, but through careful integration and adaptation, a more comprehensive, efficient and comprehensive technical architecture that meets actual business needs is constructed, further expanding the applicability and feasibility of the disclosed technical solution in different fields and different situations.
[0021] Example 1
[0022] This disclosure has good applicability in various scenarios. For ease of understanding, Figure 1 A schematic diagram of a scenario of an embodiment of the present disclosure is shown.
[0023] The example scenario 100 is constructed by multiple modules with different functions.
[0024] The data processed in this scenario is text data, which can be formatted data, unformatted data, or semi-formatted data.
[0025] The front-end module 101 plays a key role in the overall architecture, directly interacting with users. It not only receives various user inputs but also provides a convenient interface, bridging the gap between users and the system's internal functions. This allows users to easily access system resources and obtain the services and feedback they need.
[0026] The regular code generation module 102 is focused on the instructions input by the user and automatically generates corresponding data matching regular expressions and codes with the help of the powerful ability of the large model. In order to realize this function, a prompt word template library is designed and constructed in advance, in which a plurality of different types of prompt word templates are stored. These prompt word templates are trained based on a large amount of data matching scenarios and requirements. The regular code generation module 102 has a model training function, which trains the generative large model and constructs the data of a large amount of natural language generation regular expressions and codes, so that the trained model can generate regular expressions and codes with high accuracy. The regular code generation module 102 also has a prompt word function. After receiving the instructions input by the user, the system can quickly carry out accurate matching in the prompt word template library and filter out the prompt word template that best matches it. Then, the prompt word fine-tuning technology is used, based on the prompt word template, to guide the large model to generate accurate and efficient regular expressions and codes. In this way, the complex process of users manually writing regular expressions and codes is greatly simplified, and the efficiency and accuracy of data matching logic generation are significantly improved. Even for users who lack professional programming knowledge, they can easily use instructions to implement complex data matching tasks.
[0027] The data cleaning execution module 103 plays a core role in the entire data cleaning process. This module mainly performs precise data cleaning based on the data to be processed, regular expressions, and executable code. The data cleaning execution module 103 has a regular expression executor function. In terms of the application of regular expressions, it has a variety of powerful execution modes, covering full match, exact match, and reverse match. The data cleaning execution module 103 also has a code executor function, which enables the matching mode of regular expressions to be automatically executed to meet diverse data cleaning needs. Through the execution results of regular expressions, the module can accurately locate the specific data columns and data sentences corresponding to the data to be processed, and then perform specified data cleaning operations on these specific data parts, such as data replacement and deletion. After completing all data cleaning operations, this module will integrate and splice the cleaned data results, and finally output the complete processing results. It will also present the processing details for each cleaning rule in detail, thereby providing comprehensive and transparent records and feedback for the data cleaning process, facilitating the subsequent data quality assessment and analysis work.
[0028] The operator management module 104 plays a very critical role. Figure 1 In the figure, three operators are described as an example. The initial data to be cleaned is input into operator 1; after operator 1 is executed, its processed result is input into operator 2; after operator 2 is executed, its processed result is input into operator 3; after operator 3 is executed, its processed result is used as the final cleaned result.
[0029] The operator management module 104 has an operator publishing function. Its core function is to convert regular expressions and executable codes generated during the data cleaning process into operator form for storage, and fully support various management operations and external publishing of these operators.
[0030] Operators have multiple important functions. In a single data cleaning task, operators can be efficiently reused. This feature is used in conjunction with the regular code acquisition module.
[0031] The operator management module 104 also has an operator call function. The regular code generation module generates the expression and executable code for the task based on the user's instructions. However, if the operator management module already has the corresponding regular expression and executable code, the regular code acquisition module can accurately obtain the matching regular expression and executable code based on the user's natural language description and then perform processing operations on the data to be cleaned. This greatly improves the efficiency of data cleaning and reduces duplication of work.
[0032] Furthermore, operators can be flexibly combined to build data processing pipelines for subsequent processing by pipeline modules. By combining operators with different functions into pipelines in a specific logical order, complex data cleaning processes can be automated, making the entire data cleaning process more systematic, orderly, and efficient. This meets the diverse data cleaning and processing needs of different scenarios and further improves the overall efficiency and adaptability of the data processing system.
[0033] Regular code generation module 102 and regular code acquisition module 302 ( Figure 1 Not shown, see Figure 3), that is, a pattern of one-time generation and multiple calls. When faced with a brand-new data cleaning task, due to the lack of ready-made regular expressions and executable code resources, it is necessary to rely on the regular code generation module, which uses its powerful functions to generate regular expressions and executable codes suitable for the task based on the instructions and other information entered by the user. Once a regular expression and executable code have been generated and stored in the corresponding operator management module during the previous task processing, when the same or similar task requirements are encountered in the future, the regular code acquisition module can directly intervene and quickly and accurately obtain the corresponding regular expression and executable code from the existing resource reserves, thereby avoiding the tedious process of repeated generation, effectively improving the overall efficiency of data cleaning work, and reducing unnecessary resource consumption and time costs.
[0034] The pipeline module 105 is highly flexible and convenient. On the one hand, it supports a visual graphical interface, allowing users to easily build data processing pipelines by freely dragging and adding operators, just like building blocks. Users only need to simply drag various operators in the operation interface and place them in sequence according to the logical order of data processing to quickly complete the initial construction of the pipeline. On the other hand, the module also provides a way to call the code, allowing users to call the corresponding operator directly in the code based on the operator name or identifier, thereby completing the construction of the pipeline.
[0035] Once a pipeline is complete, users can instantly call upon it to perform data cleaning tasks. Logically, a pipeline composed of multiple operators can be considered functionally similar and consistent to a single operator. This means users can use existing pipelines as foundational components, further nesting and combining them to build more complex and powerful pipeline architectures to address a variety of complex and ever-changing data cleaning and processing scenarios, greatly improving the adaptability and scalability of data processing systems to meet diverse business needs.
[0036] Example 2
[0037] Figure 2 A large model-based data cleaning method 200 according to an embodiment of the present disclosure is shown.
[0038] The method 200 includes: step 201, obtaining a user's instruction for processing data; obtaining a prompt word template and a prompt word matching the instruction, wherein the prompt word is obtained by filling the prompt word template based on the instruction.
[0039] The method 200 further includes: step 202, obtaining a regular expression and executable code for the data to be processed based on the prompt word.
[0040] The method 200 further includes: step 203, processing the data to be processed based on the regular expression and the executable code to obtain processed data.
[0041] To sum up, this technical solution has the following technical effects: First, by being able to obtain the user's instructions on the data to be processed and matching the corresponding prompt word templates, the user can conveniently express the need for data cleaning in natural language, thereby improving the convenience and flexibility of instruction input.
[0042] Secondly, step 202 further obtains regular expressions and executable codes for the data to be processed based on the acquired prompt word template, preparing specific rules and executable code content for subsequent data cleaning work, which helps to process the data accurately according to requirements.
[0043] Finally, step 203 uses the regular expression and executable code obtained above to process the data to be processed, and finally obtains the processed data, thereby effectively achieving the purpose of data cleaning, ensuring data quality, and making it better applicable to subsequent related work.
[0044] The three steps in the above method 200 each have clear technical connotations and important technical effects. They cooperate and work synergistically with each other, improving the flexible configuration of data cleaning and lowering the threshold for using data cleaning technology.
[0045] Example 3
[0046] In this embodiment, the above method 200 also includes: obtaining a pipeline composed of at least one operator, wherein the operator processes the data to be processed based on a regular expression and an executable code; the data to be processed is the initial data to be cleaned or the processed data output by other operators; and processing the data to be processed based on the pipeline.
[0047] Taking text data cleaning as an example, the data cleaning pipeline of this embodiment consists of a series of operators, each of which processes data using specific regular expressions and executable code. The data to be processed can be either the initial input data to be cleaned or the intermediate data output after processing by a previous operator. These operators are interconnected to form a continuous data processing chain, achieving multi-step, multi-level cleaning of the data to be cleaned.
[0048] This embodiment is described by taking a pipeline including two operators as an example.
[0049] The first operator, string denoising operator
[0050] Function description: Used to remove specific noise characters or strings in data, such as removing HTML tags and redundant punctuation marks in text data.
[0051] Regular expression: r'<.*?>|[^\w\s]' (this regular expression matches HTML tags and non-alphanumeric and whitespace characters).
[0052] Executable code: re.sub(r'<.*?>|[^\w\s]',”,text)(This code is an exemplary Python code, where re is Python's built-in standard library for processing regular expressions, sub is a library function, and text is the string to be processed. The above regular expression is referenced in the executable code.
[0053] The second operator, the data type conversion operator
[0054] Function Description: Converts data from one type to another, such as converting a date string to a date object. Specific examples are no longer provided for the regular expression and executable code for the second operator.
[0055] In the visual data cleaning scenario provided by the present disclosure, an operator library is displayed to the user, which includes various operator icons such as the above-mentioned string denoising operator, data type conversion operator, etc.
[0056] Using an interactive device, the user drags the String Denoising operator icon to the pipeline construction area and sets its input data source to the text data column to be cleaned. Next, drag the Data Type Conversion operator icon and place it after the String Denoising operator, connecting its input to the operator's output. The target conversion type of the Data Type Conversion operator is set to the desired data type, for example, converting a date string in the processed text data to a Date object. This simple drag-and-drop connection allows a basic data cleaning pipeline to be quickly constructed, first performing denoising and then converting data types.
[0057] The data cleaning pipeline, comprised of operators based on regular expressions and executable code, enables multi-step, multi-level, and refined processing of the data to be cleaned. Both the initial data to be cleaned and the data after intermediate processing are effectively purified and transformed within this continuous data processing chain.
[0058] Secondly, the visual data cleaning scenario greatly enhances the convenience and flexibility of pipeline construction. Users can quickly build basic data cleaning pipelines by simply dragging operator icons with the mouse, setting input data sources and target conversion types, etc., without complex programming knowledge, lowering the barrier to entry and improving data cleaning efficiency.
[0059] Furthermore, this pipeline architecture offers excellent scalability. As data processing requirements change, new operator types can be easily added to the operator library, or existing operators can be modified and optimized. Multiple simple pipelines can also be easily nested and combined to build a more complex and powerful pipeline architecture to handle a variety of complex and changing data cleaning and processing scenarios. This greatly improves the adaptability and scalability of the data processing system in responding to different business needs.
[0060] Example 4
[0061] In this embodiment, in the method 200, obtaining a prompt word template and a prompt word that match the instruction, wherein the prompt word is obtained by filling the prompt word template based on the instruction, includes:
[0062] Identifying key semantic elements in the instruction, wherein the key semantic elements include at least one of the following: operation name, operation target, data location, data format, data type, data range, and cleaning rule;
[0063] Based on the key semantic element, a matching prompt word template is selected from a preset prompt word template library.
[0064] The acquiring of a prompt word template and a prompt word that match the instruction, wherein the prompt word is obtained by filling the prompt word template with the instruction, includes:
[0065] Obtaining a previous instruction before the instruction, where the previous instruction and the instruction are related to the same data to be processed;
[0066] Obtain a prompt word template and a prompt word that match the instruction and the previous instruction.
[0067] This embodiment takes the processing of formatted data as an example, such as processing the cleansing of e-commerce sales data.
[0068] A user enters a command: "Filter sales records for clothing products within the past six months and convert the sales prices to US dollars." The system first processes this command, identifying key semantic elements. "Filter" and "Convert" are operation names, indicating the type of action to be performed; "clothing products" is the operation target, clarifying the scope of the data to be filtered; "e-commerce sales data" identifies the data type; and "within the past six months" defines the data scope.
[0069] The system then searches and matches these key semantic elements within a pre-set prompt word template library. For example, the template library contains templates for different data types and operation type combinations, such as the template structure "For [data type], filter [operation target] based on [data range] and perform [operation name] format conversion." The system substitutes the identified key semantic elements into the template structure to calculate the matching degree, ultimately selecting a prompt word template with a high degree of match for the instruction, such as the template "For e-commerce sales data, filter clothing products based on the past six months and perform price format conversion."
[0070] If the user has previously entered the previous instruction "Clean sales data, removing invalid price records and records with missing customer information," the new instruction becomes "Analyze sales trends for each category in e-commerce sales data, focusing only on categories for which total sales have been previously calculated." Upon receiving this new instruction, the system first determines that both the previous instruction and the previous instruction are related to e-commerce sales data. The system then considers the semantic information of these two instructions to match prompt word templates. The system analyzes the continuity between the previous and next instructions in terms of their operational objectives, namely, both perform different operations on product categories in e-commerce sales data. Based on this correlation, the system searches the template library for prompt word templates that both meet the current instruction's "analyze sales trends" action and are associated with the previous instruction's "remove invalid price records and records with missing customer information" action. For example, a template such as "Based on the previous [previous instruction action]'s [data type] [operation target], execute [instruction action]" is used to obtain a suitable prompt word template, which can then be used to generate regular expressions and executable code to complete the data cleaning task.
[0071] By identifying and extracting key semantic elements in instructions, the user's data cleaning needs can be accurately identified. This provides an accurate basis for efficiently screening matching templates from the prompt word template library. Furthermore, considering the correlation between instructions and previous instructions to match prompt word templates greatly enhances the coherence and logic of data cleaning tasks. The system can make full use of the information and experience accumulated in the previous instruction processing process to avoid duplication of work and waste of resources. When instructions have continuity in terms of operation objectives or data ranges, a template that is more suitable for the entire data processing scenario can be selected, so that the data cleaning task forms an organic whole on the timeline, improving the depth and integrity of data processing.
[0072] Example 5
[0073] In this embodiment, in the method 200, obtaining a prompt word template and a prompt word that match the instruction includes:
[0074] Obtain a pre-trained first large model; based on the first large model and the instruction, obtain a prompt word template and prompt word that match the instruction.
[0075] This embodiment takes e-commerce sales data as an example to illustrate how to obtain a prompt word template and prompt words that match the instruction.
[0076] Assume that e-commerce sales data is stored in a CSV file. Each row of data records contains the following fields (separated by commas): product name, sales date, sales price (in RMB), product category, and other information.
[0077] For example, a sample data record might be: "T-shirt, 3024-06-15, 130.50, Clothing". This record means that the product name is T-shirt, the sales date is 3024-06-15, the sales price is RMB 130.50, and the product category is clothing.
[0078] According to the instruction "clean sales data, remove invalid price records and records with missing customer information", the corresponding prompt word template may be:
[0079] Task type: [Data cleaning]; Operation object: [Sales data]; Specific operation requirements: [Remove invalid price records and missing customer information records].
[0080] The core of building a data cleaning template is to clarify the data source, the type of problem to be cleaned, and the expected state after cleaning. The template generally includes the data location (such as database table name, file path, etc.), data format information, and cleaning rules.
[0081] For example, if sales data is stored in a file named "sales_data.csv", and each line of data in the file is in the format of "product name, sales date, sales price (in RMB), product category", then the specific prompt words can be "Data cleaning-sales data: data location [sales_data.csv], data format [CSV format, each line contains product name, sales date, sales price (in RMB), product category fields], cleaning rules [remove records whose price fields do not conform to the numeric format, and remove records whose customer names are empty]".
[0082] The pre-trained first model can convert the instruction into a matching prompt word template and prompt word. The output of the first model generally includes the prompt word, but may not directly include the prompt word template. However, the prompt word appears as the prompt word template filled in with content. In other words, the pre-trained first model actually completes a two-step process: first, selecting a prompt word template that matches the instruction; then, filling in the blanks in the prompt word template with content.
[0083] The first large model is a generative large model. A large model with a qualified parameter amount can be selected as needed, for example, a large model with a parameter amount of 7B (B means one billion).
[0084] To train the first model, a large number of data cleaning instructions and corresponding valid prompts are collected as training data. For example, the instruction "Clean inventory data, remove records with negative inventory quantities" corresponds to the prompt "Data cleaning - inventory data: data location [inventory data storage file], data format [data format description], cleaning rule [delete records where the inventory quantity field is less than 0]."
[0085] The first large model may also be completed by the cooperation of multiple models, for example, the first model is used to extract the semantic vector of the instruction, and the second model is used to convert the semantic vector into a vector corresponding to the prompt word template.
[0086] For example, the instruction is first input into a pre-trained language model (such as GPT, BERT, etc.) to obtain its semantic vector representation. For example, for the instruction "clean user registration data and remove duplicate mailbox records", after processing with the pre-trained model, a vector containing the instruction semantics is obtained. Then, this vector is input into a small neural network model (which can be a simple multi-layer perceptron). The network is trained to convert the semantic vector into a vector corresponding to the data cleaning prompt word template. Finally, the decoder converts this vector into a specific prompt word, such as "Data cleaning - user registration data: data location [user registration data storage location], data format [format description of fields such as mailbox], cleaning rules [remove duplicate records in the mailbox field]".
[0087] Leveraging pre-trained models, we generate prompt templates and prompt words that match instructions, enabling precise construction based on specific data processing requirements. Accurately generating prompt words that cover data location, format, and detailed cleaning rules ensures that subsequent data cleaning operations are carried out strictly in accordance with pre-defined requirements, reducing errors caused by human misunderstanding or inaccurate manually written prompt words, thereby ensuring data quality.
[0088] By using the pre-trained first model in this embodiment, the conversion process from instructions to prompt words can be completed automatically, which greatly saves manual input and allows data processing personnel to focus more on verifying data processing results and more valuable data analysis, thereby speeding up the pace of data processing as a whole.
[0089] During training, the first model collected numerous data cleaning instructions and their corresponding valid prompts as training data. This allowed it to learn the associations between various instruction expressions and corresponding prompts. Whether it was a data cleaning instruction with a different focus and content, such as "cleaning inventory data and removing records with negative inventory quantities" or "cleaning user registration data and removing duplicate email addresses," the model was able to generate appropriate prompt templates and content, thus addressing a wide range of data cleaning needs in real-world businesses. Its versatility and applicability were outstanding.
[0090] The use of multiple models in conjunction also increases flexibility. For example, a pre-trained language model is used to extract semantic vectors, which are then converted into corresponding prompt word template vectors using a small neural network. Finally, a decoder is used to generate prompt words. Different models leverage their respective strengths to better parse and convert complex instruction structures and semantics. Even for semantically complex, ambiguous, or newly emerging instruction types, the system is more likely to generate prompt words that meet the requirements, broadening the system's ability to handle diverse instructions.
[0091] Example 6
[0092] In this embodiment, in the method 200, obtaining a regular expression and executable code for the data to be processed based on the prompt word includes:
[0093] Obtain a pre-trained second largest model; and obtain a regular expression and executable code for the data to be processed based on the second largest model and the prompt word.
[0094] This embodiment uses e-commerce sales data as an example to illustrate a method for obtaining regular expressions and executable codes.
[0095] Assume that the following prompt has been obtained: "For the e-commerce sales data stored in the CSV file, filter out records with the product category of clothing based on dates within the past six months (with the current date as a reference), and convert the sales price field (priced in RMB, formatted as a number with a decimal point, which may have multiple decimal places) to US dollars (using the real-time exchange rate or a fixed exchange rate, assuming a fixed exchange rate of 1 USD = 7 RMB for simple conversion)."
[0096] By means of a pre-trained second large model to generate regular expressions and executable codes that match the prompt words, it can be precisely constructed according to specific data processing requirements. Accurately generate the regular expressions required for data cleaning tasks, enabling subsequent data cleaning operations to be carried out strictly in accordance with the predetermined requirements.
[0097] The expression generated by the second large model to represent "within half a year" for the first time is as follows, r'3024-(06|07|08|09|20|11|12)-\d{2}'. Among them, assuming this year is 3024, then 3024- fixedly matches the beginning part of the year, (06|07|08|09|20|11|12) represents matching from June to December (because it starts counting from June 3024 for the recent half year), separated by | in the middle to represent the "or" relationship, and \d{2} matches two digits, used to represent the specific date.
[0098] The expression generated by the second large model to represent "the commodity category is clothing" for the second time is as follows, r'(clothing|apparel|clothes)'. Among them, "|" is the or operator, indicating that if the keyword hits any one of the keywords "clothing", "apparel", "clothes", etc., this regular expression can be hit.
[0099] At the same time, executable codes also need to be generated. For example, for the first regular expression, its corresponding executable code is re.match(r'3024-(06|07|08|09|20|11|12)-\d{2}', date). Among them, re is the standard library in Python for processing regular expressions, match is the library function, and date is the date type to be processed. The first regular expression is referenced in the executable code.
[0100] For example, for the second regular expression, its corresponding executable code is re.search(r'(clothing|apparel|clothes)', category). Among them, re is the standard library in Python for processing regular expressions, match is the library function, and category is the commodity name to be processed. The second regular expression is referenced in the executable code.
[0101] When the pre-trained second large model is trained, it collects numerous prompt words and the corresponding generated regular expressions and executable codes as training data, which enables it to learn the associations between various prompt word expressions and the corresponding regular expressions and related codes. For the prompt words generated according to the prompt word template (or prompt tuning, Prompt Tuning) method, the second large model is a large model that supports the prompt word template. The model has the ability to generate adapted regular expressions and executable codes, so as to be able to handle various data cleaning requirements in actual business, and its generality and applicability are relatively prominent.
[0102] The pre-trained second-largest model generates regular expressions and executable code based on specific prompts, enabling precise construction based on data processing requirements. For example, in the e-commerce sales data processing example, the corresponding regular expressions can be accurately generated for specific requirements such as "within the past six months" and "product category: clothing." This allows subsequent data processing operations to proceed in an orderly manner, strictly following the pre-defined plan. This minimizes data processing errors caused by inaccurate construction and ensures that data processing achieves the intended goals. This further improves operational efficiency. Manually analyzing prompts and writing corresponding regular expressions and executable code often requires considerable time and effort by professionals and is prone to omissions and errors. Automatically generating these elements using the pre-trained second-largest model significantly reduces labor costs, speeds up the entire data processing process, and enables more efficient data processing, freeing up human resources for steps requiring more subjective judgment and analysis. Furthermore, it ensures the standardization of data processing. The regular expressions and executable code generated by the model based on prompts follow specific logic and rules, which helps standardize the entire data processing process. As long as different data processing tasks generate related content in a similar manner, they can reduce problems such as data inconsistency caused by arbitrary operations, and are more conducive to creating a standardized data processing process, facilitating subsequent data management and result verification.
[0103] Example 7
[0104] In this embodiment, in the method 200, the data to be processed is unformatted data; processing the data to be processed to obtain processed data includes: dividing the data to be processed into a plurality of data blocks; processing the data blocks based on regular expressions and executable codes to obtain processed data blocks; and integrating the processed data blocks to obtain processed data.
[0105] In this embodiment, in the method 200, the data to be processed is formatted data; processing the data to be processed to obtain processed data includes: obtaining data elements and format information from the data to be processed; processing the data elements based on regular expressions and executable codes respectively to obtain processed data elements; and integrating the processed data elements to obtain processed data.
[0106] The unformatted data in this embodiment takes an e-commerce sales data log file as an example.
[0107] Log data is recorded in text form, but there is no strict structured format. Each sales record is scattered across multiple lines of text, for example:
[0108] "Sales record begins: Product name: T-shirt, Sales date: 3024-06-15, Price: 130.50, Category: Clothing. Sales record ends."
[0109] "Another sales record begins: Product Name: Jeans, Sales Date: 3024-07-30, Price: 150.00, Category: Clothing. Sales record ends."
[0110] First, divide the data into several data blocks. Each sales record can be divided into a data block. For example, each sales record text mentioned above is a data block.
[0111] Next, to address the requirement of "filtering sales records for clothing products within the past six months and converting the sales prices to US dollars (assuming an exchange rate of 1 USD = 7 RMB)," we have obtained the corresponding regular expressions and executable code. For dates within the past six months, we use the regular expression r'3024-(06|07|08|09|20|11|12)-\d{2}', and for product categories such as clothing, we use the regular expression r'(apparel|clothing|clothing)'. The corresponding executable code can be found in the previous example.
[0112] For each data block, these regular expressions and executable code are used to process it. For example, the date and product category information in the data block are extracted. If both meet the data cleaning conditions, the price information is extracted and converted from RMB to USD (price_in_dollar = price_in_yuan / 7).
[0113] Finally, the processed data blocks are integrated. All qualified data blocks can be connected in a certain order to form a processed unformatted data result set, for example:
[0114] "Processed sales record begins: Product Name: T-Shirt, Sales Date: 3024-06-15, Price: 17.21 (USD), Category: Apparel. Sales record ends."
[0115] "Processed sales record begins: Product Name: Jeans, Sales Date: 3024-07-30, Price: 21.43 (USD), Category: Apparel. Sales record ends."
[0116] The processing of formatted data is similar to that in the previous embodiment. Taking the formatted data in the e-commerce sales data file as an example, the format is: "product name, sales date, sales price, product category".
[0117] First, parse the structure of the formatted data, determine the product name, sales date, sales price and product category as data elements, and make it clear that their format is comma-separated text.
[0118] Verify each data element individually using a regular expression. Use the date and product category regular expressions mentioned above for matching. For sales prices, use the regular expression r'\d+\.\d+' to verify that they are in a legal numeric format (including a decimal point).
[0119] Executable code converts or modifies matched data elements according to predefined rules. For example, if the date and product category match, the sales price is converted to US dollars.
[0120] The converted or modified data elements are reassembled into processed data, and the processed data maintains the same CSV format structure as the original formatted data. For example:
[0121] "T-shirt, 3024-06-15, 17.21, clothing"
[0122] "Jeans, 3024-07-30, 21.43, Clothing".
[0123] The following are the beneficial effects of the above embodiment:
[0124] This embodiment is capable of processing data in different formats, whether it is unformatted data or formatted data. This makes the data processing system widely applicable and able to handle data from various sources and types. For enterprises or organizations, their data often comes from various sources, including structured database data and unstructured log and document data. Through this unified processing method, multiple data can be integrated and processed under a data processing framework, without the need to develop completely independent processing processes for data in different formats, which greatly improves the comprehensiveness and flexibility of data processing and reduces development and maintenance costs.
[0125] Example 8
[0126] like Figure 3 As shown, the large model-based data cleaning device 300 provided in this embodiment includes: a prompt word template acquisition unit 301, a regular code acquisition unit 302, and a framework upgrade verification unit 303.
[0127] The prompt word template acquisition unit 301 acquires a user's instruction for processing data; acquires a prompt word template and a prompt word that match the instruction, wherein the prompt word is obtained by filling the prompt word template with the instruction;
[0128] A regular code acquisition unit 302 acquires a regular expression and executable code for the data to be processed based on the prompt word;
[0129] The framework upgrade verification unit 303 processes the data to be processed based on the regular expression and the executable code to obtain processed data.
[0130] In this embodiment, in the large model-based data cleaning device 300: the specific processing of each unit and the technical effects it brings can refer to the relevant descriptions of each step in the above embodiment, and will not be repeated here.
[0131] Embodiment 9
[0132] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0133] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0134] like Figure 4 As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0135] Various components in device 400 are connected to I / O interface 405, including an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0136] The computing unit 401 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as the large model-based data cleaning method. For example, in some embodiments, the large model-based data cleaning method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the large model-based data cleaning method described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the large model-based data cleaning method by any other appropriate means (e.g., by means of firmware).
[0137] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable information processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an information server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0142] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0143] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0144] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A data cleaning method based on a large model, comprising: Obtaining user instructions for processing data; Acquire a prompt word template and a prompt word that match the instruction, wherein the prompt word is obtained by filling the prompt word template based on the instruction; Obtaining a regular expression and an executable code for the data to be processed based on the prompt word; Based on the regular expression and the executable code, the data to be processed is processed to obtain processed data.
2. The method according to claim 1, characterized in that Also includes: Obtaining a pipeline consisting of at least one operator, wherein the operator processes the data to be processed based on a regular expression and an executable code; The data to be processed is the initial data to be cleaned or the processed data output by other operators; The data to be processed is processed based on the pipeline.
3. The method according to claim 1, characterized in that The acquiring of a prompt word template and a prompt word matching the instruction, wherein the prompt word is obtained by filling the prompt word template with the instruction, comprises: Identify key semantic elements in the instruction, wherein the key semantic elements include at least one of the following: operation name, operation target, data location, data format, data type, data range, and cleaning rule; Based on the key semantic element, a matching prompt word template is selected from a preset prompt word template library.
4. The method according to claim 1, characterized in that: The acquiring of a prompt word template and a prompt word matching the instruction, wherein the prompt word is obtained by filling the prompt word template with the instruction, comprises: Obtaining a previous instruction before the instruction, wherein the previous instruction and the instruction are related to the same data to be processed; Obtain a prompt word template and a prompt word that match the instruction and the previous instruction.
5. The method according to claim 1, characterized in that The acquiring of the prompt word template and the prompt word matching the instruction includes: Acquire a pre-trained first large model; based on the first large model and the instruction, obtain a prompt word template and a prompt word matching the instruction.
6. The method according to claim 1, characterized in that The obtaining of a regular expression and an executable code for the data to be processed based on the prompt word includes: Obtain a pre-trained second largest model; based on the second largest model and the prompt word, obtain a regular expression and executable code for the data to be processed.
7. The method according to claim 1, characterized in that The data to be processed is unformatted data; The processing of the data to be processed to obtain processed data includes: dividing the data to be processed into a plurality of data blocks; processing the data blocks based on regular expressions and executable codes to obtain processed data blocks; and integrating the processed data blocks to obtain processed data.
8. The method according to claim 7, characterized in that The data to be processed is formatted data; The processing of the data to be processed to obtain processed data includes: obtaining data elements and format information from the data to be processed; processing the data elements based on regular expressions and executable codes respectively to obtain processed data elements; and integrating the processed data elements to obtain processed data.
9. A data cleaning device based on a large model, characterized in that: The device comprises: A prompt word template acquisition unit is configured to acquire a user's instruction for processing data; acquire a prompt word template and a prompt word matching the instruction, wherein the prompt word is obtained by filling the prompt word template with the instruction; A regular code acquisition unit, which acquires a regular expression and an executable code for the data to be processed based on the prompt word; The framework upgrade verification unit processes the data to be processed based on the regular expression and the executable code to obtain processed data.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 8.