Text data construction method based on dialogue type large language model and medium

By adopting a text data construction method based on a dialogue-based large language model in the field of natural language processing, using cleaning propt and converting propt, the problem of high cost and time overhead of building large-scale text data sets is solved, and efficient and low-cost text data sets are realized. The generated data sets are of high quality and are suitable for a variety of NLP tasks.

CN119990070APending Publication Date: 2025-05-13CAS SUZHOU INSTITUTE OF INTELLIGENT COMPUTING TECHNOLOGY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062899.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology has high cost and time overhead problems in building large-scale text data sets, and the data quality is different, requiring additional manual cleaning and processing. Data augmentation technology cannot effectively provide high-quality and diverse text data.

Method used

Using a text data construction method based on a dialogue-based large language model, we design and clean the prompt and convert the prompt, and use the large language model to preprocess, clean and transform the data, and generate and organize high-quality text data sets.

Benefits of technology

It significantly improves the efficiency of data sorting and construction, reduces costs, and reduces dependence on professional knowledge. The generated text data sets are of high quality and rich diversity, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990070A_ABST
    Figure CN119990070A_ABST
Patent Text Reader

Abstract

The invention discloses a text data construction method and medium based on a dialogue type large language model.The method comprises the steps that according to the requirement of natural language processing, multi-style text data is obtained to serve as first data, and the first data forms a first data set; preprocessing all the first data to form second data, and forming a second data set by the preprocessed second data; inserting each piece of second data into a set cleaning prompt, and inputting the cleaning prompt into the large language model so as to carry out instruction evaluation and label endowing on the second data; filtering the labels to filter out part of the second data, and forming a filtered third data set; second data in the third data set are inserted into a set conversion prompt, the conversion prompt is input into the large language model, and the second data are converted into text data meeting the natural language processing requirement through the large language model. Manual processing is not needed, efficiency is improved, cost is reduced, and dependence on professional knowledge is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a text data construction method and medium based on a conversational large language model. Background Art

[0002] In the field of natural language processing (NLP), large-scale, high-quality text datasets are the key to the subsequent development and training of powerful algorithms. Traditionally, these large-scale text datasets are collected through a large amount of manual collection or questionnaires and other data collection sources, and then constructed through manual annotation. This process is not only time-consuming and labor-intensive, but also costly. For example, building a dataset for training and fine-tuning NLP models requires a lot of manual collection, and then a lot of manual work to organize, construct, and finally verify the dataset. However, NLP field models usually require millions of data to meet the needs of model training and fine-tuning. If the data is cleaned and constructed manually, a lot of resources are required.

[0003] In recent years, large language models such as OpenAI's GPT3.5 and Meta's LLAMA have achieved remarkable success in NLP tasks. These models can understand and generate natural language text, providing new possibilities for building text datasets.

[0004] However, existing technologies have several major problems in building large-scale text datasets: 1) The cost and time of manual data cleaning and labeling are high; 2) The quality of the collected data varies, requiring additional manual cleaning and processing; 3) Data augmentation technology cannot effectively provide high-quality and diverse text data. Summary of the invention

[0005] In order to overcome the above-mentioned shortcomings, the purpose of the present invention is to provide a text data construction method and medium based on a conversational large language model, which uses prompt to adjust the input and output of the large language model to generate and clean text data, significantly improving the efficiency of data sorting and data construction, reducing costs, and reducing dependence on professional knowledge.

[0006] In order to achieve the above objectives, the technical solution adopted by the present invention is: a text data construction method based on a conversational large language model, comprising: According to the requirements of natural language processing, obtaining multi-style text data as first data, wherein the first data forms a first data set; Preprocessing all the first data to form second data, wherein the preprocessed second data forms a second data set; Insert each of the second data into a set cleaning prompt, and input the cleaning prompt into a large language model to perform instruction evaluation and assign labels to the second data; Filtering the tags to filter out the portion of the second data and forming a filtered third data set; The second data in the third data set is inserted into a set conversion prompt, and the conversion prompt is input into the large language model, and the large language model converts the second data into text data that meets the natural language processing requirements. The beneficial effects of the present invention are: Combining diverse text data from a wide range of data sources with the data style and structure conversion capabilities of a conversational large language model can build a diverse text dataset and enhance the model's generalization capabilities. The designed cleaning prompt can effectively identify and filter out low-quality data, thereby significantly improving the accuracy and reliability of the final data set and providing higher-quality training and test data for NLP tasks. It reduces the need for manual intervention, saves a lot of time and human resources, and improves the efficiency of data processing.

[0007] The cleaning prompt and conversion prompt are set according to the NLP task, which can flexibly adapt to various NLP task requirements and has high applicability and scalability.

[0008] Without the need for additional model training, the existing large language model can be used to implement question generation tasks of various specifications.

[0009] Specifically, preprocessing all the first data to form second data, and forming the second data set from the preprocessed second data includes: Deleting blanks, emoticons and garbled characters in each of the first data; The first data after removing blanks, emoticons and garbled characters is stored line by line in a list to form a second data set.

[0010] This will remove interference items and initially unify the data format.

[0011] Specifically, inserting each second data into a set cleaning prompt, and inputting the cleaning prompt inserted into the second data into a large language model to perform instruction evaluation and assign a label to the second data specifically includes: Insert the second data line by line into the set cleaning prompt; The cleaning prompt is input into a large language model, and the large language model assigns a first label to the cleaning prompt input that meets the quality requirement, and assigns a second label to the cleaning prompt input that does not meet the quality requirement.

[0012] According to the requirements of natural language processing, the second data is labeled by cleaning prompts to quickly distinguish the second data that meets the quality requirements and does not meet the quality requirements.

[0013] Specifically, filtering the tag to filter out the portion of the second data and forming a filtered third data set includes: The second data corresponding to the second tag is deleted, and the remaining second data is stored row by row to form the third data set.

[0014] Use tags to quickly filter out unnecessary data.

[0015] Furthermore, the conversion prompt includes a data style conversion prompt and a data structure conversion prompt, and the data style conversion prompt and the data structure conversion prompt are sequentially input into the large language model.

[0016] Specifically, inserting the second data in the third data set into a set conversion prompt, inputting the conversion prompt into a large language model, and converting the second data into text data that meets the natural language processing requirements by the large language model specifically includes: The second data in the third data set is inserted into the data style conversion prompt, the data style conversion prompt is input into a large language model, and the large language model outputs third data; The third data is inserted into the data structure conversion prompt, the data structure conversion prompt inputs a large language model, and the large language model outputs the text data.

[0017] Through the data style conversion prompt and data structure conversion prompt designed according to the NLP task, the second data in the third data set formed after filtering can be converted into text data with a format that meets the requirements to form a text data set.

[0018] Furthermore, the large language model uses a few-shot method to constrain the large language model before processing the data style conversion prompt and the data structure conversion prompt. Prompt writing of the few-shot method is the simplest and most efficient method for constraining the large language model to be generated according to the specified content.

[0019] Furthermore, the requirements for natural language processing include the type, style, structure and quality standards of the required text data analyzed manually, and the scale, diversity requirements and any specific language or domain restrictions of the first data set are determined based on the requirements for natural language processing.

[0020] Specifically, the spaces, emoticons and garbled characters in each of the first data are deleted by using a regular expression of the Python language.

[0021] The present invention also discloses a computer-readable storage medium, on which a text data construction method program is stored. When the text data construction method program is executed by a processor, the steps of the above-mentioned text data construction method are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The process of the text data construction method in one embodiment of the present invention is as follows Figure 1 ; Figure 2 The process of the text data construction method in one embodiment of the present invention is as follows Figure 2 ; Figure 3 The process of the text data construction method in one embodiment of the present invention is as follows Figure 3 ; Figure 4 The process of the text data construction method in one embodiment of the present invention is as follows Figure 4 . DETAILED DESCRIPTION

[0023] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0024] The present invention discloses a text data construction method based on a conversational large language model, which can automatically generate high-quality, diversified text data that meets specific requirements, and can batch clean text data from multiple sources.

[0025] See attached Figure 1 As shown, the text data construction method includes: S1. According to the requirements of natural language processing, obtain multi-style text data as first data, and the first data forms a first data set.

[0026] The requirements for natural language processing include the type, style, structure, and quality standards of the required text data analyzed manually, and the size, diversity requirements, and any specific language or domain restrictions of the first data set are determined based on the requirements for natural language processing. These requirements will be used to set the subsequent conversion prompt.

[0027] In different NLP tasks, the scale and language or domain restrictions of the first data set are different. For example, in the task of news media entity name recognition, the input data style is restricted and data from multiple sources needs to be converted into news article format.

[0028] The first data includes diverse text data such as PDF and e-books, which can be purchased through third-party channels or downloaded directly.

[0029] S2. Preprocess all first data to form second data, and the preprocessed second data forms a second data set.

[0030] The preprocessing is to delete redundant symbols of the second data, and arrange and store the second data according to specified requirements.

[0031] In one embodiment, see the attached Figure 2 As shown, step S2 specifically includes: S21. Delete the blanks, emoticons and garbled characters in each first data.

[0032] At this time, the steps are completed by using the regular expression of the Python language, and redundant symbols in the first data can be deleted.

[0033] S22. Store the first data after deleting spaces, emoticons, and garbled characters in a list line by line to form a second data set.

[0034] The first data is separated by the "\n" line break symbol and stored in a list using the Python language.

[0035] S3. Insert each second data into the set cleaning prompt, and input the cleaning prompt into the large language model to perform instruction evaluation and assign labels to the second data.

[0036] The cleaning prompt is designed according to the requirements of natural language processing. Before the cleaning prompt is input into the large language model, a second piece of data has been inserted.

[0037] See attached Figure 3 As shown, step S3 specifically includes: S31, inserting the second data line by line into the set cleaning prompt.

[0038] S32: Input the clean prompt inserted into the second data into the large language model. The large language model assigns a first label to the clean prompt input that meets the quality requirement, and assigns a second label to the clean prompt input that does not meet the quality requirement.

[0039] For example, the first label is 1 and the second label is 0. When the label corresponding to a piece of second data is 1, it means that the second data meets the quality requirements. When the label corresponding to a piece of second data is 0, it means that the second data does not meet the quality requirements. Of course, the first label can also be 0 and the second label can also be 1. They can be set as needed, but the first label and the second label need to be different to facilitate subsequent filtering.

[0040] For example, the cleaning prompt is "You are a data mining expert. You only need to determine whether the input data belongs to the relevant content of patent writing. If it does, you need to answer 1, otherwise answer 0: [second data]". At this time, the cleaning prompt is input into the large language model, and the large language model will output the label category (0 or 1). The large language model can be GPT-4, which uses the dialogue model of the prompt. There is no need to train the large language model in advance, and it can be called directly.

[0041] S4. Filter the tags to filter out part of the second data, and form a filtered third data set.

[0042] Through step S3, the second data that do not meet the quality requirements have been assigned a label of 0, and these second data do not meet the requirements and therefore need to be filtered out. The labels are then filtered through a label filter written in Python.

[0043] The second data corresponding to the second tag is deleted, and the remaining second data is stored line by line to form a third data set. At this time, only the second data that meets the quality requirements is stored in the third data set.

[0044] Steps S3-S4, through the designed cleaning prompt, the present invention can effectively identify and filter out low-quality data, thereby significantly improving the accuracy and reliability of the final data set, providing higher-quality training and test data for NLP tasks, reducing the need for manual intervention, saving a lot of time and human resources, and improving the efficiency of data processing.

[0045] S5. Insert the second data in the third data set into the set conversion prompt and input it into the large language model, and the large language model converts the second data into text data.

[0046] The conversion prompt input into the large language model has inserted the second data with the required quality. Through the cooperation between the conversion prompt and the large language model, the large language model outputs text data that meets the needs of natural language processing and can meet the requirements of specific NLP tasks.

[0047] In this example, a wide range of data sets downloaded or purchased from third-party channels , combined with the data style and structure conversion capabilities of the conversational text model, the present invention can construct a diverse text data set and enhance the generalization ability of the model. Through the designed cleaning prompt, the present embodiment can effectively identify and filter out low-quality data, thereby significantly improving the accuracy and reliability of the final data set, and providing higher quality training and test data for NLP tasks. The automated data cleaning and construction process reduces the need for manual intervention, saves a lot of time and human resources, and improves the efficiency of data processing. The cleaning prompt and conversion prompt are set according to the NLP task, and can flexibly adapt to various NLP task requirements, with high applicability and scalability. At the same time, in this embodiment, various specifications of question generation tasks can be implemented without additional model training.

[0048] In one embodiment, the conversion prompt includes a data style conversion prompt and a data structure conversion prompt, and the data style conversion prompt and the data structure conversion prompt are sequentially input into the large language model.

[0049] The data style conversion prompt is used to convert the style of the second data. For example, the data style conversion prompt is "You are an expert in data style modification. Next, you only need to modify the data I input into news style type data output: [second data]".

[0050] The data structure conversion prompt is used to convert the second data with converted data style into a unified data format. For example, the data structure conversion prompt is "You are a data structure modification expert, you only need to return the input data in the format of ner: [second data]".

[0051] See attached Figure 4 As shown, step S5 specifically includes: S51. The second data in the third data set is inserted into a data style conversion prompt, the data style conversion prompt is input into a large language model, and the large language model outputs third data.

[0052] S52: insert the third data into the data structure conversion prompt, input the data structure conversion prompt into the large language model, and the large language model outputs text data.

[0053] Through the data style conversion prompt and data structure conversion prompt designed according to the NLP task, the second data in the third data set formed after filtering can be converted into text data with a format that meets the requirements to form a text data set, which can be used for subsequent large language model fine-tuning, classification and machine translation.

[0054] In one embodiment, the large language model uses a few-shot method to constrain the large language model before processing the data style conversion prompt and the data structure conversion prompt.

[0055] Through the few-shot method, the few-shot method is a method that uses a small number of styles to tell the large model how to generate data. This method does not involve training a large model. This method does not require training any model, and can allow the model to generate content according to task requirements.

[0056] In one embodiment, the target NLP task is set as sentiment analysis, which requires a large amount of product review text data. In terms of data style, diverse product review texts are required. In terms of data labels, three different labels, positive, neutral, and negative, are designed and constructed for data annotation and construction for subsequent research and other work.

[0057] Use open source review datasets covering multiple product types, or text datasets obtained from third parties. Use regular expressions to preprocess the data, remove all HTML tags, blank characters, emoticons, and garbled characters, and ensure that each review is in plain text format. Store the preprocessed data in rows, with each row representing a product review, recorded as the second dataset. For example, the second dataset is: 1 The lighting effect of the keyboard is very cool, and the typing experience is very comfortable, exceeding expectations.

[0058] 2 After using this dishwasher, I really saved a lot of time and effort, and the dishes were cleaned very cleanly.

[0059] 3 The overall quality of the product is average, with some minor flaws in appearance, but its functions can basically meet my needs.

[0060] 4 The camera's photo-taking function is average, and I was a little disappointed with the night mode effect.

[0061] 5 The packaging is good, but the product doesn't have much highlights and is similar to other similar products on the market.

[0062] Design a data cleaning prompt, for example: "For the following comments, please determine whether they contain meaningful emotional expressions: [comment text]". Insert each second data in the second data set into the cleaning prompt and input it into the conversational large language model (such as GPT-3.5). According to the output of the large language model, mark those comments that the model determines as "no meaningful emotional expression" as 0. Use the label filter to remove all comments marked as 0 to obtain the cleaned third data set. The third data set is: 1 The lighting effect of the keyboard is very cool, and the typing experience is very comfortable, exceeding expectations.

[0063] 2 After using this dishwasher, I really saved a lot of time and effort, and the dishes were cleaned very cleanly.

[0064] 3 The overall quality of the product is average, with some minor flaws in appearance, but its functions can basically meet my needs.

[0065] 4 The camera's photo-taking function is average, and I was a little disappointed with the night mode effect.

[0066] The second data in Article 5 was deleted because it was judged not to have meaningful emotional expression.

[0067] Design a data style conversion prompt, for example: "Convert the following comments to other styles, but do not change the emotional color: [comment text]", which is used to convert the style of some comments and increase data diversity. At the same time, design a data structure to build a prompt, for example: "Annotate the data of the comments I entered for sentiment analysis and return the data structure: [comment text]" Insert each second data in the third data set into the data style conversion prompt and the data structure construction prompt, and input it into the large language model. Collect the converted text output by the model to ensure that each output conforms to the formal customer feedback format. Store the converted text as a text data set for sentiment analysis NLP tasks. The text data set is: 1 Text: The lighting effect of the keyboard is very cool, and the typing experience is very comfortable, exceeding expectations.

[0068] Emotion: Positive.

[0069] 2 Text: After using this dishwasher, I really saved a lot of time and effort, and the dishes were cleaned very cleanly.

[0070] Emotion: Positive.

[0071] 3 Text: The overall quality of the product is average, there are some minor flaws in the appearance, but the functions can basically meet my needs.

[0072] Emotion:Neutral.

[0073] 4 Text: The camera's photo-taking function is average, and I was a little disappointed with the night scene mode.

[0074] Emotion: Negative.

[0075] In one embodiment, the target NLP task is set as name entity recognition (NER) for social dynamics detection on social platforms. Considering the diversity of data sources, the data style needs to be unified, and the data labels are designed and constructed: the labels used by the NER task are directly used, such as B-PER: the beginning of a person's name, I-PER: the internal part of a person's name, B-ORG: the beginning of an organization name, I-ORG: the internal part of an organization name, B-LOC: the beginning of a location, etc.

[0076] Use a review dataset purchased from a third-party channel or open source, covering 1 million reviews of various types. Use regular expressions to preprocess the data, remove all HTML tags, blank characters, emoticons, and garbled characters, and ensure that each review is in plain text format. Store the preprocessed data in rows, with each row representing a text, and record it as the second dataset.

[0077] Design a data cleaning prompt to filter out text data that does not belong to public opinion, for example: "For the following content, please determine whether it belongs to social dynamic data: [comment text]".

[0078] Each piece of second data in the second dataset is inserted into the cleaning prompt and input into the large language model (e.g., GPT-3.5). Based on the output of the model, those texts that the model determines as "not belonging to social dynamic data" are marked as 0. Use the label filter to remove all comments marked as 0 to obtain the filtered third dataset.

[0079] Design a data style conversion prompt. Social dynamics detection requires consistent data style. Here, the data style is unified into the formal news style required by the task, for example: "Convert the following comment text style into a formal news style, but do not change the emotional color: [comment text]", which is used to convert the style of some comments and increase data diversity. At the same time, design a data structure to construct a prompt, for example: "Annotate the data of the comments I entered for sentiment analysis and return the data structure: [comment text]".

[0080] Each second data in the third data set is inserted into the data style conversion prompt and the data structure construction prompt, and then input into the large language model. Collect the converted text output by the model to ensure that each output conforms to the formal customer feedback format. Store the converted text data as a text data set for subsequent social dynamics detection on social platforms.

[0081] The present invention also discloses a computer-readable storage medium, on which a text data construction method program is stored. When the text data construction method program is executed by a processor, the steps of the above-mentioned text data construction method are implemented. Based on such understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of each of the above-mentioned method embodiments when executed by a processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0082] The above implementation modes are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with this technology to understand the content of the present invention and implement it, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the protection scope of the present invention.

Claims

1. A text data construction method based on a conversational large language model, characterized by: include: According to the requirements of natural language processing, obtaining multi-style text data as first data, wherein the first data forms a first data set; Preprocessing all the first data to form second data, wherein the preprocessed second data forms a second data set; Insert each of the second data into a set cleaning prompt, and input the cleaning prompt into a large language model to perform instruction evaluation and assign labels to the second data; Filtering the tags to filter out part of the second data and forming a filtered third data set; The second data in the third data set is inserted into a set conversion prompt, and the conversion prompt is input into the large language model, and the large language model converts the second data into text data that meets the natural language processing requirements.

2. The text data construction method based on the conversational large language model according to claim 1, characterized in that: Preprocessing all the first data to form second data, wherein the preprocessed second data forms a second data set specifically includes: Deleting blanks, emoticons and garbled characters in each of the first data; The first data after removing blanks, emoticons and garbled characters is stored line by line in a list to form a second data set.

3. The text data construction method based on the conversational large language model according to claim 1, characterized in that: Inserting each second data into a set cleaning prompt, and inputting the cleaning prompt inserted into the second data into a large language model to perform instruction evaluation and assign labels to the second data specifically includes: Insert the second data line by line into the set cleaning prompt; The cleaning prompt is input into a large language model, and the large language model assigns a first label to the cleaning prompt input that meets the quality requirement, and assigns a second label to the cleaning prompt input that does not meet the quality requirement.

4. The text data construction method based on the conversational large language model according to claim 3 is characterized in that: Filtering the tag to filter out the portion of the second data and forming a filtered third data set specifically includes: The second data corresponding to the second tag is deleted, and the remaining second data is stored row by row to form the third data set.

5. The text data construction method based on the conversational large language model according to claim 1, characterized in that: The conversion prompt includes a data style conversion prompt and a data structure conversion prompt, and the data style conversion prompt and the data structure conversion prompt are sequentially input into the large language model.

6. The text data construction method based on the conversational large language model according to claim 5 is characterized in that: Inserting the second data in the third data set into the set conversion prompt, inputting the conversion prompt into the large language model, and converting the second data into text data that meets the natural language processing requirements by the large language model specifically includes: The second data in the third data set is inserted into the data style conversion prompt, the data style conversion prompt is input into a large language model, and the large language model outputs third data; The third data is inserted into the data structure conversion prompt, the data structure conversion prompt inputs a large language model, and the large language model outputs the text data.

7. The text data construction method based on the conversational large language model according to claim 1, characterized in that: The large language model uses a few-shot method to constrain the large language model before processing the data style conversion prompt and the data structure conversion prompt.

8. The text data construction method based on the conversational large language model according to claim 1, characterized in that: The requirements for natural language processing include the type, style, structure and quality standards of the required text data analyzed manually, and the scale, diversity requirements and any specific language or domain restrictions of the first data set are determined based on the requirements for natural language processing.

9. The text data construction method based on the conversational large language model according to claim 2, characterized in that: The blanks, emoticons and garbled characters in each of the first data are deleted by using regular expressions in Python language.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores a text data construction method program, and when the text data construction method program is executed by a processor, the steps of the text data construction method according to any one of claims 1 to 9 are implemented.