A text cleaning method, device and storage medium based on a large language model
By constructing example texts with strong relevance to the text in the large language model as prompt words, the problem of resource consumption caused by manual text labeling errors is solved, and the accuracy of text labeling in the large language model is improved.
Patent Information
- Application Number
- CN202410997382.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2044-07-24
AI Technical Summary
Existing methods for manually judging text label errors consume a lot of human resources, and large language models lack effective prompt words when judging whether text labels are incorrect, resulting in insufficient accuracy.
By obtaining the primary tags of the text to be cleaned and matching them in a pre-set thought chain example text library, thought chain example texts that are highly relevant to the text are constructed as prompt words and input into a trained large language model to improve the accuracy of judgment.
It improves the accuracy of large language models in judging whether text labels are incorrect, reduces human intervention, and saves human resources.
Smart Images

Figure CN118689956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric digital data processing, in particular to a text cleaning method based on a large language model, a device and a storage medium. BACKGROUND
[0002] The purpose of cleaning the text set includes cleaning the text with the wrong label in the text set. How to determine whether the label of the text is wrong, the prior art usually adopts the manual judgment method, but the manual judgment method has the problem of occupying more manual resources.
[0003] Large language models (LLM) can be used to process various natural language tasks, and large language models play an increasingly important role in various applications. A large language model can perform a new task with a small amount of examples, that is, a small amount of examples related to the task to be performed by the large language model and the task to be performed by the large language model are spliced together as prompt words of the large language model, which can improve the reasoning ability of the large language model for the task to be performed, and improve the accuracy of the reasoning result of the large language model for the task to be performed. For example, the Chinese patent with publication number CN117076649A discloses an emergency information query method and device based on a large model thinking chain. The patent selects professional problem text and professional knowledge related to the user input question from the pre-established emergency scene library based on the user input question, and constructs prompt words for the large language model based on the professional problem text, professional knowledge and user input question, so that the large language model learns the thinking related to the user input question based on the part of the prompt words corresponding to the professional problem and professional knowledge, so that the large language model has reasoning ability related to the user input question, and further improves the accuracy of the large language model in answering the user input question.
[0004] Constructing prompt words related to the content to be reasoned is beneficial to improving the accuracy of the reasoning result of the large language model. When the large language model is applied to the task of determining whether the label of the text is wrong, how to construct prompt words related to the task, and further improve the accuracy of the large language model in determining whether the label of the text is wrong, is a problem that needs to be solved. SUMMARY
[0005] The purpose of the present application is to provide a text cleaning method based on a large language model, a device and a storage medium, to construct prompt words related to the task of determining whether the label of the text is wrong, and further improve the accuracy of the large language model in determining whether the label of the text is wrong.
[0006] According to the present application, a text cleaning method based on a large language model comprises the following steps:
[0007] R100, obtaining a text set cle to be cleaned, the text set cle including a plurality of texts to be cleaned.
[0008] R200, inputting an ra-th text to be cleaned into a trained first model, obtaining a first label cle ra of the ra-th text to be cleaned. l ab 1 ra The trained first model is used to obtain the first label of the input text, and the accuracy of the trained first model in obtaining the first label of the input text is greater than or equal to tf, where tf is a preset accuracy threshold; the value range of ra is 1 to to, and to is the number of texts to be cleaned included in cle.
[0009] R300, according to l ab 1 ra Matching in a preset thinking example text library bas, obtaining num thinking example texts mod matched with l ab 1 ra The num thinking example texts mod matched. ra The bas includes a plurality of thinking example text sub-libraries, different first labels correspond to different thinking example text sub-libraries, and each thinking example text sub-library includes a plurality of thinking example texts corresponding to corresponding second labels; num is the number of corresponding thinking example texts of cle ra .
[0010] R400, according to the second labels of mod ra , cle ra and cle ra , constructing corresponding prompt words pro l ab 2 ra . ra ra .
[0011] R500, inputting pro ra into a trained large language model, and determining whether to clean cle ra away according to the output of the trained large language model.
[0012] According to a second aspect of the present application, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the above-mentioned text cleaning method based on a large language model.
[0013] According to a third aspect of the present invention, a computer-readable storage medium stores a computer program that, when executed by a processor, implements the above-described text cleaning method based on a large language model.
[0014] The present invention has at least the following beneficial effects:
[0015] This invention applies to any text cle in the text set to be cleaned. ra First, it is input into a trained first model, which has the function of obtaining the label of the input text. Therefore, the label can be obtained through the trained first model. ra The first-level tag l ab 1 ra In this invention, the first model trained on the text has high accuracy in determining the first-level labels, and the first-level labels output by the first model can be directly used as the classification criteria. ra The first-level tag l ab 1 ra After obtaining cle ra The first-level tag l ab 1 ra After that, l ab 1 ra The matching is performed in the pre-defined thought chain example text library bas, which includes sub-libraries of thought chain example texts corresponding to different first-level tags. The first-level tags can also be obtained from bas. l ab 1 ra Example text of the thought chain; based on the first-level tags in bas, it is also... l ab 1 ra The example text for constructing thought chain prompts is based on the word chain. ra The example texts with strong correlations in the thought chain are used to construct prompt words, which is beneficial for trained large language models to learn and apply cognito words. ra Relevance in thinking can enhance the semantic analysis capabilities of trained large language models, thereby improving the accuracy of their output. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of the text cleaning method based on a large language model provided for Embodiment One of the present application is shown. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0019] Embodiment One
[0020] According to the present application, a text cleaning method based on a large language model is provided, as shown in Figure 1 , comprising:
[0021] R100, obtaining a text set cle to be cleaned, cle comprising a plurality of texts to be cleaned.
[0022] In this embodiment, cle={cle1, cle2, …, cle ra ,…,cle to}, cle ra is the ra-th text to be cleaned, and the value range of ra is 1 to to, and to is the number of texts to be cleaned.
[0023] R200, inputting the ra-th text cle ra to be cleaned into a trained first model to obtain a first-level label ra of cle l ab 1 ra ; the trained first model is used to obtain a first-level label of an input text, the accuracy of the trained first model in obtaining a first-level label of an input text is greater than or equal to tf, and tf is a preset accuracy threshold; the value range of ra is 1 to to, and to is the number of texts to be cleaned included in cle.
[0024] In this embodiment, the accuracy of the trained first model in inferring the first-level label of the text is high, and the output of the trained first model when cle ra is input is taken as the first-level label ra ab l of cle 1 . ra
[0025] In this embodiment, the trained first model is used to obtain the first-level label of the input text, the first-level label is a rough label, and the first-level label further includes a plurality of second-level labels, the second-level label is a refined label, and the accuracy of the second-level label of the text inferred by the trained first model is low (i.e., less than tf). Therefore, in this embodiment, only the first-level label of the text is obtained by using the trained first model, and the second-level label of the text is not obtained by using the trained first model.
[0026] In this embodiment, the first model is a neural network model, and a supervised training manner is used to train the first model. Optionally, the process of training the first model includes: obtaining a training text set, the training text set including a plurality of training texts; obtaining a training label set, the training label set including a first-level label corresponding to each training text; and training the first model according to the training text set and the training label set to obtain the trained first model.
[0027] R300, according to l ab 1 ra In the preset thinking model example text library bas, matching is performed to obtain the thinking model example text mod corresponding to the first-level label of the input text. l ab 1 ra The matched num thinking model example texts mod ra ; The bas includes a plurality of thinking model example text sub-libraries, different thinking model example text sub-libraries correspond to different first-level labels, and each thinking model example text sub-library includes a plurality of thinking model example texts corresponding to corresponding second-level labels; num is the number of corresponding thinking model example texts of cle ra .
[0028] In this embodiment, bas=(bas1, bas2, …, bas ce , …, bas he ), bas ce is the thinking model example text sub-library corresponding to the ce-th first-level label of cle, and bas ce =(bas ce,1 ,bas ce,2 ,…,bas ce,me ,…,bas ce,te ), bas ce,me is the thinking model example text corresponding to the me-th second-level label included in the ce-th first-level label of cle, the value range of me is 1 to te, te is the number of second-level labels included in the ce-th first-level label of cle, the value range of ce is 1 to he, and he is the number of first-level labels of cle.
[0029] In this embodiment, bas ce,meIncluding several thinking chain example texts, bas ce,me In each of the included thinking chain example texts, the corresponding label is the me-th secondary label included in the ce-th primary label of cle. However, bas ce,me The corresponding judgment identifiers in the different included thinking chain example texts may be the same or different, bas ce,me The corresponding judgment identifier in the included thinking chain example text is used to indicate whether the corresponding label in this thinking chain example text is the label of the corresponding target text in this thinking chain example text. For example, if bas ce,me The corresponding judgment identifier in the included thinking chain example text is correct or yes, it indicates that the corresponding label in this thinking chain example text is the label of the corresponding target text in this thinking chain example text; if bas ce,me The corresponding judgment identifier in the included thinking chain example text is wrong or no, it indicates that the corresponding label in this thinking chain example text is not the label of the corresponding target text in this thinking chain example text.
[0030] As a specific embodiment, R300 includes:
[0031] R310, if bas ce And l ab 1 ra Match, then enter R320; bas ce Is the ce-th thinking chain example text sub-library included in bas, and the value range of ce is from 1 to he, where he is the number of primary labels corresponding to cle.
[0032] In this embodiment, if bas ce The corresponding primary label and l [[ID=Z2]]ab 1 ra Are the same, then judge bas ce And l ab 1 ra Match; otherwise, judge bas ce And l ab 1 ra Do not match.
[0033] R320, if te = num, then obtain 1 thinking chain example text from each bas ce,me To construct mod ra ; if te > num, then enter R330; if te < num, then enter R350; te is the number of secondary labels included in the primary label corresponding to bas ce Corresponding, bas ce,me Is bas ceThe corresponding first-level tag includes the set of thought chain example texts corresponding to the me-th second-level tag, where the value of me ranges from 1 to te.
[0034] R330, obtain each bas ce,me Priority.
[0035] In this embodiment, each bas ce,me The priority is known, optional, bas ce,me The more example texts of mind chains included, the better. ce,me The higher the priority, the better.
[0036] R340, retrieves one thought chain example text from each target sub-library to build a mod. ra The target sub-library is bas ce The top num example text sub-libraries of the highest priority thought chain.
[0037] R350, from each base ce,me Obtain one example text of the thought chain to construct the initial text.
[0038] R360, obtain the text deviation quantity Δq, Δq=num-te.
[0039] R370, from bas ce Δq thought chain example texts are obtained from the remaining thought chain example texts and appended to the initial text. The appended initial text is then defined as mod. ra ;bas ce The remaining thought chain example text is bas ce Example text of the thought chain excluding the initial text.
[0040] Based on R310-R370, this embodiment is based on... l ab 1 ra Matching BAS ce The code retrieved num example texts of thought chains and implemented modulo operations. ra Automatic build eliminates the need for manual mod creation. ra The process; moreover, this embodiment incorporates as much information as possible from the context of... l ab 1 ra Matching BAS ce Extract example text of the thought chain from different secondary tags to enable mod ra This includes example texts of thought chains corresponding to various secondary tags, which will be helpful for later including mods. ra The prompt word pro included raThe large language model can more comprehensively learn the thinking of judging whether the secondary label is the secondary label of the text when inputting the trained large language model, so that the trained large language model can give a more accurate judgment result, and improve the accuracy of cleaning.
[0041] In the embodiment, if the value of num is small, that is, the number of thinking examples included in the prompt word input into the trained large language model is small, the trained large language model is not conducive to comprehensively learning the thinking of judging whether the label is correct, which affects the accuracy of the judgment result; if the value of num is large, that is, the number of thinking examples included in the prompt word input into the trained large language model is large, the trained large language model takes a long time to learn the thinking of judging whether the label is correct, which is not conducive to quickly giving a judgment result.
[0042] As an optional specific embodiment, num is an empirical value, and optionally, 2≤num≤10.
[0043] As a preferred specific embodiment, the obtaining process of num includes:
[0044] R301, obtaining a preset thinking example text library bas including a plurality of bas l ab 1 ra The matching bas ce include a plurality of thinking examples ra .
[0045] R302, obtaining the number cn of initial thinking examples ra , cn ra =f(min(num max ,max(num min ,num avg ×pn ra / pn avg ))),num max is the maximum number of preset thinking examples, num max =tok max / (min(tok avg )),tok max is the total character number of the preset thinking examples, and tok avg is a sequence composed of the average number of characters of the thinking examples included in each bas ce ; num min is the minimum number of preset thinking examples, num min =tok max / (max(tok avg ));num avgnum is the preset average number of thinking example texts avg = tok max avg(tok avg ) ; pn avg is the preset number threshold of thinking example texts; max() is the maximum value, min() is the minimum value, avg() is the average value, and f() is the integer.
[0046] In this embodiment, tok avg = (tok avg,1 , tok avg,2 , …, tok avg,ce , …, tok avg,he ).
[0047] In this embodiment, the average number of thinking example texts included in each thinking example text sub-library in bas is determined as the preset number threshold of thinking example texts.
[0048] R303, if cn ra is greater than or equal to l ab 1 ra the number of secondary labels, cn ra is determined as num; if cn ra is less than l ab 1 ra the number of secondary labels, ab l is determined as num. 1 ra the number of secondary labels included in ab
[0049] In this embodiment, if cn ra is less than l ab 1 ra the number of secondary labels, mod l ab 1 ra containing thinking example texts with a smaller number of words is selected or reconstructed from bas ce matching ab ra , so that mod ra includes a number of characters not exceeding tok max .
[0050] Based on R301-R303, this embodiment can achieve the goal of setting different thought chain example texts according to different texts to be cleaned. Specifically, considering the character limit corresponding to the prompt words, the maximum, minimum and average number of thought chain example texts are used as reference factors. It can also set a larger number of thought chain example texts when the thought chain example texts corresponding to the first-level tags of the text to be cleaned contain a large number of thought chain example texts, so as to improve the diversity of the set thought chain example texts. Considering the need for the trained large language model to learn the comprehensiveness of reasoning, the number of set thought chain example texts meets the requirement of being greater than or equal to the number of second-level tags included in the first-level tags of the text to be cleaned, so that the trained large language model can learn comprehensive tag judgment thinking, thereby improving the accuracy of judgment.
[0051] In this embodiment, mod ra =(mod ra,1 ,mod ra,2 ,…,mod ra,b ,…,mod ra,num ), mod ra,b For pro ra Includes the b-th thought chain example text, where b ranges from 1 to num; each mod ra,b Includes a corresponding target text tar ra,b A corresponding second-level tag l ab 2 ra,b and a corresponding judgment text tho ra,b ,tho ra,b Including a pair l ab 2 ra,b Is it tar? ra,b The label identification identifier sig ra,b and sig ra,b The corresponding reasoning text.
[0052] In this embodiment, with l ab 1 ra The secondary tags corresponding to the matched num example texts of the thought chain are also . l ab 1 ra The sub-tags, therefore, will be based on mod ra Build pro ra When used as cue words in a trained large language model, it helps the model learn to think in relation to the text being reasoned about, further improving the accuracy of the model's output. For example, l ab 1ra For the dispute, num=4, l ab 2 ra,1 For contract disputes, l ab 2 ra,2 For property damage compensation disputes, l ab 2 ra,3 Disputes over alimony, l ab 2 ra,4 For marital and family disputes. l ab 2 ra,1 , l ab 2 ra,2 , l ab 2 ra,3 and l ab 2 ra,4 All of these are sub-tags related to disputes.
[0053] As a specific implementation method, tar ra,1 Disputes arising between boyfriends and girlfriends due to relationship issues. l ab ra,1 entangled in pursuit of what was unsuccessful, sig ra,1 This is incorrect, sig ra,b The corresponding reasoning text is: the label belongs to the dispute caused by one party pursuing the other unsuccessfully, but the text mentions boyfriend and girlfriend, which conflicts with the unsuccessful pursuit. Therefore, the label is incorrect.
[0054] R400, according to mod ra cle ra and cle ra secondary tags l ab 2 ra Build cle ra The corresponding prompt word is "pro". ra .
[0055] In this embodiment, the cle to be determined as correct ra The tag is cle ra secondary tags l ab 2 ra .
[0056] R500, Pro ra Input the trained large language model, and determine whether to add cle based on the output of the trained large language model. ra Wash it off.
[0057] Optionally, if the trained large language model has good accuracy... l ab 2 ra Is it cle ra If the result of the tag judgment is yes, then it is determined that cle will not be used. ra Clean it up; if the trained large language model is correct l ab 2 ra Is it cle ra If the result of the tag judgment is negative, then the judgment will be changed to cle ra Wash it off.
[0058] In this embodiment, if the trained large language model is l ab 2 ra Is it cle ra The label judgment result is "yes", indicating that the trained large language model has judged it as "yes". l ab 2 ra For cle ra The tags; if the trained large language model is l ab 2 ra Is it cle ra If the label judgment result is negative, it means that the trained large language model has judged... l ab 2 ra Not cle ra The tag.
[0059] This embodiment applies to any text file in the text set to be cleaned. ra First, it is input into a trained first model, which has the function of obtaining the label of the input text. Therefore, the label can be obtained through the trained first model. ra The first-level tag l ab 1 ra In this embodiment, the trained first model has high accuracy in judging the first-level labels of the text, and the first-level labels output by the trained first model can be directly used as the classification criteria. ra The first-level tag l ab 1 ra After obtaining cle ra The first-level tag l ab 1 ra After that, l ab 1ra In the preset thinking example text library bas, the bas includes a thinking example text sub-library corresponding to different first-level labels. From the bas, the first-level label is also l ab 1 ra thinking example text; based on the thinking example text in the bas, the first-level label is also l ab 1 ra thinking example text to construct the prompt word, that is, based on the thinking example text associated with the cle ra has a greater relevance, which is conducive to the trained large language model to learn the thinking associated with the cle ra has a greater relevance, which is conducive to improving the semantic analysis ability of the trained large language model, and thus improving the accuracy of the output result.
[0060] Embodiment two:
[0061] In the above embodiment one, if the judgment result of the trained large language model to l ab 2 ra is the label of the cle ra is yes, it is judged that the cle ra is not cleaned up; if the judgment result of the trained large language model to l ab 2 ra is the label of the cle ra is no, it is judged that the cle ra is cleaned up; the above embodiment one does not consider how accurate the output result of the trained large language model is, if the accuracy of the output result of the trained large language model is low, then the accuracy of the cleaning of the cle is also low. In order to solve this problem, the R500 of the present embodiment comprises: if the judgment result of the trained large language model to l ab 2 ra is the label of the cle ra is yes, enter R510.
[0062] R510, obtaining the accuracy tpr of the judgment result of the trained large language model to
[0063] Optionally, the accuracy of the judgment result of the trained large language model to is based on the test text, when the accuracy tpr of the judgment result of the trained large language model to is 90% and the trained large language model to l ab 2 ra is the label of the cle raWhen the judgment result of the label is yes, it means l ab 2 ra is cle ra The probability of the label is 90%; when the true positive rate (tpr) of the trained large language model for the judgment result of yes is 78% and the trained large language model for l ab 2 ra whether it is cle ra When the judgment result of the label is yes, it means l ab 2 ra is cle ra The probability of the label is 78%.
[0064] R520, if tpr ≥ tf, then judge not to clean cle ra ; tf is a preset accuracy threshold.
[0065] In this embodiment, tf is an empirical value. Optionally, tf is 90% or 95%.
[0066] Based on R510 - R520, in this embodiment, when the accuracy of the trained large language model for the judgment result of yes is relatively high, directly take the judgment result of the trained large language model as the standard. Thus, the text with the judgment result of yes output by the trained large language model corresponding to cle can be retained, realizing the rapid cleaning of these texts.
[0067] In this embodiment, R5还 includes: if tpr < tf, then enter R521. <Based on R521 - R522, in this embodiment, when the accuracy rate of the trained large - language model for a judgment result of "yes" is relatively low, instead of directly taking the judgment result of the trained large - language model as the standard, the trained large - language model is used to make multiple judgments on the prompt words corresponding to the text, and the judgment result with the most occurrences in the multiple judgments is used as the final judgment result. Thus, this embodiment can reduce the judgment deviation triggered by factors such as hallucinations in the trained large - language model, and further improve the accuracy of cleaning.
[0071] As a preferred specific implementation manner, R500 further includes: If the judgment result of the trained large - language model on l ab 2 ra whether it is the label of cle ra is "no", then enter R530.
[0072] R530, obtain the false positive rate fpr of the trained large - language model for a judgment result of "no".
[0073] Optionally, the false positive rate of the trained large - language model for a judgment result of "no" is obtained based on test texts. When the false positive rate fpr of the trained large - language model for a judgment result of "no" is 68% and the trained large - language model's judgment result on l ab 2 ra whether it is the label of cle<000whether the label is cle ra , qua is a preset number of repeated reasoning, and qua is an odd number and qua is greater than or equal to 3.
[0077] In this embodiment, qua is an empirical value, and optionally, qua is 3 or 5.
[0078] R560, if there are more than or equal to (qua+1) / 2 times of the judgment results of the qua times of judgment results as no, it is judged that the cle ra is cleaned off; otherwise, it is judged that the cle ra is not cleaned off.
[0079] In this embodiment, if there are more than or equal to (qua+1) / 2 times of the judgment results of the qua times of judgment results as no, it means that the majority of the qua times of judgment results are no; if there are less than (qua+1) / 2 times of the judgment results of the qua times of judgment results as no, it means that the majority of the qua times of judgment results are yes.
[0080] Based on R530-R560, in the case that the accuracy of the trained large language model in judging the result as no is low, this embodiment does not directly take the judgment result of the trained large language model as the standard, but instead uses the trained large language model to make multiple judgments on the text corresponding to the prompt word, and takes the judgment result with the highest frequency in the multiple judgments as the final judgment result. Thus, this embodiment can also achieve recall of the text from the text corresponding to the preliminary judgment result as no, reduce the judgment deviation of the trained large language model due to the hallucination factor, and improve the accuracy of cleaning.
[0081] Embodiment three:
[0082] This embodiment provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the following steps when executing the computer program:
[0083] R100, obtaining a text set cle to be cleaned, the cle including a plurality of texts to be cleaned.
[0084] R200, inputting the ra-th text cle ra to be cleaned to a trained first model to obtain a first label of cle ra l ab 1 ra The trained first model is used to obtain a first-level label of the input text, and an accuracy of the trained first model in obtaining the first-level label of the input text is greater than or equal to tf, where tf is a preset accuracy threshold; and ra is in a range of 1 to to, where to is a quantity of texts to be cleaned in cle.
[0085] R300, according to l ab 1 ra In a preset thinking example text library bas, matching is performed to obtain a thinking example text mod corresponding to the first-level label ab of the first text cle l ab 1 ra The matched num thinking example texts mod ra The bas includes a plurality of thinking example text sub-libraries, different thinking example text sub-libraries correspond to different first-level labels, and each thinking example text sub-library includes a plurality of thinking example texts corresponding to corresponding second-level labels; num is the number of texts to be cleaned in cle ra corresponding to the second-level label.
[0086] R400, according to mod ra , cle ra and the second-level label of cle ra l ab 2 ra Constructing cle ra corresponding to the prompt word pro ra .
[0087] R500, inputting the pro ra into a trained large language model, and determining whether to clean the cle ra away according to an output of the trained large language model.
[0088] Embodiment Four
[0089] The embodiment provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0090] R100, obtaining a text set to be cleaned cle, and the cle includes a plurality of texts to be cleaned.
[0091] R200, inputting the ra text to be cleaned cle ra into a trained first model to obtain a first-level label ab of the cle ra l 1 ra The trained first model is used to obtain a first-level label of the input text, and an accuracy of the trained first model in obtaining the first-level label of the input text is greater than or equal to tf, where tf is a preset accuracy threshold; and ra is in a range of 1 to to, where to is a quantity of texts to be cleaned in cle.
[0092] R300, according to l ab 1 ra In a preset thinking example text library bas, matching is performed to obtain a thinking example text mod corresponding to the first-level label of the input text. l ab 1 ra The matched num thinking example texts mod ra The bas includes a plurality of thinking example text sub-libraries, different thinking example text sub-libraries correspond to different first-level labels, and each thinking example text sub-library includes a plurality of thinking example texts corresponding to corresponding second-level labels; num is the number of texts to be cleaned in cle ra .
[0093] R400, according to mod ra , cle ra and the second-level label of cle ra . l ab 2 ra The cle ra corresponding prompt pro ra is constructed.
[0094] R500, the pro ra is input into a trained large language model, and whether the cle ra is cleaned out is determined according to an output of the trained large language model.
[0095] Although some specific embodiments of the present application have been described in detail through examples, those skilled in the art should understand that the above examples are only for illustration, and are not intended to limit the scope of the present application. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present application. The scope of the present application is defined by the appended claims.
Claims
1. A text cleaning method based on a large language model, characterized in that, The method comprises the following steps: R100, obtaining a text set cle to be cleaned, the text set cle comprising a plurality of texts to be cleaned; R200, the first ra to be cleaned text cle ra Input into the trained first model, get cle ra The first level label l ab 1 ra ; the trained first model is used to obtain the first level label of the input text, the accuracy of the trained first model in obtaining the first level label of the input text is greater than or equal to tf, tf is a preset accuracy threshold; the value range of ra is 1 to to, to is the number of texts to be cleaned included in cle; R300, according to l ab 1 ra Match within the pre-defined thought chain example text library bas to obtain... l ab 1 ra Matching num example texts of thought chains mod ra ;bas includes several sub-libraries of thought chain example texts. Different sub-libraries correspond to different first-level tags, and each sub-library includes several thought chain example texts corresponding to second-level tags;num represents the number of thought chains. ra The number of corresponding thought chain example texts; mod ra =(mod ra,1 ,mod ra,2 ,…,mod ra,b ,…,mod ra,num ), mod ra,b For pro ra Includes the b-th thought chain example text, where b ranges from 1 to num; each mod ra,b Includes a corresponding target text tar ra,b A corresponding second-level tag l ab 2 ra,b and a corresponding judgment text tho ra,b ,tho ra,b Including a pair l ab 2 ra,b Is it tar? ra,b The label identification identifier sig ra,b and sig ra,b The corresponding reasoning text; R400, according to mod ra , cle ra and cle ra of the secondary label l ab 2 ra constructing cle ra corresponding prompt word pro ra ; R500, cle ra inputting the trained large language model, and determining whether to remove cle ra from the input according to the output of the trained large language model. l ab 2 ra whether the determination result of the label of cle ra is yes, it is determined not to remove cle ra from the input; if the determination result of the label of cle l ab 2 ra whether the determination result of the label of cle ra is no, it is determined to remove cle ra from the input.
2. The text cleaning method based on a large language model according to claim 1, characterized in that, R300 comprises: R310, if bas ce With l ab 1 ra Match, go to R320; bas ce The ce-th thinking model example text sub-library included in bas, ce is in the range of 1 to he, and he is the number of first-level labels corresponding to cle. R320, if te=num, then from each bas ce,me Get one example text of the thought chain to build a mod ra ;te is bas ce The number of second-level tags included in the corresponding first-level tag, bas ce,me for bas ce The corresponding first-level tag includes the set of thought chain example texts corresponding to the me-th second-level tag, where the value of me ranges from 1 to te.
3. The text cleaning method based on a large language model according to claim 2, characterized in that, R320 further comprises: if te>num, entering R330; R330, get the priority of each bas ce,me ; R340, from each target sub-library, obtain 1 thinking model example text construction mod ra , the target sub-library is bas ce The first num thinking model example text sub-libraries with the highest priority in the target sub-library.
4. The text cleaning method based on a large language model according to claim 2, characterized in that, R320 further comprises: if te<num, entering R350; R350, from each bas ce,me An example of a thought model is shown in FIG.
3. The initial text is constructed from the text of each bas R360, obtaining a text deviation quantity Δq, Δq=num-te; R370, from bas ce acquire Δq examples of thinking tokens from the remaining examples of thinking tokens and append them to the initial text, and determine the appended initial text as mod ra .
5. The text cleaning method based on a large language model according to claim 2, characterized in that, The obtaining process of num comprises: R301, the preset thinking model example text library bas in which l ab 1 ra Matching bas ce The number of thinking model examples pn included ra ; R302, the number of initial thought example texts cn is obtained ra , cn ra = f(min(num max , max(num min , num avg × pn ra / pn avg ))) num max is the preset maximum number of thought example texts, num max = tok max / (min(tok avg )) tok max is the preset total character number of thought example texts, tok avg is a sequence composed of the average number of characters of thought example texts included in each bas ce ; num min is the preset minimum number of thought example texts, num min = tok max / (max(tok avg )) num avg is the preset average number of thought example texts, num avg = tok max / (avg(tok avg )) pn avg is the preset number threshold of thought example texts; max() is the maximum value, min() is the minimum value, avg() is the average value, and f() is the integer value. R303, if cn ra greater than or equal to l ab 1 ra the number of secondary tags included, cn is determined as num; if cn ra is determined as num; if cn ra less than l ab 1 ra the number of secondary tags included, cn is determined as num; if cn l ab 1 ra the number of secondary tags included, cn is determined as num.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the text cleaning method based on the large language model in any one of claims 1-5.
7. A computer-readable storage medium storing a computer program, wherein the computer program comprises the following steps of: The computer program is executed by the processor to realize the text cleaning method based on the large language model in any one of claims 1-5.
Citation Information
Patent Citations
Emergency information query method and device based on large model thinking chain
CN117076649A
Text data label optimization method and device, equipment and storage medium
CN117891945A
Method for processing large-scale domain question annotation through large language model
CN118210919A