A text cleaning method based on text tags

By obtaining and using prompt words from the large language model, combining multiple judgments and accuracy thresholds, we automatically judge and clean the error tags in the text set, solving the problem of manual judgment occupies a lot of resources and improving the efficiency and accuracy of text cleaning.

CN118690021BActive Publication Date: 2025-07-08BEIJING RUIQI INFORMATION TECH CO LTD +3
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410997259.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-07-08
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

In the prior art, since the large language model has the ability to classify text and give classification reasons, how to achieve automatic judgment of whether the corresponding labels of text are correct based on the large language model to complete the cleaning task of text sets has not been solved, resulting in a large amount of manual resources.

Method used

By obtaining the prompt words of the text to be cleaned, including the thinking chain sample text and the text to be reasoned, the correctness of the text tag is judged using the trained large language model, and combining the multiple judgment results and accuracy thresholds, the wrong tags are automatically judged and cleaned.

Benefits of technology

Automatic judgment of text labels is realized, the occupation of human resources is reduced, and the efficiency and accuracy of cleaning tasks are improved, especially when processing large-scale text sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118690021B_ABST
    Figure CN118690021B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electronic digital data processing, and in particular to a text cleaning method based on text labels. The method comprises: obtaining a text set cle to be cleaned; obtaining the rath text cle to be cleaned in the cle ra Corresponding prompt word pro ra , pro ra Includes thought chain example text mod ra and the text to be inferred ans ra ; will pro ra Input the trained large language model to be judged by the trained large language model l ab ra Is it cle ra The label of l ab ra Is it cle ra The result of the label judgment determines whether to cle ra The present invention realizes automatic judgment on whether the label corresponding to the text is correct, thereby completing the task of cleaning the text set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular, to a text cleaning method based on text tags. Background Art

[0002] A text set is a set composed of several texts, and each text corresponds to a labeled tag. However, the labeled tags may be incorrect, that is, there may be a situation where the tag labeled for a text does not match the text. The cleaning task of the text set includes cleaning the texts in the text set whose corresponding tags are incorrect tags and retaining the texts in the text set whose corresponding tags are correct tags; when the text set includes a large number of texts, if it is manually determined whether the tags corresponding to the texts included in the text set are correct, there is a problem of consuming a large amount of human resources.

[0003] A large language model refers to a language model trained on a large amount of text corpus and containing tens of billions of parameters or more. The large language model can better understand and generate human language. The large language model can be used to classify and reason about texts and give corresponding reasoning reasons. For example, Chinese Patent Application No. CN117951572A discloses a text classification method, device, medium, and electronic device based on a large language model. This patent application realizes information extraction of target texts, classification of target texts, and giving corresponding classification reasons through a large language model.

[0004] The task of judging whether the tag of a text is correct is essentially also a text classification task. In the existing technology, when the large language model already has the ability to classify texts and give classification reasons, how to automatically judge whether the tag corresponding to a text is correct based on the large language model, and then complete the cleaning task of the text set is an urgent problem to be solved. Summary of the Invention

[0005] The purpose of the present invention is to provide a text cleaning method based on text tags to automatically judge whether the tag corresponding to a text is correct based on a large language model, and then complete the cleaning task of the text set.

[0006] According to the present invention, a text cleaning method based on text tags includes the following steps:

[0007] L100, obtaining a text set cle to be cleaned, where cle includes several texts to be cleaned.

[0008] L200, obtaining the ra-th text cle to be cleaned ra corresponding prompt word pro ra , pro ra including a chain-of-thought example text set mod ra and text to be inferred ansra , mod ra includes num chain-of-thought example texts, where num is the ra number of corresponding chain-of-thought example texts; mod ra the b-th chain-of-thought example text included in mod ra,b includes a corresponding target text tar ra,b , a corresponding label l ab ra,b and a corresponding judgment text tho ra,b , tho ra,b includes a pair of l ab ra,b whether it is the label of tar ra,b judgment identifier sig ra,b and the reasoning text corresponding to sig ra,b ; ans ra includes cle ra and cle ra 's label l ab ra ; The value range of ra is from 1 to to, where to is the number of texts to be cleaned included in cle; the value range of b is from 1 to unm.

[0009] L300, input pro ra into the trained large language model to be judged by the trained large language model l ab ra whether it is the label of cle ra .

[0010] L400, according to the judgment result of the trained large language model on l ab ra whether it is the label of cle ra , judge whether to clean cle ra .

[0011] The present invention has at least the following beneficial effects:

[0012] The present invention intends to use the large language model to judge whether the label of the text to be cleaned is incorrect. Specifically, for each text to be cleaned, first obtain its corresponding prompt word, which includes num chain-of-thought example texts and the text to be reasoned. The num chain-of-thought example texts are used to enable the trained large language model to better learn the thinking of judging whether the label is the label of the text. Each chain-of-thought text includes a target text, a label, a judgment identifier indicating whether the label is the target text, and the reason for making this judgment. Based on the num chain-of-thought example texts, the trained large language model can learn the thinking of judging whether the label is the corresponding label of the text and has the ability to judgel ab ra Whether it is cle ra The semantic analysis ability of the label; thus, by inputting the prompt word corresponding to the text to be judged into the trained large language model, the automatic judgment of whether the label of the text is correct can be realized, solving the problem of consuming a large amount of human resources in the prior art for manually judging whether the label of the text is correct. Brief Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0014] Figure 1 It is a flowchart of the text cleaning method based on text labels provided in Embodiment 1 of the present invention. Detailed Embodiments

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0016] Embodiment 1:

[0017] According to the present invention, a text cleaning method based on text labels is provided, as Figure 1 shown, including:

[0018] L100, obtaining a text set cle to be cleaned, where cle includes several texts to be cleaned.

[0019] In this embodiment, cle = {cle1, cle2,..., cle ra ,..., cle to}, and cle ra is the ra-th text to be cleaned, and the value range of ra is from 1 to to, where to is the number of texts to be cleaned.

[0020] L200, obtaining the prompt word pro ra corresponding to the ra-th text cle to be cleaned in cle ra , where pro ra includes a thinking chain example text set mod ra and the text to be inferred ans ra , and modra including num chain-of-thought example texts, where num is cle ra the quantity of corresponding chain-of-thought example texts; mod ra the b-th chain-of-thought example text included mod ra,b including a corresponding target text tar ra,b and a corresponding label l ab ra,b and a corresponding judgment text tho ra,b tho ra,b including a pair of l ab ra,b judgment identifier sig of whether it is the label of tar ra,b and the reasoning text corresponding to sig ra,b and ans ra,b corresponding to sig ra including cle ra and cle ra labels l ab ra ; the value range of ra is from 1 to to, where to is the quantity of texts to be cleaned included in cle; the value range of b is from 1 to unm.

[0021] In this embodiment, mod ra ={mod ra,1 , mod ra,2 , …, mod ra,b , …, mod ra,num}.

[0022] In this embodiment, if the value of num is small, that is, the quantity of chain-of-thought example texts included in the prompt input to the trained large language model is small, it is not conducive to the trained large language model to comprehensively learn the thinking of judging whether the label is correct, affecting the accuracy of the judgment result; if the value of num is large, that is, the quantity of chain-of-thought example texts included in the prompt input to the trained large language model is large, the time for the trained large language model to learn the thinking of judging whether the label is correct is long, which is not conducive to quickly giving the judgment result; as an optional specific implementation manner, num is an empirical value, and optionally, 2 ≤ num ≤ 10.

[0023] As a preferred specific implementation manner, each l ab ra,b and l ab ra are sub-labels under the same superior label. Thus, pro raWhen used as a prompt word for a trained large language model, it is helpful for the trained large language model to learn the thinking that is more relevant to the text to be inferred, and further improve the accuracy of the output results of the trained large language model. For example, num=4, l ab ra For marriage and family disputes, l ab ra,1 For contract disputes, l ab ra,2 Disputes over compensation for property damage, l ab ra,3 For maintenance disputes, l ab ra,4 For marriage and family disputes, l ab ra , l ab ra,1 , l ab ra,2 , l ab ra,3 and l ab ra,4 They are all secondary tags. l ab ra , l ab ra,1 , l ab ra,2 , l ab ra,3 and l ab ra,4 The parent tags (i.e. first-level tags) of are all disputes. l ab ra , l ab ra,1 , l ab ra,2 , l ab ra,3 and l ab ra,4 All are sub-tags of the dispute tag.

[0024] As a specific implementation, tar ra,1 If there is a dispute between the two parties regarding the delivery date in the signed contract, l ab ra,1 For contract disputes, sig ra,1 To be correct, sig ra,b The corresponding reasoning text is: The two parties have signed a contract, and they have a dispute over the delivery date in the contract, which is a contractual dispute. Therefore, the label is correct.

[0025] L300, the pro ra Input the trained large language model to be judged by the trained large language model lab ra Whether it is cle ra 's label.

[0026] Those skilled in the art know that any large language model in the prior art falls within the protection scope of the present invention.

[0027] In this embodiment, pro ra is a prompt word based on the chain of thought. In this embodiment, pro ra is used as the input of the trained large language model, which is beneficial for the trained large language model to learn the thought of judging whether the labeled label is the label of the text, improve the accuracy of the output result of the trained large language model, and further improve the accuracy of cleaning.

[0028] L400, according to the trained large language model for l ab ra Whether it is cle ra 's label, judge whether to clean cle ra off.

[0029] This embodiment intends to use a large language model to judge whether the label of the text to be cleaned is incorrect. Specifically, for each text to be cleaned, first obtain its corresponding prompt word, which includes num chain-of-thought example texts and the text to be inferred. The num chain-of-thought example texts are used to enable the trained large language model to better learn the thought of judging whether the label is the label of the text. Each chain-of-thought text includes a target text, a label, a judgment identifier indicating whether the label is the target text, and the reason for making this judgment. Based on the num chain-of-thought example texts, the trained large language model can learn the thought of judging whether the label is the corresponding label of the text and has the semantic analysis ability to judge l ab ra Whether it is cle ra 's label; thus, by inputting the prompt word corresponding to the text to be judged into the trained large language model, the automatic judgment of whether the label of the text is correct can be realized, solving the problem of consuming a large amount of human resources in the prior art for manually judging whether the label of the text is correct.

[0030] As a preferred specific implementation manner, L400 includes: if the judgment result of the trained large language model for l ab ra Whether it is cle ra 's label is yes, then enter L410.

[0031] L410, obtain the accuracy rate tpr of the trained large language model for the judgment result of yes.

[0032] In this embodiment, if the trained large language model l ab ra is determined to be the label of cle ra as yes, it means that the trained large language model determines that l ab ra is the label of cle ra .

[0033] Optionally, the accuracy rate of the trained large language model for a yes judgment result is obtained based on the test text. When the accuracy rate tpr of the trained large language model for a yes judgment result is 90% and the trained large language model l ab ra is determined to be the label of cle ra as yes, it means that l ab ra is the label of cle ra with a probability of 90%; when the accuracy rate tpr of the trained large language model for a yes judgment result is 78% and the trained large language model l ab ra is determined to be the label of cle ra as yes, it means that l ab ra is the label of cle ra with a probability of 78%.

[0034] L420, if tpr ≥ tf, then it is determined not to wash out cle ra ; tf is a preset accuracy rate threshold.

[0035] In this embodiment, tf is an empirical value. Optionally, tf is 90% or 95%.

[0036] Based on L410 - L420, in this embodiment, when the accuracy rate of the trained large language model for a yes judgment result is relatively high, the judgment result of the trained large language model is directly used as the standard. Thus, the text with a yes judgment result output by the trained large language model corresponding to cle can be retained, achieving fast cleaning of these texts.

[0037] In this embodiment, L420 further includes: if tpr < tf, then go to L421.

[0038] L421, obtain the qua - th judgment result of the trained large language model on l ab ra whether it is the label of cle ra . qua is a preset number of repeated inferences, qua is an odd number and qua ≥ 3.

[0039] L422, if there are more than or equal to (qua + 1) / 2 negative judgment results among the qua - level judgment results, then it is judged that cle will be ra washed away; otherwise, it is judged that cle will not be ra washed away.

[0040] Based on L421 - L422, in this embodiment, when the accuracy rate of the judgment result being yes by the trained large - language model is relatively low, instead of directly taking the judgment result of the trained large - language model as the standard, the trained large - language model is used to make multiple judgments on the prompt words corresponding to the text, and the judgment result with the most occurrences among the multiple judgments is used as the final judgment result. Thus, this embodiment can reduce the judgment deviation triggered by factors such as hallucinations in the trained large - language model, and further improve the accuracy of cleaning.

[0041] As a preferred specific implementation manner, L400 further includes: if the trained large - language model l ab ra judgment result on whether it is the label of cle ra is negative, then enter L430.

[0042] L430, obtain the false positive rate fpr of the trained large - language model for the negative judgment result.

[0043] In this embodiment, if the trained large - language model l ab ra judgment result on whether it is the label of cle ra is negative, it means that the trained large - language model judges l ab ra is not the label of cle ra .

[0044] Optionally, the false positive rate of the trained large - language model for the negative judgment result is obtained based on the test text. When the false positive rate fpr of the trained large - language model for the negative judgment result is 68% and the trained large - language model l ab ra judgment result on whether it is the label of cle ra is negative, it means that l ab ra is not the label of cle ra with a probability of 68%; when the false positive rate fpr of the trained large - language model for the negative judgment result is 92% and the trained large - language model l ab ra judgment result on whether it is the label of cle ra is negative, it means that l ab ra is not the label of cle raThe probability of the label is 92%.

[0045] L440, if fpr < tf, then enter L450; otherwise, cle ra is washed away.

[0046] In this embodiment, tf is an empirical value. Optionally, tf is 90% or 95%.

[0047] L450, obtain the judgment result of whether the text corresponding to the prompt word of the trained large language model is cle l ab ra is cle ra The qua - th judgment result, where qua is a preset number of repeated inferences, qua is an odd number and qua ≥ 3.

[0048] In this embodiment, qua is an empirical value. Optionally, qua is 3 or 5.

[0049] L460, if there are ≥ (qua + 1) / 2 judgment results of no among the qua - th judgment results, then judge that cle ra is washed away; otherwise, judge that cle ra is not washed away.

[0050] In this embodiment, if there are ≥ (qua + 1) / 2 judgment results of no among the qua - th judgment results, it means that most of the qua - th judgment results are no; if there are < (qua + 1) / 2 judgment results of no among the qua - th judgment results, it means that most of the qua - th judgment results are yes.

[0051] Based on L430 - L460, in this embodiment, when the accuracy of the judgment result of no by the trained large language model is relatively low, instead of directly taking the judgment result of the trained large language model as the standard, the trained large language model is used to make multiple judgments on the prompt word corresponding to the text, and the judgment result with the most occurrences among the multiple judgments is used as the final judgment result. Thus, this embodiment can also recall the text from the text with a preliminary judgment result of no, reduce the judgment deviation triggered by factors such as hallucinations of the trained large language model, and improve the accuracy of cleaning.

[0052] Embodiment Two:

[0053] If the prompt word used for input to the trained large language model in Embodiment One is constructed manually, there is still the problem of relatively high occupation of human resources. To further reduce the problem of occupation of human resources, on the basis of Embodiment One, this embodiment also includes the process of automatically obtaining pro ra of ra The process of obtaining pro includes:

[0054] L210, will cle ra Input to the first trained model to obtain cle ra First-level label l ab 1 ra ; The trained first model is used to obtain the primary label of the input text, and the accuracy of the trained first model in obtaining the primary label of the input text is greater than or equal to tf, where tf is a preset accuracy threshold.

[0055] In this embodiment, the accuracy of the first-level label of the text inferred by the trained first model is relatively high. ra The output of the first model trained with ra First-level label l ab 1 ra .

[0056] In this embodiment, the trained first model is used to obtain the first-level label of the input text. The first-level label is a rough label. The first-level label also includes several second-level labels. The second-level label is a refined label. In this embodiment, the accuracy of the second-level label of the text inferred by the trained first model is low (i.e., less than tf). Therefore, this embodiment only uses the trained first model to obtain the first-level label of the text, and does not use the trained first model to obtain the second-level label of the text.

[0057] In this embodiment, the first model is a neural network model, and a supervised training method is used when training the first model. Optionally, the process of training the first model includes: obtaining a training text set, the training text set includes a plurality of training texts; obtaining a training label set, the training label set includes a primary label corresponding to each training text; training the first model according to the training text set and the training label set to obtain a trained first model.

[0058] L220, according to l ab 1 ra Match in the preset thinking chain sample text library bas to obtain l ab 1 ra Matching num thought chain example text mod ra ; bas=(bas1,bas2,…,bas ce ,…,bas he ),bas ce The sub-library of example texts of the thinking chain corresponding to the ceth first-level label corresponding to cle, bas ce =(basce,1 , bas ce,2 , …, bas ce,me , …, bas ce,te ), bas ce,me is the example text of the thinking chain corresponding to the me-th secondary label included in the ce-th primary label corresponding to cle. The value range of me is from 1 to te, where te is the number of secondary labels included in the ce-th primary label corresponding to cle, and the value range of ce is from 1 to he, where he is the number of primary labels corresponding to cle; num is the number of example texts of the thinking chain corresponding to cle ra corresponding to the thinking chain example text.

[0059] In this embodiment, bas ce,me includes several example texts of the thinking chain, bas ce,me In each example text of the thinking chain included, the corresponding label is the me-th secondary label included in the ce-th primary label corresponding to cle. However, bas ce,me the corresponding judgment identifiers in different example texts of the thinking chain included may be the same or different, bas ce,me the judgment identifier corresponding to the example text of the thinking chain included is used to indicate whether the corresponding label in this example text of the thinking chain is the label of the target text corresponding to this example text of the thinking chain. For example, if bas ce,me the judgment identifier corresponding to the example text of the thinking chain included is correct or yes, it indicates that the corresponding label in this example text of the thinking chain is the label of the target text corresponding to this example text of the thinking chain; if bas ce,me the judgment identifier corresponding to the example text of the thinking chain included is wrong or no, it indicates that the corresponding label in this example text of the thinking chain is not the label of the target text corresponding to this example text of the thinking chain.

[0060] Optionally, L220 includes:

[0061] L221, if bas ce matches with l ab 1 ra then enter L222.

[0062] In this embodiment, if bas ce the corresponding primary label matches with l ab 1 ra is the same, then it is judged that bas ce matches with l ab 1 ra ; otherwise, it is judged that bas ce matches with l ab 1 raMismatch.

[0063] L222, if te = num, obtain 1 example text of the thinking chain from each bas ce,me to construct mod ra ; if te > num, go to L223; if te < num, go to L225.

[0064] L223, obtain the priority of each bas ce,me

[0065] In this embodiment, the priority of each bas ce,me is known. Optionally, the more the number of example texts of the thinking chain included in bas ce,me , the higher the priority of bas ce,me .

[0066] L224, obtain 1 example text of the thinking chain from each target sub-library to construct mod ra , where the target sub-library is the top num example text sub-libraries with the highest priority in bas ce .

[0067] L225, obtain 1 example text of the thinking chain from each bas ce,me to construct the initial text.

[0068] L226, obtain the text deviation quantity Δq, Δq = num - te;

[0069] L227, obtain Δq example texts of the thinking chain from the remaining example texts of the thinking chain in bas ce and append them to the initial text, and determine the appended initial text as mod ra ; the remaining example texts of the thinking chain in bas ce are the example texts of the thinking chain in bas ce except the initial text.

[0070] Based on L221 - L227, this embodiment obtains num example texts of the thinking chain from bas l ab 1 ra matching ce , realizes the automatic construction of mod ra , and saves the process of manually constructing mod ra ; moreover, this embodiment obtains example texts of the thinking chain with different secondary labels from bas l ab 1 ra matching ce as much as possible to make mod ra ​Including thinking chain example texts corresponding to multiple secondary tags, which is beneficial for later inputting prompt words pro ra including mod ra into the trained large language model. The large language model can more comprehensively learn the thinking of judging whether the secondary tag is a secondary tag of the text, so that the trained large language model can give a more accurate judgment result and improve the accuracy of cleaning.

[0071] L230, according to mod ra 、cle ra and cle ra of the secondary tag l ab 2 ra Construct the prompt word pro corresponding to cle ra 。 ra 。

[0072] Based on L210-L230, this embodiment can automatically construct pro ra , eliminating the process of manually constructing prompt words, improving the efficiency of constructing prompt words, and also reducing the human resources occupied in the text cleaning process.

[0073] As a preferred specific implementation manner, the acquisition process of num includes:

[0074] L201, obtain the number pn of thinking chain example texts included in the preset thinking chain example text library bas that matches l ab 1 ra 。 ce 。 ra 。

[0075] L202, obtain the number cn of initial thinking chain example texts ra , cn ra =f(min(num max ,max(num min ,num avg ×pn ra / pn avg ))), num max is the maximum number of preset thinking chain example texts, num max =tok max / (min(tok avg,1 ,tok avg,2 ,…,tok avg,ce ,…,tok avg,he ))), tok max is the total character number of the preset thinking chain example texts, tok avg,ce is bas ceAverage number of characters of the included chain-of-thought example text; num min Is the minimum number of the preset chain-of-thought example text, num min =tok max / (max(tok avg,1 ,tok avg,2 ,…,tok avg,ce ,…,tok avg,he )); num avg Is the average number of the preset chain-of-thought example text, num avg =tok max / (avg(tok avg,1 ,tok avg,2 ,…,tok avg,ce ,…,tok avg,he )); pn avg Is the quantity threshold of the preset chain-of-thought example text; max( ) is to take the maximum value, min( ) is to take the minimum value, avg( ) is to take the average value, and f( ) is to take the integer part.

[0076] In this embodiment, the average number of the chain-of-thought example texts included in each chain-of-thought example text sub-library in bas is determined as the quantity threshold of the preset chain-of-thought example text.

[0077] L203, if cn ra is greater than or equal to l ab 1 ra the number of secondary labels included, then cn ra is determined as num; if cn ra is less than l ab 1 ra the number of secondary labels included, then l ab 1 ra the number of secondary labels included is determined as num.

[0078] In this embodiment, if cn ra is less than l ab 1 ra the number of secondary labels included, then preferentially select or reconstruct mod l ab 1 ra from the bas ce matching to include the chain-of-thought example text with fewer words to form mod ra , so that the number of characters included in mod ra does not exceed tok max that's all.

[0079] Based on L201-L203, this embodiment can achieve the purpose of setting different chain-of-thought example texts according to different texts to be cleaned. Specifically, considering the character number limit corresponding to the prompt words, taking the maximum number, minimum number, and average number of the chain-of-thought example texts as reference factors, it is also possible to set a larger number of chain-of-thought example texts when the number of chain-of-thought example texts included in the chain-of-thought example texts corresponding to the first-level label of the text to be cleaned is relatively large, so as to improve the diversification of the set chain-of-thought example texts; considering the need for the trained large language model to learn comprehensive reasoning, the number of the set chain-of-thought example texts meets the requirement of being greater than or equal to the number of second-level labels included in the first-level label of the text to be cleaned, so that the trained large language model can learn the comprehensive label judgment thinking and have stronger semantic analysis ability, thereby improving the accuracy of judgment.

[0080] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A text cleaning method based on text tags, characterized in that, including the following steps: L100, obtaining a text set cle to be cleaned, where cle includes several texts to be cleaned; L200, obtain the ra-th text cle to be cleaned in cle ra The corresponding prompt pro ra , pro ra includes the thought chain example text set mod ra and the text ans to be inferred ra , mod ra includes num thought chain example texts, where num is the number of thought chain example texts corresponding to cle ra ; mod ra includes the b-th thought chain example text mod ra,b includes a corresponding target text tar ra,b 、a corresponding label l ab ra,b and a corresponding judgment text tho ra,b , tho ra,b includes a pair l ab ra,b whether it is tar ra,b the judgment flag sig of the label ra,b and the inference text corresponding to sig ra,b ; ans ra includes cle ra and cle ra 's label l ab ra ; The value range of ra is from 1 to to, where to is the number of texts to be cleaned included in cle; The value range of b is from 1 to num; L300, take the pro ra input to the trained large language model for the trained large language model to judge l ab ra whether it is the label of cle ra ; L400, based on the trained large language model for l ab ra whether it is cle ra Based on the judgment result of the label, determine whether to ra clean cle The obtaining process of num includes: L201, obtain the number pn of the thought chain example texts corresponding to l ab 1 ra in the preset thought chain example text library ra , l ab 1 ra is the first-level label of cle ra ; L202, obtain the number cn of initial thought chain example texts ra , cn ra = f(min(num max , max(num min , num avg × pn ra / pn avg ))), num max is the maximum number of preset thought chain example texts, num min is the minimum number of preset thought chain example texts, num avg is the average number of preset thought chain example texts, pn avg is the quantity threshold of the preset thought chain example texts; max( ) is to take the maximum value, min( ) is to take the minimum value, and f( ) is to round; L203, if cn ra is greater than or equal to l ab 1 ra the number of secondary tags included, then set cn ra as num; if cn ra is less than l ab 1 ra the number of secondary tags included, then l ab 1 ra the number of secondary tags included is set as num.

2. The text cleaning method based on text tags according to claim 1, wherein L400 includes: If the judgment result of whether the trained large language model is l ab ra the label of cle ra is yes, then proceed to L410; L410, obtaining the accuracy rate tpr of the trained large language model for the judgment result of yes; L420, if tpr ≥ tf, then it is determined not to wash off cle ra ; tf is a preset accuracy threshold.

3. The text cleaning method based on text tags according to claim 2, wherein L400 further includes: If the judgment result of whether the trained large language model is l ab ra the label of cle ra is no, then it enters L430; L430, obtaining the accuracy rate fpr of the trained large language model for the judgment result of no; L440, if fpr < tf, then enter L450; L450, obtain the judgment result of the large language model trained on whether l ab ra is cle ra for the qua-th time, where qua is the preset number of repeated inferences, qua is an odd number and qua ≥ 3; L460, if there are more than or equal to (qua + 1) / 2 negative judgment results among the qua judgment results, then it is judged that cle will be ra washed away; otherwise, it is judged that cle will not be ra washed away.

4. The text cleaning method based on text tags according to claim 2, characterized in that L420 further includes: if tpr < tf, then enter L421; L421, obtain the judgment result of the trained large language model on l ab ra whether it is cle ra for the qua-th time, where qua is the preset number of repeated inferences, qua is an odd number and qua ≥ 3; L422, if there are judgments with results of no in more than or equal to (qua + 1) / 2 times among the qua - level judgment results, then it is judged that cle will be ra washed away; otherwise, it is judged that cle will not be ra washed away.

5. The text cleaning method based on text tags according to claim 3, wherein L440 further includes: if fpr ≥ tf, then cle ra is washed off.

6. The text cleaning method based on text tags according to claim 1, wherein Each l ab ra,b and l ab ra are sub - tags under the same parent tag.

Citation Information

Patent Citations

  • Text classification method and device based on large language model, medium and electronic equipment

    CN117951572A

  • Text classification method and device, computer equipment and storage medium

    CN113011533A

  • Text data label optimization method and device, equipment and storage medium

    CN117891945A