Domain corpus data auditing and automatic correcting method based on generative model
By adopting the audit and automatic correction methods based on generative models in the corpus data review, the problems of artificial dependence and insufficient model adaptability in the existing technology are solved, and efficient and accurate corpus review and automatic correction processes are achieved.
Patent Information
- Application Number
- CN202510137831.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art relies on manual labor in corpus data review, is inefficient and susceptible to subjective influence, lacks self-review and update mechanisms, limits the adaptability and flexibility of the model, and lacks automatic correction measures, resulting in an increase in manual intervention.
The domain corpus data review and automatic correction methods based on generative models are adopted, and through preprocessing, LLM audit, manual review, automatic correction and model update steps, manual dependence is reduced, model self-review and update capabilities are enhanced, and the corpus content is automatically corrected.
It improves the accuracy and efficiency of multi-field corpus review, enhances the adaptability and flexibility of the model, reduces manual intervention, and achieves a more efficient and accurate intelligent corpus review and automatic correction process.
Smart Images

Figure CN120012764A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, specifically to the subfield of generative artificial intelligence technology and natural language processing technology, and provides a domain corpus data review and automatic correction method based on a generative model. Background Art
[0002] At present, generative artificial intelligence technology has gradually become popular in various fields such as education, medical care, finance, and law. The demand for high-quality domain-specific corpus data for large language models (LLM) has become increasingly urgent. However, the current generation and review of corpus data is overly dependent on experts, which is inefficient and susceptible to subjective influences. However, professional fields have very high requirements for terminology and data accuracy, and large data scale requirements, which are difficult for experts to fully meet, and there may be omissions in the review results. The generation of domain data will limit the further development of the industry. At present, there is an urgent need for more efficient domain data review methods to improve the quality of corpus and meet the needs of large models in different industry fields.
[0003] Disadvantages of the prior art: With the development and promotion of information technology and artificial intelligence technology, there are currently multiple methods for corpus organization and information review based on LLM, as follows: Patent CN118797536A "A Comprehensive Audit Processing Method for Network Information Data Based on a Big Model" describes a comprehensive audit processing method for network information data based on a big model, which aims to improve the efficiency and accuracy of content auditing. This method combines big model technology and historical data to improve the accuracy and adaptability of the model, and is suitable for processing large-scale network information data. However, it does not specify the format and type of the audit content, and does not propose a method for correcting the audit content.
[0004] Patent CN118551045A "A method and device for reviewing test reports based on a large model" This invention reduces manual workload and subjective misjudgment through automated review processes, improves review efficiency and accuracy; uses large models and natural language processing technology to process large-scale professional data and optimize anomaly detection; ensures the accuracy of project information by comparing with databases; and provides data-driven insights to help manage and understand infrastructure conditions. However, the patent lacks a self-review and update mechanism and relies entirely on pre-trained language models, which limits its transferability to other fields. In addition, this technology is mainly aimed at the review of test reports, and does not involve more complex corpus generation review and automatic correction.
[0005] Patent CN118551046A "A method for enhancing document processing flow based on a large language model" This patent proposes a method for enhancing document processing flow based on a large language model and Transformer architecture, aiming to improve the efficiency and accuracy of processing a large number of documents by building an intelligent and automated document processing system. The method covers multiple stages such as document capture, classification, extraction, review, enrichment and data integration, and integrates a large language model to achieve document content understanding, automatic extraction of key data, context enrichment and data synthesis. However, this patent requires a manual real-time supervision module, which has high labor cost requirements and fails to fully reduce the manual burden.
[0006] Although the current corpus review technology based on large language models (LLMs) has made progress in improving review efficiency and accuracy, it still relies on manual review, especially in multi-domain corpus review, which requires manual feedback to improve accuracy. In addition, the existing technology lacks a self-review and update mechanism and over-relies on pre-trained models, which limits the adaptability and flexibility of the model. At the same time, the lack of automatic correction measures for the review content means that even if problems are identified, manual intervention is required to correct them, which increases the workload and may affect the review efficiency. Therefore, future research needs to focus on reducing reliance on manual review, enhancing the model's self-learning and update capabilities, and developing automated correction mechanisms to achieve more efficient and accurate corpus review. Summary of the invention
[0007] The purpose of this invention is to gradually reduce the reliance on manual labor in domain data review, enhance the self-review and update capabilities of large language models, and realize automatic correction of audit content. It will improve the accuracy of multi-domain corpus review, enhance the adaptability and flexibility of the model, and gradually reduce manual intervention through automated correction measures, thereby achieving a more efficient and accurate corpus intelligent review and automatic correction process, and provide a method for rapid domain data review and automatic correction for large models in various fields.
[0008] In order to achieve the above purpose, the present invention adopts the following technical solutions: The present invention provides a method for reviewing and automatically correcting domain corpus data based on a generative model, comprising the following steps: Step 1: Preprocess the original corpus to obtain standardized corpus; Step 1.1: Standardize the format of the original corpus from the Internet, manually collected, or generated by a large language model and convert it into a JSON format object; Step 1.2: Classify and count the JSON format objects obtained in step 1.1, and classify them according to the corpus type, including single-round dialogue, multi-round dialogue and statement; Step 1.3, using the LLM-based fuzzy recognition technology, correct the word order and punctuation of the classified corpus obtained in step 1.2; Step 1.4, perform syntactic analysis, part-of-speech tagging, semantic analysis and sentence reorganization on the corrected corpus obtained in step 1.3 to obtain the reorganized text; Step 1.5: Map the reorganized text into JSON format again and output it as standardized corpus.
[0009] Step 2: Conduct LLM review on the standardized corpus and obtain the scoring results; Step 2.1, multiple LLMs generate opinions on the current corpus according to the prompt project, each LLM generates K opinions, and a total of K×M opinions are obtained; Step 2.1.1, select multiple LLMs that have been fine-tuned by instructions to ensure that they have good performance on the domain corpus; Step 2.1.2: Segment the standardized corpus into paragraphs or sentences suitable for LLM processing; Step 2.1.3: Design a specific prompting project for each LLM to guide the model to generate opinions about the correctness of the corpus facts; Step 2.1.4: Input the segmented corpus into each LLM to generate multiple different viewpoints ; Step 2.1.5: Repeat step 2.1.4 to ensure that each LLM generates K different views to increase the diversity of the results; Step 2.1.6: Collect all the opinions generated by LLM, totaling K × M; Step 2.2: Perform cluster analysis on the opinions obtained in step 2.1 and calculate the generation probability of each semantically equivalent class; Step 2.2.1. Use the pre-trained BERT model to transform each opinion Encoded as a semantic vector of fixed dimension; Step 2.2.2, calculate the cosine similarity between all semantic vectors as the similarity measure for cluster analysis; Step 2.2.3, apply K-means clustering algorithm to divide the semantic vectors into multiple clusters according to cosine similarity, i.e., semantic equivalence classes; Step 2.2.4: Count the number of opinions in each cluster and calculate the generation probability of each semantically equivalent class : in is a semantic equivalence class, For this semantic equivalence class The viewpoint in It is a reminder engineering standard set by humans. It is manually set by the reviewer according to the different review contents, and is the standard for reviewing the corpus or the keywords that the corpus must meet.
[0010] Step 2.3: Calculate the semantic entropy based on the generation probability obtained in step 2.2 to determine the factual correctness of the corpus; Step 2.3.1: Based on the generation probability , the calculation formula of application semantic entropy is: ; Step 2.3.2, calculate the semantic entropy value of each semantic equivalence class; Step 2.3.3, summarize the semantic entropy values of all semantically equivalent classes to obtain the semantic entropy measure of the entire corpus; Step 2.3.4: Analyze the semantic entropy value to determine the factual correctness of the corpus. A lower semantic entropy value indicates a higher factual correctness. Step 2.3.5: Output the semantic entropy value of each corpus.
[0011] Step 2.4: According to the preset qualified standard k1 and unqualified standard k2, the corpus is divided into three categories: high score, medium score and low score, and the scoring results are output. Step 2.4.1, set the initial pass standard k1 and fail standard k2, these standards are adjusted according to domain knowledge and experience; Step 2.4.2, compare the semantic entropy value of each corpus with the qualified standard k1 and the unqualified standard k2; Step 2.4.3: Classify the corpus with semantic entropy value lower than k1 as high-scoring corpus; Step 2.4.4: Classify the corpus with semantic entropy values higher than k2 as low-scoring corpus, and the low-scoring corpus is considered unqualified corpus; Step 2.4.5: Classify the corpus with semantic entropy values between k1 and k2 as medium-point corpus; Step 2.4.6: Arrange the classification results and generate a JSON object containing the original corpus paragraph, semantic entropy value and scoring category for each corpus; Step 2.4.7: Output the scoring results of all corpora.
[0012] Step 3: Manually review the corpus with a medium score in the LLM review result and obtain the final score feedback; Step 3.1: Input the intermediate-scoring corpus output by LLM review into the public review module and have multiple reviewers perform scoring; Step 3.2: Statistically analyze the reviewers' scores, remove the significant outliers, and calculate the mean; Step 3.3, identify the disputed items and input them into the expert review module for re-examination; Step 3.4: The expert review module reviews the disputed items and determines whether they are qualified or not; Step 3.5: Integrate the original corpus and the final score, and output the score feedback as a JSON object.
[0013] Step 4: Automatically correct the unqualified corpus to obtain the corrected corpus; Step 4.1: Distribute the unqualified corpus to the manual rewriting and LLM rewriting modules according to the preset ratio; Unqualified corpora enter the manual rewriting module and LLM rewriting module according to the ratio of p and 1-p, and p obeys the following formula: Where m is the total number of corpus items reviewed, and the parameter It is used to adjust the rate of change of p according to the total number of corpora. According to the characteristics of the sigmoid function, as m gradually increases, p will follow a smooth decreasing trend and remain basically unchanged when m is greater than a certain threshold. That is, as the corpus accumulates, the manual workload in the manual rewriting module will gradually decrease. At the end of the process, the manual rewriting module and the LLM rewriting module jointly output the modified corpus. Step 4.2: Multiple people simultaneously rewrite the corpus that has entered the manual rewriting module to form modified corpus pairs, which are then entered into the modified corpus pair warehouse; Step 4.3: When the number of modified corpus pairs in the warehouse reaches a certain amount, they are input into the LLM update module to fine-tune the LLM; Step 4.4: Automatically rewrite the corpus entering the LLM rewriting module and output the corrected corpus.
[0014] Step 5: Review the corrected corpus again until it passes the review and becomes qualified corpus; Step 5.1: Re-enter the corrected corpus into the LLM review module for a new round of review; Step 5.2: Judge the audit result obtained in step 5.1. If it is qualified, output the qualified corpus. If it is unqualified, return to step 4 for further correction; Step 5.3: Repeat steps 5.1 and 5.2 until all corpora are reviewed and qualified.
[0015] Step 6: Collect scoring feedback and update the LLM audit model; Step 6.1, collect a certain amount of rating feedback, and calculate the relationship between semantic entropy and the final rating; Step 6.2, recalculate the qualified standard k1 and the unqualified standard k2 according to the statistical results; Step 6.3, when the target corpus review task is changed or the update is not ideal, modify the specified context x; Step 6.4: Use the updated k1, k2, and x to update the LLM audit model to improve the accuracy of subsequent audits.
[0016] Because the present invention adopts the above technical means, it has the following beneficial effects: 1. Corpus review module: This module improves the efficiency and accuracy of corpus review by combining LLM review and manual review. The LLM review part inputs the pre-processed corpus into the core review module and uses the deep analysis capabilities of the large language model for scoring. The manual review conducts secondary verification of the LLM review results through public review and expert review to ensure the accuracy of the review.
[0017] 2. Corpus correction module: This module automatically corrects unqualified corpora. Through pre-trained models and semantic analysis, it finds qualified corpora that are closest in semantics to the unqualified corpora as a reference for rewriting. This method retains the core semantics of the original corpus while generating more semantically accurate text.
[0018] 3. LLM audit model update process: By collecting scoring feedback, LLM is fine-tuned to continuously enhance its audit capabilities. This method makes LLM scoring more polarized, reduces the intermediate scoring corpus that requires manual second review, thereby reducing labor costs and improving audit efficiency and accuracy.
[0019] In summary, the present invention effectively improves the accuracy of multi-domain corpus review by combining LLM and manual review, and develops an automated correction mechanism, enhances the adaptability and flexibility of the model, and gradually reduces manual intervention through automated correction measures, thereby achieving a more efficient and accurate corpus intelligent review and automatic correction process. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 : Invention patent system block diagram; Figure 2 :Description diagram of the corpus review module; Figure 3 : LLM audit module illustration; Figure 4 : Manual review module; Figure 5 :Corpus correction module process; Figure 6 : LLM update module process. DETAILED DESCRIPTION
[0021] The following is a detailed description of the embodiments of the present invention. Although the present invention will be described and illustrated in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, modifications or equivalent substitutions made to the present invention should all be included in the scope of the claims of the present invention.
[0022] In addition, in order to better illustrate the present invention, numerous specific details are given in the following specific embodiments. It will be understood by those skilled in the art that the present invention can also be implemented without these specific details.
[0023] The purpose of this invention is to develop a corpus review and automatic correction method based on a large language model (LLM) and human-computer coupling, aiming to gradually reduce the dependence of domain data review on manual work, enhance the self-review and update capabilities of large language models, and realize automatic correction of review content. This method will improve the accuracy of multi-domain corpus review, enhance the adaptability and flexibility of the model, and gradually reduce human intervention through automated correction measures, thereby achieving a more efficient and accurate corpus intelligent review and automatic correction process, and providing a method for rapid review and automatic correction of domain data for large models in various fields.
[0024] like Figure 1 As shown, the overall framework of the present invention includes a corpus review module and a corpus correction module.
[0025] The corpus review module mainly reviews and scores the corpus, including LLM review and manual review. The corpus correction module corrects unqualified corpus (including low-scoring corpus) and inputs it into the review module again after correction until it passes the review and obtains qualified corpus. In addition, the review module will also feed back the scoring feedback of unqualified corpus as output to the correction module.
[0026] Corpus review module: like Figure 2 As shown in the figure, the corpus review module includes two modules: LLM review and manual review. The LLM review module includes the corpus preprocessing module and the core review module. The core review module is a large scoring model, which will divide the corpus into three categories according to the preset scoring criteria: high-scoring corpus, medium-scoring corpus and low-scoring corpus. Qualified corpus includes high-scoring corpus, and unqualified corpus includes low-scoring corpus. In addition, manual review is divided into public review, expert review and scoring analysis processing.
[0027] like Figure 3As shown in the figure, the LLM audit module includes two main modules: the corpus preprocessing module and the core audit language model. First, the original corpus from the Internet, manual collection or generated by its large language model is input into the corpus preprocessing module for preprocessing. The module standardizes the original corpus into JSON format objects and classifies and counts them according to the corpus type (such as single-turn dialogue, multi-turn dialogue and statement, etc.). Based on LLM's fuzzy recognition technology, the module is able to identify and correct the correct word order and punctuation when the text is damaged or incomplete. This fuzzy recognition technology allows the module to understand the intention of the text even when the information is incomplete or ambiguity exists. Subsequently, LLM conducts in-depth analysis and reorganization of the sentences using techniques such as syntactic analysis, part-of-speech tagging and semantic analysis. The reorganized text is further mapped into JSON format and finally output as standardized corpus.
[0028] The standardized corpus and the corrected corpus output by the correction module are used as input to the core review module. The core review module is a number of LLMs that have been fine-tuned by instructions to calculate the semantic entropy of the input corpus. Since the preprocessing module standardizes the format of the corpus, the main criterion for judging the quality of the corpus is the factual correctness of the corpus. The solution uses the calculation of semantic entropy to measure the factual correctness of the corpus.
[0029] Specific steps: Through the prompting project, the M large models let each LLM obtain the viewpoint of the current corpus. Since the process of generating corpus by the large model is random due to sampling, each LLM outputs K times and will get K different results. Specifically, the viewpoint output by model i (i=1,...,K) at the jth (j=1,...,M) is recorded as , then a total of K×M different opinions can be obtained. Here we introduce the concept of semantic entropy, which is used to quantify the uncertainty of the language model when generating text with specific meanings, and help evaluate the factual correctness of the corpus. After obtaining K×M opinions, each opinion Encode into semantic vectors through pre-trained models. Specifically, the pre-trained BERT model is selected to encode each viewpoint into a 768-dimensional semantic vector. Then, the K-means clustering algorithm is used to divide this batch of corpora into different semantic equivalence classes based on semantics (cosine similarity of embedding vectors). For each semantic equivalence class, the generation probability of each corpus in the class is calculated. It is estimated by the token probability given by the pre-trained model, that is, the conditional probability of the model for each corpus: . Add the conditional probabilities of each cluster to get the semantic likelihood of that cluster .have: Where c is the semantic equivalence class, s is the viewpoint in the semantic equivalence class (c), and x is the manually set prompt engineering standard. x is manually set by the reviewer according to the different review contents, and is the standard of the review corpus or the keyword that the corpus needs to meet. The semantic entropy (SE) of the semantic set is calculated based on the following formula: The semantic entropy is used to measure the factual correctness of the corpus. The lowest calculated value of semantic entropy is 0. The higher the semantic entropy, the worse the factual correctness of the corpus. Set the qualification standard k 1 and failure standard k 2 When the semantic entropy is lower than k 1 When the semantic entropy is higher than k 2 When the corpus is considered to be a low-scoring corpus, the semantic entropy is between k 1 With k 2 The corpus in between is considered as the middle score corpus and is input into the manual review module. 1 With k 2 It is necessary to set the extreme value at the beginning of the audit, that is, k 1 As small as possible and k 2 Try to make it as large as possible to facilitate subsequent adjustments. The output format of each corpus is a JSON object containing two keys: the original corpus paragraph and the semantic entropy. For example, {"corpus": "Sample corpus", "SE": 0.68}.
[0030] like Figure 4 As shown in the figure, the manual review module includes public review, expert review and score analysis processing module. Specifically, the intermediate review corpus output by the LLM review module is used as input and first output to the public review module. The public review outputs the reviewer score and enters the score analysis processing module. The module sends the items with large score differences to the expert review module for re-examination. Finally, the qualified and unqualified corpus and the corresponding score feedback are output.
[0031] Scoring process of the public review module: Each corpus is scored by n reviewers, and n is generally considered to be at least 5; the scoring is done using an X-point system, which is the same as the LLM review scoring system. In order to meet the requirements of strict review, the reviewers score S H Need to meet S H ≤X×korS H ≥X×(1-k). Taking X = 10, k = 0.2 as an example, the reviewers can only give 0-2 points and 8-10 points on a 10-point scale: 0 points means that it does not meet the requirements of the corpus at all or has obvious errors; 10 points means that it fully meets the requirements of the corpus and is innovative.
[0032] Rating analysis and processing module process: The module receives the reviewer rating S from the public review H. Remove significant outlier scores (for example, if there are 5 reviewers with 4 high scores and 1 low score, remove the low score) and calculate the mean . Identify controversial items (taking 5 reviewers as an example, such as 3 high scores and 2 low scores) and input them into the expert review module for re-examination. The large model has powerful mathematical processing capabilities and can effectively identify the above situations. For non-controversial items, the average of the reviewers' scores is considered to be the final score; for controversial items, the re-examination score is taken as the final score. The analysis and processing module integrates the content including the original corpus and the final score into a JSON object and outputs it as score feedback, for example {"corpus": "sample corpus",, "FinalScore": 1.6}.
[0033] Expert review module process: The expert review module receives disputed items marked by the scoring analysis and processing module. Each corpus is reviewed by m experts. For labor cost considerations, it is generally believed that m < n. Experts only judge whether it is qualified or not, and the evaluation is N (No, no / unqualified) or Y (Yes, yes / qualified). q is the re-examination pass rate parameter, and p is the expert pass rate, that is, the proportion of m experts who are evaluated as Y. For q, 0.5≤q≤1. When q takes the minimum value of 0.5, at least half of the experts pass. The value of q increases with the requirements for corpus quality. For corpora with extremely strict requirements, q takes the maximum value of 1, that is, the expert has a veto. For the X-point system, when q≠1, the re-examination score S E It can be expressed as: When q=1, that is, the expert has a veto power, the re-examination score S E It can be expressed as: Corpus correction module: The corpus correction module consists of a manual rewriting module, a modified corpus pair warehouse, an LLM update module, and an LLM rewriting module. Figure 5 As shown: Unqualified corpus enters the manual rewriting module and LLM rewriting module according to the ratio of p and 1-p. p obeys the following formula: Where m is the number of corpora reviewed, and parameter k is used to adjust the rate of change of p according to the total number of corpora. Referring to the characteristics of the sigmoid function, as m gradually increases, p will follow a smooth decreasing trend and remain basically unchanged when m is large enough. That is, as the corpus accumulates, the manual workload in the manual rewriting module will gradually decrease. At the end of the process, the manual rewriting module and the LLM rewriting module jointly output the modified corpus.
[0034] The manual rewriting module requires multiple people to rewrite the same corpus at the same time according to the corpus requirements. The requirements can be the prompt word engineering standards in the review module. After rewriting, the original corpus + modified corpus are entered into the modified corpus pair warehouse in pairs. When the modified corpus pair warehouse accumulates a certain number, it is entered into the LLM update module, which fine-tunes the LLM that performs the LLM rewriting module task and aligns the LLM modification standards with the manual standards. In this way, the workload of manual rewriters is gradually reduced.
[0035] The LLM update module executes the LLM update module process when the modified corpus pair warehouse reaches a certain number. The modified corpus pairs are provided by the modified corpus pair warehouse to enter the semantic clustering module. The semantic clustering module adopts the same strategy as the semantic clustering in the review module, and the object is the plural modified corpus rewritten by manual rewriters for the same corpus. The semantic clustering module will output the same original corpus and the rewritten corpus classified by style. After fine-tuning multiple pre-trained models according to the rewritten corpus of different styles, the LLMs of different corpus language styles are weighted and fused according to the modification goals to update the LLM rewriting module.
[0036] The LLM rewriting module is a module that is specifically used to rewrite the corpus after several fine-tuning. After inputting unqualified corpus, it will automatically output the modified corpus.
[0037] LLM audit model update process: The format of the scoring feedback is a JSON object containing the original corpus, semantic entropy, and final score, for example {"corpus": "sample corpus", "SE":1.3,"FinalScore": 1.6}. After collecting i (the size of i depends on the amount of review tasks, usually more than 50) scoring feedback, the review process is interrupted and the update begins. Based on the scoring feedback of all the medium-scoring corpora, the relationship between semantic entropy and the final score is statistically calculated: the trimmed mean of the semantic entropy of the final qualified corpus is calculated and compared with the qualified standard k 1 After re-weighted averaging, we get a new k 1 Similarly, the semantic entropy trimmed mean value of the final unqualified corpus is calculated and compared with the unqualified standard k 2 After re-weighted averaging, we get a new k 2 , for example, when the weighting ratio is 1:1, the new standard is a compromise between the old standard and the trimmed mean of semantic entropy. The weighting weight can be adjusted as appropriate according to the calculation effect. At the same time, when the target corpus review task is changed, or when several updates are not ideal, the reviewer should modify the specified context x. 1 and k 2 The concept describes that k should be used in the initial round 1 As small as possible and k 2 As large as possible, so that k 1 and k2 After the update, k 1 -k 2 The value of is reduced, that is, the number of medium-scoring corpora input to manual review in the next round of review tasks is reduced, which improves the accuracy of the review and reduces the workload of manual review. Considering the ideal situation, that is, there is no obvious difference in the quality of the corpus in each round, after a certain number of rounds of updates, k 1 -k 2 The value of tends to a stable value, with slight fluctuations and lower than the initial ratio. It is believed that the workload of manual review is reduced, and a more efficient and accurate corpus intelligent review and automatic correction process is achieved.
Claims
1. A method for reviewing and automatically correcting domain corpus data based on a generative model, characterized in that: The following steps are involved: Step 1: Preprocess the original corpus to obtain standardized corpus; Step 2: Perform LLM review on the standardized corpus and obtain the scoring results. LLM stands for Large Language Model. Step 3: Manually review the corpus with a medium score in the LLM review result and obtain the final score feedback; Step 4: Automatically correct the unqualified corpus to obtain the corrected corpus; Step 5: Review the corrected corpus again until it passes the review and becomes qualified corpus; Step 6: Collect scoring feedback and update the LLM audit model.
2. According to the method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 1, it is characterized in that: Step 1 includes the following steps: Step 1.1: Standardize the format of the original corpus from the Internet, manually collected, or generated by a large language model and convert it into a JSON format object; Step 1.2: Classify and count the JSON format objects obtained in step 1.1, and classify them according to the corpus type, including single-round dialogue, multi-round dialogue and statement; Step 1.3: Using the LLM-based fuzzy recognition technology, correct the word order and punctuation of the classified corpus obtained in step 1.2; Step 1.4, performing syntactic analysis, part-of-speech tagging, semantic analysis and sentence reorganization on the revised corpus obtained in step 1.3 to obtain a reorganized text; Step 1.5: Map the reorganized text into JSON format and output it as standardized corpus.
3. According to the method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 1, it is characterized in that: Step 2 includes the following steps: Step 2.1, multiple LLMs generate opinions on the current corpus according to the prompt project, each LLM generates K opinions, and a total of K×M opinions are obtained; Step 2.2: Perform cluster analysis on the opinions obtained in step 2.1 and calculate the generation probability of each semantically equivalent class; Step 2.3: Calculate the semantic entropy based on the generation probability obtained in step 2.2 to determine the factual correctness of the corpus; Step 2.4: According to the preset qualified standard k1 and unqualified standard k2, the corpus is divided into three categories: high score, medium score and low score, and the scoring results are output.
4. According to the method of claim 3, the method is characterized in that: Step 2.1 includes the following steps: Step 2.1.1, select multiple LLMs that have been fine-tuned by instructions; Step 2.1.2: Segment the standardized corpus into paragraphs or sentences suitable for LLM processing; Step 2.1.3: Design a prompt project for each LLM to guide the model to generate opinions about the correctness of the corpus facts; Step 2.1.4: Input the segmented corpus into each LLM to generate multiple different viewpoints ; Step 2.1.5: Repeat step 2.1.4, and each LLM generates K different views to increase the diversity of the results; Step 2.1.6: Collect all the opinions generated by LLM, totaling K×M.
5. According to claim 4, a method for reviewing and automatically correcting domain corpus data based on a generative model is characterized in that: Step 2.2 includes the following steps: Step 2.2.
1. Use the pre-trained BERT model to transform each opinion Encoded as a semantic vector of fixed dimension; Step 2.2.2, calculate the cosine similarity between all semantic vectors as the similarity measure for cluster analysis; Step 2.2.3, apply K-means clustering algorithm to divide the semantic vectors into multiple clusters according to cosine similarity, i.e., semantic equivalence classes; Step 2.2.4: Count the number of opinions in each cluster and calculate the generation probability of each semantically equivalent class : in is a semantic equivalence class, For this semantic equivalence class The viewpoint in It is a reminder engineering standard set by humans. Manually set by the reviewer based on the review content, which is the standard for reviewing the corpus or the keywords that the corpus must meet; Step 2.3 includes the following steps: Step 2.3.1: Based on the generation probability , the calculation formula of application semantic entropy is: ; Step 2.3.2, calculate the semantic entropy value of each semantic equivalence class; Step 2.3.3, summarize the semantic entropy values of all semantically equivalent classes to obtain the semantic entropy measure of the entire corpus; Step 2.3.4: Analyze the semantic entropy value to determine the factual correctness of the corpus. The lower the semantic entropy value, the higher the factual correctness. Step 2.3.5: Output the semantic entropy value of each corpus.
6. A method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 5, characterized in that: Step 2.4 includes the following steps: Step 2.4.1, set the initial pass standard k1 and fail standard k2, these standards are adjusted according to domain knowledge and experience; Step 2.4.2, compare the semantic entropy value of each corpus with the qualified standard k1 and the unqualified standard k2; Step 2.4.3: Classify the corpus with semantic entropy value lower than k1 as high-scoring corpus; Step 2.4.4: Classify the corpus with semantic entropy values higher than k2 as low-scoring corpus, and the low-scoring corpus is considered unqualified corpus; Step 2.4.5: Classify the corpus with semantic entropy values between k1 and k2 as medium-point corpus; Step 2.4.6: Arrange the classification results and generate a JSON object containing the original corpus paragraph, semantic entropy value and scoring category for each corpus; Step 2.4.7: Output the scoring results of all corpora.
7. The method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: Input the intermediate-scoring corpus output by LLM review into the public review module and have multiple reviewers perform scoring; Step 3.2: Statistically analyze the reviewers' scores, remove the significant outliers, and calculate the mean; Step 3.3, identify the disputed items and input them into the expert review module for re-examination; Step 3.4: The expert review module reviews the disputed items and determines whether they are qualified or not; Step 3.5: Integrate the original corpus and the final score, and output the score feedback as a JSON object.
8. The method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 1, characterized in that: Step 4 includes the following steps: Step 4.1: Distribute the unqualified corpus to the manual rewriting and LLM rewriting modules according to the preset ratio; Unqualified corpora enter the manual rewriting module and LLM rewriting module according to the ratio of p and 1-p, and p obeys the following formula: in m The cumulative number of corpus items reviewed, parameter Used to adjust according to the total number of corpora p The rate of change, according to the characteristics of the sigmoid function, as m Gradually increase, p will follow a smooth decreasing trend and m When the value is greater than a certain threshold, it remains basically unchanged. That is, as the corpus accumulates, the manual workload in the manual rewriting module will gradually decrease. At the end of the process, the manual rewriting module and the LLM rewriting module jointly output the modified corpus. Step 4.2: Multiple people simultaneously rewrite the corpus that has entered the manual rewriting module to form modified corpus pairs, which are then entered into the modified corpus pair warehouse; Step 4.3: When the number of modified corpus pairs in the warehouse reaches the preset number, it is input into the LLM update module to fine-tune the LLM; Step 4.4: Automatically rewrite the corpus entering the LLM rewriting module and output the corrected corpus.
9. The method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 1, characterized in that: Step 5 includes the following steps: Step 5.1: Re-enter the corrected corpus into the LLM review module for a new round of review; Step 5.2: Judge the audit result obtained in step 5.
1. If it is qualified, output the qualified corpus. If it is unqualified, return to step 4 for further correction; Step 5.3: Repeat steps 5.1 and 5.2 until all corpora are reviewed and qualified.
10. A method for reviewing and automatically correcting domain corpus data based on a generative model according to claim 1, characterized in that: Step 6 includes the following steps: Step 6.1, collect a preset number of rating feedbacks, and calculate the relationship between semantic entropy and the final rating; Step 6.2, recalculate the qualified standard k1 and the unqualified standard k2 according to the statistical results; Step 6.3, when the target corpus review task is changed or the update is not ideal, modify the specified context x; Step 6.4: Use the updated k1, k2, and x to update the LLM audit model to improve the accuracy of subsequent audits.
Citation Information
Patent Citations
Document auditing method, device and equipment and storage medium
CN116663525A
Large language model data processing method based on deep understanding of user semantics
CN117556827A
Method for generating double-model corpus
CN118350459A
Systems and methods for using image scoring for an improved search engine
US20250037422A1
Cited By
Fault removal agent training method and device and electronic equipment
CN122241242A