Evaluation feedback reinforcement learning-based error suppression protection method and system, and storage medium
Through the method of reinforcement learning based on evaluation feedback, the multi-iteration optimization method is adopted to significantly improve the output quality and reliability of the model in response to the problem that generative artificial intelligence models are prone to generate false information and inaccurate references.
Patent Information
- Application Number
- CN202510208375.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
AI Technical Summary
The existing generative artificial intelligence model is prone to generating false information and inaccurate references, which limits its application in high accuracy demand scenarios.
The error suppression protection method based on evaluation feedback reinforcement learning is adopted to generate natural language feedback from multiple evaluation dimensions, continuously iterate and optimize the model, and improve the reliability and security of generative artificial intelligence.
Through multiple iteration optimization, the output quality of the generative artificial intelligence model is significantly improved, the generation of false information is reduced, the accuracy of reference is improved, and the reliability and security of the model are enhanced.
Smart Images

Figure CN120124602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and more specifically, to an error suppression and protection method, system, and storage medium based on evaluation feedback reinforcement learning. Background Art
[0002] Currently, with the booming development of artificial neural networks and the widespread application of generative artificial intelligence and large language models, traditional text generation and information retrieval technologies have gradually exposed problems such as inconsistent generated content, false information, and inaccurate citations. These defects limit the widespread application of large models in many practical scenarios. To address this challenge, in recent years, researchers have proposed various techniques to improve the output quality of large models, including retrieval-enhanced generation methods and feedback-driven self-optimization methods.
[0003] Retrieval enhancement techniques significantly improve the accuracy and citation quality of generated text by combining the model's generation process with external knowledge sources. This technique can reduce the generation of false information by retrieving relevant evidence from an external knowledge base and incorporating it into the model generation. However, existing retrieval enhancement techniques still face problems such as citation errors and generated information not being included in the retrieval evidence, thus limiting their application in scenarios with high accuracy requirements.
[0004] Feedback optimization techniques, on the other hand, provide step-by-step optimized feedback for model generation by mimicking the way humans improve text. Current mainstream methods are mostly based on reinforcement learning and self-optimization strategies, using the language model to improve the intrinsic evaluation of the initial output. However, due to the limitations of the language model's self-evaluation ability, this intrinsic evaluation method is difficult to effectively identify and correct factual errors in the generated content, resulting in limited self-optimization effects.
[0005] Therefore, there is an urgent need to provide an error suppression and protection method based on evaluation feedback reinforcement learning to solve the problems of easy generation of false information and inaccurate citations faced by existing generative artificial intelligence models. Summary of the Invention
[0006] In view of this, the present invention aims to solve the problems of easy generation of false information and inaccurate citations faced by existing generative artificial intelligence models, and provides an error suppression and protection method, system, and storage medium based on evaluation feedback reinforcement learning, generating natural language feedback from multiple evaluation dimensions, continuously iterating and optimizing the model, improving the reliability and security of generative artificial intelligence, and providing new ideas and technical support for the highly reliable application of generative artificial intelligence.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An error suppression and protection method based on evaluation feedback reinforcement learning, comprising:
[0009] Step 1, Initial Response Generation Phase: Receive the initial input sequence and relevant materials, and generate an initial output based on the task-specific generation prompt;
[0010] Step 2, Output Quality Evaluation Phase: Based on the evaluation metrics, use an explicit model to compare the initial output with the correct facts to obtain a quantitative quality evaluation;
[0011] Step 3, Evaluation Feedback Phase: Based on the evaluation metrics, the initial input sequence, and the task-specific feedback prompt, use a language large model to generate evaluation feedback that is easy for the large model to understand for each evaluation metric;
[0012] Step 4, Iterative Learning Output Phase: Based on the initial input sequence, relevant materials, initial output, evaluation feedback, and task-specific output prompt, use a language large model to generate a new output, and then use this output as the initial output, repeat Steps 2 to 4 iteratively until the evaluation metrics meet the requirements to generate an improved output.
[0013] Optionally, the method for generating the initial output in Step 1 specifically includes:
[0014] Y t = M([p init ; x; e])
[0015] where Y t is the initial output, M is the language large model, p init is the task-specific generation prompt, x is the initial input sequence, e is the relevant materials, and [·; ·] represents concatenation.
[0016] Optionally, obtaining the quantitative quality evaluation in Step 2 specifically includes:
[0017] Based on a correct fact y and evaluation metrics δ = {ε 1 , ε 2 , …, ε n}, then there is:
[0018] S t = [ε 1 (Y t , y), …, ε n (Y t , y)] ∈ R n
[0019] where S t is the quantitative quality evaluation, ε n is the model output quality metric, ε n (Y t , y) indicates that the comparison object of the metric is the initial output and the correct fact, and R nis an n-dimensional real space.
[0020] Optionally, the method for generating the evaluation feedback in step 3 specifically includes:
[0021] Based on the task-specific feedback prompt P fb , then:
[0022] F t = M([p fb ; S t ; x; e; Y t )
[0023] where F t is the evaluation feedback in natural language form, M is a language large model, S t is the quantified quality evaluation, x is the initial input sequence, e is the relevant material, and Y t is the initial output.
[0024] Optionally, the specific output method for each iteration in step 4 includes:
[0025] Given the task-specific refined output prompt p refine , then:
[0026] Y t+1 = M([p refine ; x; e; Y t ; F t )
[0027] where Y t+1 is the output after one iteration improvement of Y t , M is a language large model, x is the initial input sequence, e is the relevant material, Y t is the initial output, and F t is the evaluation feedback in natural language form.
[0028] Optionally, the method for generating the improved output in step 4 specifically includes:
[0029] When the evaluation index reaches the expectation or the number of iterations reaches the set upper limit in the iteration, the improved input H is defined as:
[0030] H = [p refine ; x; e; Y 0 ; F 0 , …, Y k , F k
[0031] where p refine is the task-specific refined output prompt, Y k is the output after k iterations, and F kThe evaluation feedback in natural language form after k iterations, and finally the improved output y is obtained. output :
[0032] y output = M(H)
[0033] Where M is a language large model and H is the improved input.
[0034] An error suppression and protection system based on evaluation feedback reinforcement learning, including:
[0035] Initial response generation module: Receiving an initial input sequence and relevant materials, and generating an initial output based on task-specific generation prompts.
[0036] Quality evaluation output module: Based on evaluation metrics, using an explicit model to compare the initial output with the correct facts to obtain a quantified quality evaluation.
[0037] Evaluation feedback module: Based on evaluation metrics, the initial input sequence, and task-specific feedback prompts, using a language large model to generate evaluation feedback that is easy for the large model to understand for each evaluation metric.
[0038] Iterative learning output module: Based on the initial input sequence, relevant materials, initial output, evaluation feedback, and task-specific output prompts, using a language large model to generate a new output, and then using this output as the initial output, repeating the functions of the quality evaluation output module and the evaluation feedback module iteratively until the evaluation metrics meet the requirements, and generating an improved output.
[0039] A storage medium stores a computer program, and when the computer program runs on a computer, it causes the computer to execute the error suppression and protection method based on evaluation feedback reinforcement learning described above.
[0040] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an error suppression and protection method, system and storage medium based on evaluation feedback reinforcement learning, including: Step 1, initial response generation stage: receiving an initial input sequence and related materials, and generating an initial output based on task-specific generation prompts; Step 2, output quality evaluation stage: based on evaluation metrics, using an explicit model to compare the initial output with the correct facts to obtain a quantified quality evaluation; Step 3, evaluation feedback stage: based on evaluation metrics, the initial input sequence and task-specific feedback prompts, using a language large model to generate evaluation feedback that is easy for the large model to understand corresponding to each evaluation metric; Step 4, iterative learning output stage: based on the initial input sequence, related materials, initial output, evaluation feedback and task-specific output prompts, using a language large model to generate a new output, and then using this output as the initial output, repeating Steps 2 to 4 iteratively until the evaluation metrics meet the requirements, and generating an improved output. The present invention aims to solve the problems of easy generation of false information and inaccurate citations faced by existing generative artificial intelligence models, generate natural language feedback from multiple evaluation dimensions, continuously iterate and optimize the model, improve the reliability and security of generative artificial intelligence, and provide a new idea and technical support for the high-reliability application of generative artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0042] Figure 1 It is a schematic flowchart of the method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0044] The embodiments of the present invention disclose an error suppression and protection method based on evaluation feedback reinforcement learning, as Figure 1 shown, including:
[0045] Step 1, initial response generation stage: receiving an initial input sequence and related materials, and generating an initial output based on task-specific generation prompts;
[0046] Step 2, Output Quality Evaluation Phase: Based on the evaluation metrics, use an explicit model to compare the initial output with the correct facts to obtain a quantitative quality evaluation;
[0047] Step 3, Evaluation Feedback Phase: Based on the evaluation metrics, the initial input sequence, and task-specific feedback prompts, use a large language model to generate evaluation feedback that is easy for the large model to understand for each evaluation metric;
[0048] Step 4, Iterative Learning Output Phase: Based on the initial input sequence, relevant materials, the initial output, the evaluation feedback, and task-specific output prompts, use a large language model to generate a new output, and then use this output as the initial output, repeating Steps 2 to 4 iteratively until the evaluation metrics meet the requirements to generate an improved output.
[0049] In a specific embodiment, the method for generating the initial output in Step 1 specifically includes:
[0050] Y t = M([p init ; x; e])
[0051] where Y t is the initial output, M is the large language model, p init is the task-specific generation prompt, x is the initial input sequence, e is the relevant material, and [·; ·] represents concatenation.
[0052] In this embodiment, Table 1 gives a schematic diagram of p init The content that the generation prompt should include is: input format, output format, emphasizing features such as objectivity and accuracy that the output should possess.
[0053] Table 1 Schematic Table of Task-Specific Generation Prompts
[0054]
[0055] In a specific embodiment, the specific process of obtaining a quantitative quality evaluation in Step 2 includes:
[0056] Based on a correct fact y and evaluation metrics δ = {ε 1 , ε 2 , …, ε n}, then:
[0057] S t = [ε 1 (Y t , y), …, ε n (Y t , y)] ∈ R n
[0058] where St For the quantified quality evaluation, ε n For the quality metric of the model output, ε n (Y t , y) indicates that the comparison object of the metric is the initial output and the correct fact.
[0059] In this embodiment, ε 1 The accuracy rate is adopted, ε 2 The citation reliability rate is adopted, ε 3 The citation precision rate is adopted. The accuracy rate calculates the similarity between the initial output and the correct fact through indicators such as BLEU and ROUGE to quantify the accuracy. The citation reliability rate is quantified through source authority analysis and knowledge base comparison, etc. The citation precision rate extracts the semantic features of the citation content generated by the model and the source text, and conducts semantic comparison to complete the quantification.
[0060] In a specific embodiment, the method for generating the evaluation feedback in step 3 specifically includes:
[0061] Based on the task-specific feedback prompt p fb , then there is:
[0062] F t = M([p fb ; S t ; x; e; Y t )
[0063] Among them, F t is the evaluation feedback in natural language form, M is the language large model, S t is the quantified quality evaluation, x is the initial input sequence, e is the relevant material, Y t is the initial output.
[0064] In this embodiment, Table 2 gives the schematic diagram of p fb . The content that the feedback prompt should include is: input format, feedback description, and output format. In the feedback description, each metric is separated, and the output content when the quality evaluation corresponding to each metric meets the requirements and does not meet the requirements is pointed out.
[0065] Table 2 Schematic Table of Task-Specific Feedback Prompt
[0066]
[0067] In a specific embodiment, the specific output method for each iteration in step 4 includes:
[0068] Given the task-specific refined output prompt p refine , then:
[0069] Y t+1 = M([prefine ; x; e; Y t ; F t )
[0070] Among them, Y t+1 is the output after one iteration of improvement, M is the language large model, x is the initial input sequence, e is the relevant material, and Y t is the initial output, and F t is the evaluation feedback in natural language form. t
[0071] In this embodiment, p refine should include the following content: input format, output format, and emphasize that the initial output and its corresponding evaluation feedback should be mainly referred to when outputting.
[0072] In a specific embodiment, the method for generating the improved output in step 4 specifically includes:
[0073] When the evaluation index reaches the expectation or the number of iterations reaches the set upper limit in the iteration, the improved input H is defined as:
[0074] H = [p refine ; x; e; y 0 ; F 0 , …, Y k , F k
[0075] Among them, p refine is the refined output prompt specific to the task, Y k is the output after k iterations, F k is the evaluation feedback in natural language form after k iterations, and finally the improved output y output :
[0076] y output = M(H)
[0077] Among them, M is the language large model and H is the improved input.
[0078] In this embodiment, the improved output traces back all the evaluations and iterations, and the optimal result is not achieved in an independent attempt.
[0079] This embodiment aims at the problems of easy generation of false information and inaccurate citation faced by current generative large models, and proposes an error suppression and protection method based on evaluation feedback reinforcement learning. This method can generate natural language feedback from multiple evaluation dimensions to guide the iterative optimization of the model and improve the reliability and security of generative artificial intelligence.
[0080] An error suppression and protection system based on evaluation feedback reinforcement learning includes:
[0081] Initial response generation module: Receives an initial input sequence and related materials, and generates an initial output based on a task-specific generation prompt;
[0082] Quality evaluation output module: Based on evaluation metrics, uses an explicit model to compare the initial output with the correct facts to obtain a quantified quality evaluation;
[0083] Evaluation feedback module: Based on evaluation metrics, the initial input sequence, and task-specific feedback prompts, uses a large language model to generate evaluation feedback that is easy for the large model to understand for each evaluation metric;
[0084] Iterative learning output module: Based on the initial input sequence, related materials, initial output, evaluation feedback, and task-specific output prompts, uses a large language model to generate a new output, and then uses this output as the initial output, repeating the functions of the quality evaluation output module and the evaluation feedback module iteratively until the evaluation metrics meet the requirements, generating an improved output.
[0085] A storage medium stores a computer program, and when the computer program runs on a computer, it causes the computer to execute an error suppression protection method based on evaluation feedback reinforcement learning.
[0086] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0087] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An error suppression protection method based on evaluation feedback reinforcement learning, characterized in that: include: Step 1: Initial response generation phase: receiving the initial input sequence and related materials, and generating the initial output based on task-specific generation prompts; Step 2: Output quality evaluation phase: Based on the evaluation indicators, the explicit model is used to compare the initial output with the correct facts to obtain a quantitative quality evaluation; Step 3: Evaluation and feedback phase: Based on the evaluation indicators, the initial input sequence, and task-specific feedback prompts, the language model is used to generate evaluation feedback that is easy for the model to understand for each evaluation indicator. Step 4, iterative learning output phase: Based on the initial input sequence, relevant materials, initial output, evaluation feedback and task-specific output prompts, use the language model to generate a new output, and then use this output as the initial output, repeat steps 2 to 4, iterate until the evaluation indicators meet the requirements, and generate an improved output.
2. According to claim 1, the error suppression protection method based on evaluation feedback reinforcement learning is characterized in that: The method for generating the initial output in step 1 specifically includes: Y t =M([p init ;x;e]) Among them, Y t is the initial output, M is the language model, p init is the task-specific generation prompt, x is the initial input sequence, e is the related material, and [·;·] indicates a connection.
3. The error suppression protection method based on evaluation feedback reinforcement learning according to claim 1 is characterized in that: The quantitative quality evaluation obtained in step 2 specifically includes: Based on a correct fact y and evaluation index δ={ε1,ε2,···,ε n }, then: S t =[ε1(Y t ,and),···,ε n (AND t ,and)]∈R n Among them, S t For quantitative quality evaluation, ε n is the model output quality metric, ε n (Y t ,y) indicates that the comparison object of the indicator is the initial output and the correct fact, R n is an n-dimensional real number space.
4. The error suppression protection method based on evaluation feedback reinforcement learning according to claim 1 is characterized in that: The method for generating evaluation feedback in step 3 specifically includes: Based on task-specific feedback prompts fb , then: F t =M([p fb ;S t ;x;e;Y t ]) Among them, F t is the evaluation feedback in natural language form, M is the language model, S t is a quantitative quality evaluation, x is the initial input sequence, e is the related material, Y t is the initial output.
5. The error suppression protection method based on evaluation feedback reinforcement learning according to claim 1 is characterized in that: The specific output methods for each iteration in step 4 include: Given a task-specific refinement output hint p refine ,but: Y t+1 =M([p refine ;x;e;Y t ;F t ]) Among them, Y t+1 Y t After an iterative improvement, M is the language model, x is the initial input sequence, e is the relevant material, and Y t is the initial output, F t Evaluation feedback in natural language form.
6. The error suppression protection method based on evaluation feedback reinforcement learning according to claim 1 is characterized in that: The method for generating the improved output in step 4 specifically includes: When the evaluation index reaches the expected value or the number of iterations reaches the upper limit in the iteration, the improved input H is defined as: H=[p refine ;x;e;Y0;F0,···,Y k ,F k ] Among them, p refine is a task-specific refined output prompt, x is the initial input sequence, e is the related material, and Y k is the output after k iterations, F k is the evaluation feedback in natural language after k iterations, and the final improved output y output : y output =M(H) Among them, M is the language model and H is the improved input.
7. An error suppression protection system based on evaluation feedback reinforcement learning, characterized in that: The error suppression protection method based on evaluation feedback reinforcement learning according to any one of claims 1 to 6 is applied, comprising: Initial response generation module: receives the initial input sequence and related materials, and generates initial output based on task-specific generation prompts; Quality evaluation output module: Based on the evaluation indicators, the explicit model is used to compare the initial output with the correct facts to obtain a quantitative quality evaluation; Evaluation and feedback module: Based on the evaluation indicators, the initial input sequence and the task-specific feedback prompts, the language model is used to generate evaluation feedback that is easy for the model to understand for each evaluation indicator; Iterative learning output module: Based on the initial input sequence, relevant materials, initial output, evaluation feedback and task-specific output prompts, a large language model is used to generate a new output, and then the output is used as the initial output to repeatedly execute the functions of the quality evaluation output module and the evaluation feedback module, iterating until the evaluation indicators meet the requirements and generating an improved output.
8. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program runs on a computer, the computer is enabled to execute the error suppression protection method based on evaluation feedback reinforcement learning as described in any one of claims 1 to 6.