Reinforcement learning training method, system and device for government affair industry large model and medium

By building a government industry model with basic language understanding ability and government affairs knowledge, setting up a government affairs environment simulator, defining multi-dimensional reward functions, using reinforcement learning algorithms for strategy learning, and combining user feedback optimization training, the government affairs model has been solved inadequate understanding of professional knowledge and the accuracy of generated content, and efficient and accurate government affairs task processing is achieved.

CN120494032APending Publication Date: 2025-08-15SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510524826.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing large-scale model has insufficient professional knowledge in the field of government affairs, the accuracy and reliability of the generated content need to be improved, and there is a lack of effective mechanisms for self-optimization and adjustment, making it difficult to adapt to the diversity and complexity of government affairs tasks.

Method used

Build a government industry big model with basic language understanding ability and government affairs knowledge, set up a government affairs environment simulator, define state information, action information and multi-dimensional reward functions, adopt reinforcement learning algorithms for strategy learning, and optimize training combined with user feedback, and regularly evaluate model performance.

Benefits of technology

The performance of the government affairs industry big model in various government affairs scenario tasks has been improved, and it can handle government affairs tasks more efficiently and accurately, reduce manual intervention, and improve the efficiency and information quality of government affairs work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494032A_ABST
    Figure CN120494032A_ABST
Patent Text Reader

Abstract

The invention provides a reinforcement learning training method, system and device for a government affair industry large model and a medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: constructing a large government affair industry model with basic language understanding ability and government affair knowledge, and initializing the large government affair industry model; setting a government affair environment simulator for the government affair industry large model to simulate various government affair scene tasks; defining state information, action information and a reward function for the government affair industry large model to guide the model to generate output meeting requirements; learning the strategy of the government affair industry large model by using a reinforcement learning algorithm so as to learn the optimal strategy through interaction with the government affair environment; collecting user feedback in the government affair scene as reinforcement learning training of an additional reward signal optimization model; and regularly evaluating the performance of the government affair industry large model, and adjusting and optimizing the model according to an evaluation result. According to the method, efficient learning and performance optimization of the government affair industry large model in various government affair scene tasks are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and more specifically relates to a reinforcement learning training method, system, equipment and medium for a large model of the government industry. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, large language models have been widely applied in various fields, including government affairs. Large models can be used for tasks such as intelligent customer service, policy interpretation, and government document processing, improving the efficiency and quality of government work. However, existing large models have some limitations in government scenarios, such as insufficient understanding of government domain expertise and the need to improve the accuracy and reliability of generated content. These issues have limited the widespread application and promotion of large models in government affairs.

[0003] The government sector is uniquely characterized by numerous policies, regulations, specialized terminology, and complex business processes. Existing large models often lack sufficient training and optimization in these areas, leading to potential biases or errors when handling government tasks. For example, when interpreting policies, the model may fail to accurately grasp the policy's intent and details, resulting in content that does not conform to policy requirements. Similarly, when writing official documents, the model may fail to adhere to strict document formatting and language standards, impacting the document's quality and credibility.

[0004] Furthermore, the diversity and complexity of government tasks place higher demands on model adaptability. Different government tasks may involve different domain knowledge and objectives, requiring models to flexibly adjust their strategies to meet diverse task requirements. Existing large models also fall short in this regard, lacking effective mechanisms for self-optimization and adjustment based on task feedback. Summary of the Invention

[0005] In response to the above problems, the purpose of the present invention is to provide a reinforcement learning training method, system, equipment and medium for large models of the government affairs industry. By constructing a model with government affairs knowledge, setting up a simulator, defining state actions and reward functions, using reinforcement learning algorithm training, combining user feedback optimization and regular evaluation and adjustment, efficient learning and performance optimization of large models of the government affairs industry in various government affairs scenario tasks are achieved.

[0006] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: In a first aspect, an embodiment of the present application provides a reinforcement learning training method for a large model of the government affairs industry, including: Build a large model of the government affairs industry with basic language understanding capabilities and government affairs knowledge, and initialize the large model of the government affairs industry; Set up a government environment simulator for the government industry large model to simulate various government scenario tasks; Define state information, action information, and reward functions for large models in the government sector to guide the model to generate outputs that meet requirements; Use reinforcement learning algorithms to learn the strategies of large models in the government sector, so as to learn the optimal strategies through interaction with the government environment; Collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model; Regularly evaluate the performance of large models in the government affairs industry, and adjust and optimize the models based on the evaluation results.

[0007] In an optional embodiment, the step of constructing a government affairs industry model with basic language comprehension capabilities and government affairs knowledge, and initializing the government affairs industry model, includes: Use a large language model to build a large model for the government affairs industry with basic language understanding capabilities and government affairs knowledge; Use preset government-related data to pre-train the government industry model to learn general knowledge and background information in the government field; Use specific data from the government affairs field to train the government affairs industry big model, and adjust the model based on the training results so that the government affairs industry big model can generate content that meets the requirements.

[0008] In an optional embodiment, a government environment simulator is provided for the government industry large model to simulate various government scenario tasks, including: Create a government affairs environment simulator, and use it to simulate government affairs scenario tasks for the government affairs industry large model. The government affairs scenario tasks include policy interpretation, official document writing, intelligent question answering, and government affairs document processing; Set goals and evaluation criteria for each government scenario task through the government environment simulator; Setting input information for the government environment simulator, the input information including task type, background context, user input, and government domain knowledge; Output information is set for the government environment simulator, where the output information includes task results, reward signals, and feedback information.

[0009] In an optional embodiment, the state information, action information, and reward function are defined for the large model of the government affairs industry to guide the model to generate output that meets the requirements, including: Define input state information for the government industry large model, including context information of the current task, user input information, and model internal state information; Define output action information for the government industry large model, including policy interpretation text, official document writing, and intelligent question and answer responses generated by the model automatically; Define a multi-dimensional reward function for the government industry model, which includes an accuracy reward function R accuracy , relevance reward function R relevance , compliance reward function R compliance , format reward function R format ; Based on the multi-dimensional reward function, the comprehensive reward R is calculated by the following formula total :

[0010] in, 、 、 、 are the weights of each multi-dimensional reward function.

[0011] In an optional embodiment, the use of a reinforcement learning algorithm to learn the strategy of the large model of the government affairs industry to learn the optimal strategy by interacting with the government affairs environment includes the following steps: Initialize the policy network π θ (a t ∣s t ) and value network V(s t ), and set the parameters of the initialization policy network θ and the parameters of the value network ; Among them, a t is the action information at time t, s t is the status information at time t; Use the current policy network π θ (a t ∣s t ) State information s at time t in the government environment t , action information a t 、Comprehensive Reward R t and the state information s at the next moment t+1 ; The advantage function A at time t is calculated using the following formula t :

[0012] Where γ is the discount factor; Use the PPO algorithm to update the parameters θ of the policy network and minimize the following first loss function:

[0013] Among them, ϵ is a hyperparameter used to control the probability ratio range of truncation; π θold (a t ∣s t ) is the old policy in state st Next select action a t probability; Update the parameters of the value network using the following formula , and minimize the following second loss function:

[0014] Repeat the above process until the strategy converges or the preset number of training rounds is reached.

[0015] In an optional embodiment, collecting user feedback in government scenarios as an additional reward signal to optimize the reinforcement learning training of the model includes: Collect real user feedback in government scenarios through the government environment simulator, including user satisfaction and task completion rate; Incorporate real user feedback as an additional reward signal into the reinforcement learning training process to optimize the strategy of large models in the government industry.

[0016] In an optional embodiment, the periodic evaluation of the performance of the government affairs industry large model and the adjustment and optimization of the model based on the evaluation results include: Regularly evaluate the task completion rate, user satisfaction, content accuracy and reliability of the government industry big model; Adjust model parameters, multi-dimensional reward function weights, or training data based on the evaluation results to continuously optimize model performance.

[0017] In a second aspect, the present application also provides a reinforcement learning training system for a large model of the government affairs industry, including: The model initialization module is used to build a large government industry model with basic language understanding capabilities and government affairs knowledge, and initialize the large government industry model; The government environment simulator setting module is used to set up a government environment simulator for the government industry large model to simulate various government scenario tasks; The parameter information definition module is used to define state information, action information, and reward functions for the large model of the government industry to guide the model to generate output that meets the requirements; The strategy learning module is used to learn the strategy of the large-scale model of the government industry using reinforcement learning algorithms, so as to learn the optimal strategy through interaction with the government environment; Feedback and optimization module, used to collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model; The evaluation module is used to regularly evaluate the performance of large models in the government affairs industry and adjust and optimize the models based on the evaluation results.

[0018] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the reinforcement learning training method for the large model of the government affairs industry as described in any one of the above items are implemented.

[0019] In a fourth aspect, an embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the reinforcement learning training method for the large model of the government affairs industry as described in any of the above items are implemented.

[0020] It can be seen from the above technical solutions that the present invention has the following advantages: In the reinforcement learning training method for the large model of the government affairs industry provided in this application, a large model of the government affairs industry with basic language comprehension capabilities and government affairs knowledge is constructed, a government affairs environment simulator is set up to simulate various government affairs scenario tasks, state information, action information and multi-dimensional reward functions are defined to guide the model output, and the reinforcement learning algorithm is used to interact with the government affairs environment to learn the optimal strategy. At the same time, user feedback in the government affairs scenario is used as an additional reward signal to optimize model training, and the model performance is regularly evaluated to adjust and optimize the model. This method significantly improves the performance of the large model of the government affairs industry in various government affairs scenario tasks, enabling it to complete tasks such as policy interpretation, official document writing, intelligent question and answer, and government documents processing more efficiently and accurately.

[0021] This application uses reinforcement learning to optimize the government industry's large-scale model, enabling faster and more accurate processing of government tasks, reducing manual intervention and improving the efficiency of government work. The efficiency of processing government tasks is one of the key indicators of government work. The model optimized through reinforcement learning can respond to user needs more quickly and complete tasks more accurately, thereby reducing manual intervention and improving the overall efficiency of government work.

[0022] This application can make the government content generated by the model more accurate and compliant, reduce errors and misleading information, and improve the quality and credibility of government information. The quality and credibility of government information are important guarantees for government work. Through reinforcement learning training methods, the government content generated by the model can be made more accurate and compliant, reduce errors and misleading information, and thus improve the quality and credibility of government information.

[0023] This application uses feedback from government scenarios for training, enabling the model to better adapt to various tasks and needs in the government sector, with greater adaptability and flexibility. The tasks and needs in the government sector are diverse and complex. By training with feedback from government scenarios, the model can better adapt to different tasks and needs, with greater adaptability and flexibility.

[0024] This application can reduce the workload of manual review and modification, lowering the labor cost of government work while improving the overall quality and efficiency of work. Labor cost is one of the important costs of government work. By reducing the workload of manual review and modification, it can reduce the labor cost of government work while improving the overall quality and efficiency of work. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 Flowchart of the reinforcement learning training method for the large-scale model of the government industry provided in this application.

[0027] Figure 2 Schematic diagram of the process of strategy learning provided in this application.

[0028] Figure 3 Schematic diagram of the structure of the reinforcement learning training system for the large-scale model of the government industry provided in this application.

[0029] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION

[0030] The various embodiments of the present disclosure will be described in more detail below in the specific steps of the reinforcement learning training method for the government industry large model. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather that the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.

[0031] The core of this invention is to provide a reinforcement learning training method for a large model in the government affairs industry. In the related art, the following problems exist: Insufficient understanding of professional knowledge in the government field: Existing large models are insufficient in understanding professional knowledge in the government field and are unable to accurately understand and process professional terminology and policies and regulations in government tasks.

[0032] Issues with the accuracy and reliability of generated content: Existing large models may have deviations or errors when generating government content, affecting the quality and credibility of government information.

[0033] Insufficient model adaptability: Existing large models lack effective mechanisms to self-optimize and adjust based on task feedback when processing government tasks, making it difficult to adapt to the diversity and complexity of government tasks.

[0034] To solve the above problems, the present invention proposes a reinforcement learning training method for a large-scale model in the government affairs industry, which achieves the purpose of the invention through the following methods: 1. Integrate knowledge in the government field: During the reinforcement learning process, professional knowledge and policies and regulations in the government field are integrated into the model training, enabling the model to better understand and handle government tasks.

[0035] 2. Design a multi-dimensional reward function: Based on the characteristics of government scenarios, a reward function that comprehensively considers multiple dimensions such as accuracy, relevance, and compliance is designed to guide the model to generate high-quality government content.

[0036] 3. Introduce a user feedback mechanism: By collecting real user feedback in government scenarios, user satisfaction is taken as an important optimization goal, so that the model can better meet user needs.

[0037] 4. Optimize model strategies: Use reinforcement learning algorithms to learn and optimize the strategies of large models, so that the model can adjust its behavior based on task feedback, improving the adaptability and flexibility of the model.

[0038] Through the above-mentioned technical means, the present invention aims to improve the performance and application effect of the big model of the government affairs industry in government affairs scenarios, so that it can handle government affairs tasks more efficiently and accurately, and meet the strict requirements of government affairs work for efficiency, accuracy and compliance.

[0039] Hereinafter, the terms "include" or "may include" as used in various embodiments of the present disclosure indicate the presence of disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "include," "have," and their cognates are intended only to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing, and should not be understood as excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or the possibility of adding one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing.

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0041] See also Figure 1 The figure is a flowchart of a method for reinforcement learning training of a large model of the government affairs industry in a specific embodiment, the method comprising: S1: Build a government affairs industry big model with basic language comprehension capabilities and government affairs knowledge, and initialize the government affairs industry big model.

[0042] In the specific implementation, a large language model (such as the DeepSeek R1 model) is first used to build a large government sector model with basic language understanding capabilities and government knowledge. This model has been preliminarily trained in general and government domains and has a certain knowledge base.

[0043] Next, the government sector model is pre-trained using pre-set government-related data to learn general knowledge and background information in the government sector. The selection and initialization of the pre-trained model is fundamental to the entire training process and directly impacts model performance and training results. The pre-trained model should possess strong language understanding and generation capabilities and have a sufficient amount of accumulated knowledge in the government sector to better adapt to the needs of government tasks during the subsequent reinforcement learning process.

[0044] Specifically, during the pre-training phase, the model is trained on a large amount of government-related data to learn general knowledge and background information in the government sector. This data includes policy documents, official documents, laws and regulations, and service guidelines. Through pre-training, the model is able to understand the language patterns and knowledge systems of the government sector.

[0045] Furthermore, the government sector's large model is trained using data specific to the government sector. Based on the training results, the model is adjusted to ensure that it generates content that meets the requirements. Specifically, during the fine-tuning phase, the model, building on pre-training, is further trained using data specific to the government sector. This data can include annotated policy interpretations, official document writing samples, and intelligent question-and-answer records. This fine-tuning process enables the model to more accurately handle government tasks and generate content that meets policy requirements.

[0046] S2: Set up a government environment simulator for the government industry large model to simulate various government scenario tasks.

[0047] In a specific implementation, a government environment simulator is first created. This simulator then simulates government scenarios for a large government industry model. These scenarios include policy interpretation, document writing, intelligent question-and-answer (Q&A), and government document processing. Furthermore, the simulator sets goals and evaluation criteria for each task.

[0048] Specifically, a variety of government scenarios are set up through the government environment simulator, including policy interpretation, official document writing, and intelligent question-and-answering. Each task has clear objectives and evaluation criteria. The design of the government environment simulator is crucial. It needs to be able to accurately simulate the various tasks and environments in government scenarios and provide a rich training sample for the model. The design of government scenario tasks should cover all aspects of government work, including but not limited to policy interpretation, official document writing, intelligent question-and-answering, and government document processing. Each task should have clear objectives and evaluation criteria to ensure accurate evaluation and feedback on the model's performance during the reinforcement learning process.

[0049] It's important to note that the role of the government environment simulator in this method is to simulate various tasks and environments in government scenarios, providing a rich set of training samples for the model. It accurately simulates the context, processes, and user interactions of government tasks, helping the model better understand and handle them. The simulator allows the model to undergo extensive training in a virtual environment, thereby improving its performance in real-world government scenarios.

[0050] Then, input information is set for the government environment simulator, where the input information includes task type, background context, user input, and government domain knowledge.

[0051] For example, the input information of the government environment simulator mainly includes the following categories: Task types: such as policy interpretation, official document writing, intelligent question and answer, government document processing, etc.

[0052] Task background and contextual information: including policy documents, official document templates, historical cases, etc.

[0053] User input: A query, request, or question from a user.

[0054] Government domain knowledge and policies and regulations: used to ensure that the content generated by the model complies with policy requirements.

[0055] Finally, output information is set for the government environment simulator, and the output information includes task results, reward signals and feedback information.

[0056] For example, the output information of the government environment simulator mainly includes the following categories: Task results: The model generates policy interpretations, official document drafts, question-and-answer responses, etc. based on the input.

[0057] Reward signal: A reward value generated based on the quality and compliance of the model’s output.

[0058] Feedback information: Evaluation results of model output, including accuracy, relevance, compliance, etc.

[0059] The purpose of the output information is to provide feedback to the model, helping it optimize its strategy. Reward signals are used to update the strategy during reinforcement learning, while feedback information is used to evaluate the model's performance and guide its improvement.

[0060] In a specific implementation, the infrastructure of a government environment simulator typically includes a task generator, a user simulator, an evaluation module, and a feedback mechanism. The task generator generates specific government tasks based on predefined task types and background information; the user simulator simulates user behavior and input, providing an interactive training environment for the model; the evaluation module evaluates the model's output according to predefined evaluation criteria and generates a reward signal; and the feedback mechanism feeds the evaluation results and reward signal back to the model for policy updates. Its working principle is as follows: First, the task generator generates specific government tasks based on government domain knowledge and policies and regulations; then, the user simulator simulates user behavior and asks questions or requests to the model; then, the model generates corresponding actions based on the input state information, such as generating text or recommending operations; then, the evaluation module evaluates the model's output, generating a reward signal and feedback information; finally, the model updates its policy based on the reward signal and feedback information to optimize future behavior.

[0061] S3: Define state information, action information, and reward functions for large models in the government sector to guide the model to generate outputs that meet the requirements.

[0062] In this specific implementation, we first define the input state information for the large government industry model. This state information includes the context of the current task, user input, and the model's internal state. The state definition should be comprehensive and accurate, reflecting all aspects of the current task so that the model can make reasonable decisions based on this state information. The acquisition and processing of state information is a key step in the reinforcement learning process, and its integrity and accuracy must be ensured.

[0063] For example, the specific content of the input status information is as follows: 1. Contextual information of the current task: Task type: Clarify the category of the current task, such as policy interpretation, official document writing, intelligent question and answer, etc.

[0064] Task context: This includes the source, purpose, and related background information of the task, helping the model understand the overall framework of the task.

[0065] Historical interaction records: Record previous interactions with users, including user questions, model responses, and user feedback, so that the model can make more reasonable decisions based on historical information.

[0066] 2. User input: Text input: The text questions or requirements raised by the user, which is the main basis for the model to process and generate answers.

[0067] Voice input: If voice interaction is supported, the user's voice input will also be converted into text or other processable formats.

[0068] Other forms of input: may include files and pictures uploaded by users. These inputs need to be preprocessed before they can be input into the model as state information.

[0069] 3. Internal state of the model: Model parameters: The parameter status of the current model, including weights and biases, which determine the behavior of the model.

[0070] Cache information: The model may store intermediate results or cached data during processing. This information can help the model respond quickly.

[0071] Status of processed tasks: records the status of tasks that the model has processed, including task completion status, user satisfaction, etc.

[0072] Next, define output action information for the government industry big model. This action information represents the big model's response to user input, including generated text and recommended actions. The definition of action information should align with the task objectives and meet the specific needs of the government task. The results of the action execution directly impact task completion and user satisfaction, so the definition of action information must fully consider the characteristics and requirements of the government task.

[0073] For example, the specific content of the output action information is as follows: Policy interpretation: Generate detailed interpretations of policy documents, including policy background, specific content, implementation details, etc., to help users better understand the policies.

[0074] Official document writing: Generate official documents that comply with government regulations, such as notices, reports, requests, etc., ensuring correct format and accurate content.

[0075] Furthermore, a multi-dimensional reward function is defined for large models in the government sector. Specifically, the reward function is designed based on the characteristics of the government scenario, taking into account multiple dimensions such as accuracy, relevance, and compliance. For example, in a policy interpretation task, if the content generated by the model is accurate and meets policy requirements, a positive reward is given; if the generated content contains errors or does not meet policy requirements, a negative reward is given. The design of the reward function is a core part of the reinforcement learning process, as it determines the model's behavioral direction and optimization goals. The reward function should comprehensively consider all aspects of the government task, including accuracy, relevance, and compliance, to guide the model to generate high-quality government content.

[0076] In reinforcement learning training for large models in the government sector, the design of the reward function is a core step, determining the model's behavior and optimization goals. The following is a detailed reward function design, which comprehensively considers multiple dimensions such as accuracy, relevance, compliance, and format.

[0077] For example, the multi-dimensional reward function specifically includes: Accuracy reward function R accuracy : Accuracy is a core requirement for government tasks, especially in policy interpretation and document writing. The accuracy reward function is as follows:

[0078] Relevance reward function R relevance : Relevance reward is used to evaluate how well the model output matches the user input. The functional form of relevance reward is as follows:

[0079] Compliance reward function R compliance : Compliance rewards are used to assess whether the model output complies with government regulations and policy requirements. The functional form of compliance rewards is as follows:

[0080] Format reward function R format : The format reward is used to evaluate whether the model output conforms to the standard format of government documents. The format reward function is as follows:

[0081] Finally, based on the multi-dimensional reward function, the comprehensive reward R is calculated by the following formula total :

[0082] in, 、 、 、 are the weights of each multi-dimensional reward function.

[0083] S4: Use reinforcement learning algorithms to learn the strategies of large models in the government sector, so as to learn the optimal strategies through interaction with the government environment.

[0084] In this step, a reinforcement learning algorithm (such as Q-learning and PPO) is used to learn the policy of the large model. At each time step, the model selects an action based on its current state and updates its policy based on the rewards fed back by the environment. Policy learning is a key step in the reinforcement learning process. Through policy learning, the model can adjust its behavior based on the reward signals fed back by the environment to achieve better task performance. The selection and implementation of the policy learning algorithm must be optimized and adjusted according to the characteristics of the specific task and model to ensure that the model can effectively learn the optimal policy.

[0085] S5: Collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model.

[0086] In a specific implementation, real user feedback in government scenarios, including user satisfaction and task completion rate, is first collected through a government environment simulator; then the real user feedback is incorporated into the reinforcement learning training process as an additional reward signal to optimize the strategy of the large model of the government industry.

[0087] During training, a feedback mechanism collects real user feedback and task results from government scenarios. This feedback serves as an additional reward signal to further optimize model training. User feedback is an important basis for evaluating model performance. By collecting and analyzing user feedback, we can understand the model's performance in real-world applications, identify any problems or deficiencies, and thus provide a basis for model optimization and improvement. The feedback mechanism should be designed to accurately collect and process user feedback information and effectively integrate it into the model training process.

[0088] S6: Regularly evaluate the performance of large models in the government affairs industry and adjust and optimize the models based on the evaluation results.

[0089] In a specific implementation method, the task completion rate, user satisfaction, content accuracy and reliability of the large model of the government affairs industry are regularly evaluated; the model parameters, multi-dimensional reward function weights or training data are adjusted according to the evaluation results to continuously optimize the model performance.

[0090] For example, on the one hand, trained models are regularly evaluated, using metrics such as task completion rate, user satisfaction, and the accuracy and reliability of generated content. Based on the evaluation results, the model is adjusted and optimized to improve its performance in government scenarios. Model evaluation is a critical step in the training process. It allows us to understand the model's performance and identify any problems or deficiencies, thus providing a basis for model optimization and improvement. The evaluation metrics selected should be able to fully reflect the model's performance in government scenarios, including task completion rate, user satisfaction, and the accuracy and reliability of generated content. Based on the evaluation results, the model is adjusted and optimized to improve its performance and performance.

[0091] On the other hand, by collecting real user feedback and taking user satisfaction as a key optimization objective, the model can better meet user needs. User feedback is a crucial basis for evaluating model performance. By collecting and analyzing user feedback, we can understand the model's performance in real applications, identify any problems or deficiencies, and thus provide a basis for model optimization and improvement. User feedback-driven optimization can help the model better meet user needs, improving its practicality and user satisfaction.

[0092] In this embodiment, by constructing a large industry model that integrates government knowledge and language comprehension capabilities, combined with the realistic reproduction of multi-scenario tasks such as policy interpretation and official document writing by a government environment simulator, and the precise guidance of output quality by a multi-dimensional reward function, a closed-loop training system of "knowledge pre-training-scenario simulation-strategy reinforcement-feedback optimization" is formed. Based on the PPO reinforcement learning algorithm, the policy iteration and value network collaborative optimization enable the model to achieve dynamic balance in key indicators such as accuracy, compliance, and format specifications. At the same time, the continuous injection of real feedback signals such as user satisfaction further enhances the model's adaptability to complex government needs. The parameter tuning and weight adjustment mechanism driven by regular evaluation ensures the continuous optimization of the model in core dimensions such as task completion rate and content reliability, ultimately achieving a significant improvement in the efficiency and quality of government scenario processing, effectively promoting the intelligent transformation of government services.

[0093] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0094] like Figure 2 As shown, this embodiment discloses a strategy learning method, which specifically includes the following steps: S401: Initialize the policy network and value network.

[0095] Initialize the policy network π θ (a t ∣st ) and value network V(s t ), and set the parameters of the initialization policy network θ and the parameters of the value network ; Among them, a t is the action information at time t, s t is the status information at time t; S402: Collect status information, action information and comprehensive rewards.

[0096] Use the current policy network π θ (a t ∣s t ) State information s at time t in the government environment t , action information a t 、Comprehensive Reward R t and the state information s at the next moment t+1 ; S403: Calculate the advantage function.

[0097] The advantage function A at time t is calculated using the following formula t :

[0098] Where γ is the discount factor; S404: Update the policy network.

[0099] Use the PPO algorithm to update the parameters θ of the policy network and minimize the following first loss function:

[0100] Among them, ϵ is a hyperparameter used to control the probability ratio range of truncation; π θold (a t ∣s t ) is the old policy in state s t Next select action a t The probability of π θ (a t ∣s t ) is the current strategy in state s t Next select action a t probability.

[0101] It is important to note that the PPO algorithm is a policy gradient-based reinforcement learning algorithm that maximizes cumulative rewards by optimizing the policy network. The core idea of the PPO algorithm is to limit the amplitude of policy updates by truncating the probability ratio, thereby improving the stability and efficiency of training.

[0102] S405: Update the value network.

[0103] Update the parameters of the value network using the following formula , and minimize the following second loss function:

[0104] S406: Repeat the above steps until the strategy converges or the preset number of training rounds is reached.

[0105] like Figure 3 As shown, the following is an embodiment of the reinforcement learning training system for the government affairs industry big model provided by the embodiment of the present disclosure. This system and the reinforcement learning training method for the government affairs industry big model in the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiment of the reinforcement learning training system for the government affairs industry big model, please refer to the embodiment of the reinforcement learning training method for the government affairs industry big model.

[0106] A reinforcement learning training system for a large model of the government affairs industry includes: a model initialization module, a government affairs environment simulator setting module, a parameter information definition module, a strategy learning module, a feedback and optimization module, and an evaluation module.

[0107] The model initialization module is used to build a large model of the government affairs industry with basic language comprehension capabilities and government affairs knowledge, and to initialize the large model of the government affairs industry.

[0108] The government affairs environment simulator setting module is used to set up a government affairs environment simulator for the government affairs industry large model to simulate various government affairs scenario tasks.

[0109] The parameter information definition module is used to define state information, action information and reward functions for the large model of the government industry to guide the model to generate output that meets the requirements.

[0110] The strategy learning module is used to learn the strategy of the large model of the government industry using the reinforcement learning algorithm, so as to learn the optimal strategy through interaction with the government environment.

[0111] The feedback and optimization module is used to collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model.

[0112] The evaluation module is used to regularly evaluate the performance of large models in the government affairs industry and adjust and optimize the models based on the evaluation results.

[0113] The reinforcement learning training system for the government affairs industry big model provided in this embodiment constructs a government affairs industry big model with basic language comprehension capabilities and government affairs knowledge, and uses a government affairs environment simulator to simulate various government affairs scenario tasks, defines state information, action information and multi-dimensional reward functions, and uses a reinforcement learning algorithm for strategy learning. At the same time, it combines user feedback in government affairs scenarios as an additional reward signal for optimization training, and regularly evaluates model performance for adjustment and optimization, thereby effectively improving the task processing capabilities, user satisfaction, and content accuracy and reliability of the government affairs industry big model in government affairs scenarios, and realizing efficient learning and performance improvement of the government affairs industry big model.

[0114] Figure 4 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.

[0115] The reinforcement learning training method for the large model of the government industry provided in the embodiment of the present application can be applied to electronic devices. Those skilled in the art will understand that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently. In the embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0116] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a button, a camera, a display, and a SIM card interface, etc.

[0117] A processor may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0118] The processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on the instruction opcode and timing signal to complete the control of instruction fetching and execution.

[0119] The processor may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.

[0120] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of an electronic device. The external memory card communicates with the processor through the external memory interface, enabling data storage. For example, files such as music and videos can be stored on the external memory card.

[0121] Internal memory can be used to store computer-executable program code, which includes instructions. The processor executes the instructions stored in the internal memory to perform various functional applications and data processing of the electronic device. The internal memory can include a program storage area and a data storage area. The internal memory can include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0122] The wireless communication function of an electronic device can be implemented through an antenna, a wireless communication module, a modem processor, and a baseband processor.

[0123] Wireless communication modules can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.

[0124] Electronic devices can implement audio functions through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.

[0125] Electronic devices can achieve shooting functions through ISP, camera, video codec, GPU, display and application processor.

[0126] Electronic devices can achieve display functions through GPU, display screen and application processor.

[0127] A GPU is a microprocessor for image processing that connects the display screen to the application processor. The GPU performs mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information.

[0128] The display screen is used to display images, videos, etc. The display screen includes a display panel.

[0129] The above-mentioned electronic device implements the reinforcement learning training method of the government affairs industry big model of this application by constructing a government affairs industry big model and adopting the reinforcement learning training method, combined with the government affairs environment simulator, multi-dimensional reward function and user feedback optimization, to achieve the beneficial effects of improving the task processing capabilities of government affairs scenarios, enhancing content accuracy and compliance, improving user satisfaction and realizing efficient learning and continuous optimization of the model.

[0130] The storage medium provided in this application stores a program product that can implement a reinforcement learning training method for a large model of the government affairs industry.

[0131] Reinforcement learning training methods for large models in the government sector include: Build a large model of the government affairs industry with basic language understanding capabilities and government affairs knowledge, and initialize the large model of the government affairs industry; Set up a government environment simulator for the government industry large model to simulate various government scenario tasks; Define state information, action information, and reward functions for large models in the government sector to guide the model to generate outputs that meet requirements; Use reinforcement learning algorithms to learn the strategies of large models in the government sector, so as to learn the optimal strategies through interaction with the government environment; Collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model; Regularly evaluate the performance of large models in the government affairs industry, and adjust and optimize the models based on the evaluation results. In some possible implementations, the reinforcement learning training method for the government affairs industry large model disclosed herein can be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above "Exemplary Method" section of this specification according to various exemplary implementations of the present disclosure.

[0132] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0133] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A reinforcement learning training method for a large model in the government industry, characterized by: include: Build a large model of the government affairs industry with basic language comprehension capabilities and government affairs knowledge, and initialize the large model of the government affairs industry; Set up a government environment simulator for the government industry large model to simulate various government scenario tasks; Define state information, action information, and reward functions for large models in the government sector to guide the model to generate outputs that meet requirements; Use reinforcement learning algorithms to learn the strategies of large models in the government sector, so as to learn the optimal strategies through interaction with the government environment; Collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model; Regularly evaluate the performance of large models in the government affairs industry and adjust and optimize the models based on the evaluation results.

2. The reinforcement learning training method for the government affairs industry large model according to claim 1 is characterized in that: The aforementioned construction of a government affairs industry big model with basic language comprehension capabilities and government affairs knowledge, and initialization of the government affairs industry big model, includes: Use a large language model to build a large model for the government affairs industry with basic language understanding capabilities and government affairs knowledge; Use preset government-related data to pre-train the government industry model to learn general knowledge and background information in the government field; Use specific data from the government affairs field to train the government affairs industry big model, and adjust the model based on the training results so that the government affairs industry big model can generate content that meets the requirements.

3. The reinforcement learning training method for the government affairs industry large model according to claim 1 is characterized in that: The government affairs environment simulator is set up for the government affairs industry large model to simulate various government affairs scenario tasks, including: Create a government affairs environment simulator, and use it to simulate government affairs scenario tasks for the government affairs industry large model. The government affairs scenario tasks include policy interpretation, official document writing, intelligent question answering, and government affairs document processing; Set goals and evaluation criteria for each government scenario task through the government environment simulator; Setting input information for the government environment simulator, the input information including task type, background context, user input, and government domain knowledge; Output information is set for the government environment simulator, and the output information includes task results, reward signals and feedback information.

4. The reinforcement learning training method for the government affairs industry large model according to claim 1 is characterized in that: The aforementioned large-scale model for the government sector defines state information, action information, and reward functions to guide the model to generate outputs that meet the requirements, including: Define input state information for the government industry large model, including context information of the current task, user input information, and model internal state information; Define output action information for the government industry large model, including policy interpretation text, official document writing, and intelligent question and answer responses generated by the model automatically; Define a multi-dimensional reward function for the government industry model, which includes an accuracy reward function R accuracy , relevance reward function R relevance , compliance reward function R compliance , format reward function R format ; Based on the multi-dimensional reward function, the comprehensive reward R is calculated by the following formula total : in, 、 、 、 are the weights of each multi-dimensional reward function.

5. The reinforcement learning training method for the government affairs industry large model according to claim 4 is characterized in that: The use of reinforcement learning algorithms to learn the strategies of the large model of the government industry, so as to learn the optimal strategies through interaction with the government environment, includes: Initialize the policy network π θ (a t ∣s t ) and value network V(s t ), and set the parameters of the initialization policy network θ and the parameters of the value network ; Among them, a t is the action information at time t, s t is the status information at time t; Use the current policy network π θ (a t ∣s t ) State information s at time t in the government environment t , action information a t 、Comprehensive Reward R t and the state information s at the next moment t+1 ; The advantage function A at time t is calculated using the following formula t : Where γ is the discount factor; Use the PPO algorithm to update the parameters θ of the policy network and minimize the following first loss function: Among them, ϵ is a hyperparameter used to control the probability ratio range of truncation; π θold (a t ∣s t ) is the old policy in state s t Next select action a t probability; Update the parameters of the value network using the following formula , and minimize the following second loss function: Repeat the above process until the strategy converges or the preset number of training rounds is reached.

6. The reinforcement learning training method for the government affairs industry large model according to claim 1 is characterized in that: The collection of user feedback in government scenarios as an additional reward signal to optimize the reinforcement learning training of the model includes: Collect real user feedback in government scenarios through the government environment simulator, including user satisfaction and task completion rate; Incorporate real user feedback as an additional reward signal into the reinforcement learning training process to optimize the strategy of large models in the government industry.

7. The reinforcement learning training method for the government affairs industry large model according to claim 4 is characterized in that: The aforementioned regular evaluation of the performance of the government affairs industry large model and adjustment and optimization of the model based on the evaluation results include: Regularly evaluate the task completion rate, user satisfaction, content accuracy and reliability of the government industry big model; Adjust model parameters, each multi-dimensional reward function, or training data based on the evaluation results to continuously optimize model performance.

8. A reinforcement learning training system for a large model of the government industry, characterized by: The system adopts the reinforcement learning training method of the large government industry model as described in any one of claims 1 to 7; The system comprises: The model initialization module is used to build a large government industry model with basic language understanding capabilities and government affairs knowledge, and initialize the large government industry model; The government environment simulator setting module is used to set up a government environment simulator for the government industry large model to simulate various government scenario tasks; The parameter information definition module is used to define state information, action information, and reward functions for the large model of the government industry to guide the model to generate output that meets the requirements; The strategy learning module is used to learn the strategy of the large model of the government industry using reinforcement learning algorithms, so as to learn the optimal strategy through interaction with the government environment; The feedback and optimization module is used to collect user feedback in government scenarios and use it as an additional reward signal to optimize the reinforcement learning training of the model; The evaluation module is used to regularly evaluate the performance of large models in the government affairs industry and adjust and optimize the models based on the evaluation results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the steps of the reinforcement learning training method for the government affairs industry large model as described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the reinforcement learning training method for the government affairs industry large model as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Official document writing optimization method and device, electronic equipment and storage medium

    CN121859849A