Information generation method and device, equipment and storage medium

By determining the task type and evaluation dimension, and combining the asynchronous mechanism with the evaluation dimension priority ladder reward fusion strategy, the problem of inaccurate reward signals in large-scale reinforcement learning models is solved, the accuracy and efficiency of information generation are improved, and the accuracy and pertinence of reward information are ensured.

CN120671734APending Publication Date: 2025-09-19BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510812187.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing large-scale reinforcement learning models have deficiencies in the accuracy, stability, and real-time performance of reward signals, resulting in inaccurate model optimization direction and prone to problems such as over-optimization of rewards and poor robustness.

Method used

By determining the task type and task evaluation dimension of the target task, combining multiple evaluation dimensions and evaluation results to generate target information, and adopting an asynchronous mechanism and evaluation dimension priority ladder reward fusion strategy, the accuracy and efficiency of information generation are improved, and GPU resource utilization is maximized.

Benefits of technology

It alleviates the transition optimization problem of fixed evaluation dimensions, improves the accuracy and efficiency of information generation, ensures the accuracy and pertinence of reward information, and improves user experience and model training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671734A_ABST
    Figure CN120671734A_ABST
Patent Text Reader

Abstract

The invention provides an information generation method, and relates to the technical field of artificial intelligence, in particular to the technical fields of natural language processing, deep learning, large language models and the like. The method comprises the steps of determining a task type of a target task; determining a task evaluation dimension corresponding to the target task according to the task cue word and the task type of the target task; according to the task evaluation dimension and a task result, an evaluation result corresponding to the task evaluation dimension is generated, and the task result is generated by the large language model according to the target task and the task cue word; and determining target information of the target task according to the task evaluation dimension and the evaluation result. The method improves the comprehensiveness and accuracy of the determined information result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as natural language processing, deep learning, and large language models, and especially to an information generation method, apparatus, device, and storage medium. Background Art

[0002] In recent years, large-scale reinforcement learning (RL) has demonstrated groundbreaking potential in fields like natural language processing. By providing models with accurate information (such as reward signals), RL can effectively guide them to produce responses that cater to human preferences, significantly improving their performance on a variety of tasks. The accuracy, stability, and real-time nature of reward information are crucial to effective training. Only by providing a stable, accurate, and reasonable reward signal can the model optimize in the desired direction. Summary of the Invention

[0003] The present disclosure provides an information generation method, apparatus, device, and storage medium.

[0004] According to a first aspect of the present disclosure, an information generation method is provided, including: determining a task type of a target task; determining a task evaluation dimension corresponding to the target task based on a task prompt word and a task type of the target task; generating an evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result, wherein the task result is generated by a large language model based on the target task and the task prompt word; and determining target information of the target task based on the task evaluation dimension and the evaluation result.

[0005] According to a second aspect of the present disclosure, an information generating device is provided, comprising: a task type determining module, configured to determine the task type of a target task; a task dimension determining module, configured to determine a task evaluation dimension corresponding to the target task based on a task prompt word and a task type of the target task; an evaluation result determining module, configured to generate an evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result, wherein the task result is generated by a large language model based on the target task and the task prompt word; and an information determining module, configured to determine target information of the target task based on the task evaluation dimension and the evaluation result.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner in the first aspect.

[0007] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method described in any implementation manner of the first aspect.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method described in any implementation manner of the first aspect when executed by a processor.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention. Figure 1 is an exemplary system architecture diagram in which the present disclosure may be applied; Figure 2 is a flow chart of a first embodiment of an information generating method according to the present disclosure; Figure 3 is a flow chart of a second embodiment of the information generating method according to the present disclosure; Figure 4 is a flow chart of a third embodiment of the information generating method according to the present disclosure; Figure 5 is a flowchart of a fourth embodiment of the information generating method according to the present disclosure; Figure 6 is a flowchart of a fifth embodiment of the information generating method according to the present disclosure; Figure 7 is a structural diagram of an embodiment of an information generating device according to the present disclosure; Figure 8 It is a block diagram of an electronic device used to implement the information generating method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0011] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0012] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0013] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the information generating method or information generating apparatus disclosed herein can be applied.

[0014] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0015] Users can use terminal devices 101, 102, 103 to interact with server 105 via network 104 to receive or send information, etc. Various client applications can be installed on terminal devices 101, 102, 103.

[0016] Terminal devices 101, 102, and 103 can be either hardware or software. When hardware is used, terminal devices 101, 102, and 103 can be various electronic devices, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When software is used, terminal devices 101, 102, and 103 can be installed in any of the aforementioned electronic devices. These devices can be implemented as multiple software programs or software modules, or as a single software program or software module. This is not specifically limited here.

[0017] The server 105 can provide various services. For example, the server 105 can analyze and process the target tasks and task prompt words obtained from the terminal devices 101, 102, and 103, and generate processing results (such as target information).

[0018] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (for example, to provide distributed services), or as a single software program or software module. This is not specifically limited here.

[0019] It should be noted that the information generating method provided in the embodiment of the present disclosure is generally executed by the server 105 , and accordingly, the information generating device is generally provided in the server 105 .

[0020] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0021] Continue to refer Figure 2 , which shows a process 200 of a first embodiment of the information generation method according to the present disclosure. The information generation method comprises the following steps: Step 201: Determine the task type of the target task.

[0022] In this embodiment, the execution subject of the information generation method (for example Figure 1 The server 105 shown in the figure will first determine the task type of the target task. Specifically, the execution entity will first obtain the target task and related information about the target task, such as the task content and task prompts. After obtaining the relevant information about the target task, the execution entity will parse the relevant information of the target task, for example, parsing the task prompts corresponding to the target task, and then determine the task type of the target task based on the parsed results. The task type here can be a translation task, a creative task, etc. Specifically, a translation task refers to the task of translating content from one language to another. A creative task generally refers to a literary creation task, that is, the process of creating literary works for readers to enjoy through artistic processing.

[0023] Step 202: Determine the task evaluation dimension corresponding to the target task according to the task prompt word and task type of the target task.

[0024] In this embodiment, the execution entity determines the task evaluation dimensions corresponding to the target task based on the task prompt and task type. The task prompt here refers to the prompt word corresponding to the target task, which refers to the key instructions in the input sample during large-scale model fine-tuning and task processing. For example, if the target task type is a creative task, the task prompt word for the target task might be "Please create a lyrical text based on the following content."

[0025] Furthermore, after determining the target task's task type, the execution entity will also determine the target strategy corresponding to that task type from a pre-built task strategy library. Specifically, the task strategy library in this embodiment contains task strategies corresponding to multiple tasks, with different tasks corresponding to different task strategies. Therefore, after determining the target task's task type, the execution entity can determine the target strategy corresponding to that task type from the task strategy library based on the correspondence between task types and task strategies.

[0026] Then, the task prompt word is dimensionally evaluated according to the target strategy to obtain the task evaluation dimension corresponding to the target task. Here, there can be at least one task evaluation dimension. If d is used to represent the task evaluation dimension, then the task evaluation dimension corresponding to the target task can be expressed as 、 , where n≥1.

[0027] As an example, when the task type of the target task is a creative task, the task evaluation dimensions of the target task determined may be: high quality, word count compliance, and instruction compliance.

[0028] Step 203: Generate an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result.

[0029] In this embodiment, the execution entity generates an evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result, wherein the task result is generated by the large language model based on the target task and the task prompt word.

[0030] It should be noted that in the field of artificial intelligence, large language models (abbreviated as large models) refer to deep neural networks with over a billion parameters. They are capable of processing massive amounts of data and performing various complex tasks, such as natural language processing, computer vision, and speech recognition. Generative large models are generative models based on large corpora. They are large-scale neural network models that can generate, understand, and reason about natural language in an end-to-end manner. By training on large amounts of text data, they can perform a wide range of tasks, including text summarization, translation, and sentiment analysis. Large models refer to deep neural networks with over a billion parameters, capable of processing massive amounts of data and performing various complex tasks. With the continuous improvement of computer hardware performance and the continuous optimization of deep learning algorithms, the development of large models is accelerating. While the parameter size of large models continues to expand, training time is also increasing, but performance is also improving. Large models are often based on deep learning architectures, such as Transformers, which enable them to demonstrate impressive capabilities in various natural language processing tasks. Common large models include, but are not limited to, ChatGPT, GPT-4, and ERNIE.

[0031] Specifically, the above-mentioned execution entity will first input the target task and task prompt words into the big model, and then output the task result corresponding to the target task. For example, when the target task is a creative task and the task prompt words are "Please create a lyrical text based on the following content", the target task and task prompt words are input into the big model, and the output is the generated lyrical text.

[0032] The execution entity will then generate an evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result. This evaluation result is used to assess whether the task result meets the requirements. For example, if the task type is a translation task, the task result is the translation result, and the evaluation result is used to assess whether all words in the task result are accurately translated.

[0033] Step 204: Determine target information of the target task based on the task evaluation dimension and the evaluation result.

[0034] In this embodiment, the execution entity determines the target information for the target task based on the task evaluation dimensions and evaluation results. The target information here can refer to the reward information for the target task. Rewards in reinforcement learning are the core signals that determine how an agent behaves in an environment. Rewards provide timely feedback on the agent's behavior, used to assess the effectiveness of an action in a given state, and thus influence future decisions.

[0035] Specifically, for the evaluation result of a particular task evaluation dimension, the execution entity will determine whether the evaluation criteria for the current dimension are met based on the evaluation result, and then determine the target information based on the judgment result. For example, the evaluation result may be compared with a preset threshold (e.g., 0), and then the comparison result may be used to determine whether the evaluation criteria for the current dimension are met.

[0036] For example, for Task Evaluation Dimension 2, if the evaluation result for this dimension is equal to 0, it indicates that the current dimension's evaluation is not satisfied. In this case, the target information for the previous Task Evaluation Dimension will be directly returned. If the evaluation result for this dimension is greater than 0, it indicates that the current dimension's evaluation result is satisfied or partially satisfied. In this case, the information for the current dimension will be calculated. The result evaluation process for the next Task Evaluation Dimension will then proceed until the final information, i.e., the target information (e.g., reward information), is obtained.

[0037] Currently, the reward signal for large-scale reinforcement learning is typically determined by a reward model. This reward model memorizes human-annotated preference information to provide a holistic, unified score as feedback for the large-scale model's responses. The current mainstream reward calculation method uses a unified reward model to assign a scalar reward to text, which serves as the basis for further model optimization. However, reward models have the following flaws: They are limited by the distribution of training data and subjective biases, making generalization difficult. They can easily lead to reward over-optimization during iterations, resulting in limited robustness. Furthermore, the rewards provided by the reward model are poorly interpretable, and the correlation between score and response quality is inconsistent.

[0038] The information generation method provided by the embodiment of the present disclosure first determines the task type of the target task, then determines the task evaluation dimension corresponding to the target task based on the task prompt word and the task type of the target task, then generates the evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result, and finally determines the target information of the target task based on the task evaluation dimension and the evaluation result. The information generation method in this embodiment first determines the evaluation dimension corresponding to the task prompt word of the target task, combines multiple evaluation dimensions and the evaluation result corresponding to each evaluation dimension to generate the target information (such as reward information) corresponding to the target task, thereby alleviating the transition optimization problem caused by the fixed evaluation dimension and improving the accuracy of the determined target information. In addition, the information generation method in this embodiment improves the information generation efficiency and computing efficiency by adopting an asynchronous mechanism, and the method can maximize the resource utilization of the GPU (Graphics Processing Unit) to avoid waiting on the training end.

[0039] In addition, in the technical solutions involved in this disclosure, the acquisition, storage, use, processing, transportation, provision and disclosure of user personal information involved (such as the target tasks and task prompts involved in this disclosure) are in compliance with relevant laws and regulations and do not violate public order and good morals.

[0040] Continue to refer Figure 3 , Figure 3 A process 300 of a second embodiment of the information generation method according to the present disclosure is shown. The information generation method comprises the following steps: Step 301: Determine the task type of the target task.

[0041] In this embodiment, the execution subject of the information generation method (for example Figure 1 The server 105 shown in the figure will first determine the task type of the target task. Specifically, the execution entity will first obtain the target task and related information about the target task, such as the task content and task prompts. After obtaining the relevant information about the target task, the execution entity will parse the relevant information of the target task, for example, parsing the task prompts corresponding to the target task, and then determine the task type of the target task based on the parsed results. The task type here can be a translation task, a creative task, etc. Specifically, a translation task refers to the task of translating content from one language to another. A creative task generally refers to a literary creation task, that is, the process of creating literary works for readers to enjoy through artistic processing.

[0042] Step 302: Determine the task evaluation strategy corresponding to the task type.

[0043] In this embodiment, after determining the target task's task type, the execution entity also determines the target strategy corresponding to that task type from a pre-built task strategy library. This means that the task strategy library in this embodiment contains task strategies corresponding to multiple tasks, with different tasks corresponding to different task strategies. Therefore, after determining the target task's task type, the execution entity can determine the target strategy corresponding to that task type, i.e., the task evaluation strategy, from the task strategy library based on the correspondence between task types and task strategies.

[0044] Step 303: extract at least one task keyword from the task prompt words.

[0045] In this embodiment, the execution entity analyzes the task prompt word to determine at least one keyword corresponding to the task prompt word, ie, the task keyword.

[0046] Step 304: Match at least one task keyword with an evaluation dimension keyword corresponding to the task evaluation strategy, and determine at least one task evaluation dimension corresponding to the target task based on the matching result.

[0047] In this embodiment, since different task evaluation strategies correspond to different evaluation dimensions, the execution entity will first determine the evaluation dimension keywords corresponding to the task evaluation strategy, and then match at least one task keyword with the evaluation dimension keywords, thereby determining the task evaluation dimension corresponding to the target task based on the matching results. The task evaluation dimensions here are generally multiple. If d is used to represent the task evaluation dimension, then the task evaluation dimension corresponding to the target task can be expressed as 、 , where n≥1.

[0048] As an example, when the task type of the target task is a creative task, the task evaluation dimensions of the target task determined may be: high quality, word count compliance, and instruction compliance.

[0049] Therefore, the task evaluation strategy corresponding to the task type is first determined, and then the task evaluation dimension of the target task is determined based on the task evaluation strategy, so that the corresponding task evaluation dimension is determined according to different task types, thereby improving the accuracy of the task evaluation dimension.

[0050] Step 305: Input the task evaluation dimension and the task result into the large language model, and output the evaluation result corresponding to the task evaluation dimension.

[0051] In this embodiment, the above-mentioned execution entity will first input the target task and task prompt words into the big model, and then output the task result corresponding to the target task. For example, when the target task is a creative task and the task prompt words are "Please create a lyrical text based on the following content", the target task and task prompt words are input into the big model, and the output is the generated lyrical text.

[0052] The execution entity then inputs the task evaluation dimensions and task results into the larger model, which then outputs the evaluation results corresponding to the task evaluation dimensions. This evaluation result is used to assess whether the task results meet the requirements. For example, if the task type is a translation task, the task result is the translation result, and the evaluation result is used to assess whether all the words in the task result are accurately translated.

[0053] Therefore, the large model can accurately determine the evaluation results corresponding to the task evaluation dimensions, and then the target reward can be calculated based on the evaluation results, thereby improving the accuracy of the information results.

[0054] Step 306: Determine target information of the target task based on the task evaluation dimension and the evaluation result.

[0055] Step 306 is basically the same as step 204 in the aforementioned embodiment. For the specific implementation method, reference can be made to the aforementioned description of step 204 and will not be repeated here.

[0056] from Figure 3 It can be seen that Figure 2 Compared to the corresponding embodiments, the information generation method in this embodiment highlights the steps of determining the task evaluation dimensions corresponding to the target task and the evaluation results corresponding to the task evaluation dimensions. This method first determines the task evaluation strategy corresponding to the task type, and then determines the task evaluation dimensions for the target task based on the task evaluation strategy. This determines the corresponding task evaluation dimensions based on different task types, thereby improving the accuracy of the task evaluation dimensions. Furthermore, the large model accurately determines the evaluation results corresponding to the task evaluation dimensions, and then the target information can be calculated based on the evaluation results, further improving the accuracy of the determined information.

[0057] Continue to refer Figure 4 , Figure 4 A process 400 of a third embodiment of the information generation method according to the present disclosure is shown. The information generation method comprises the following steps: Step 401: Determine the task type of the target task.

[0058] Step 402: Determine the task evaluation dimension corresponding to the target task according to the task prompt word and task type of the target task.

[0059] Step 403: Generate an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result.

[0060] Steps 401-403 are basically the same as steps 201-203 of the aforementioned embodiment. For specific implementation methods, reference can be made to the aforementioned description of steps 201-203, which will not be repeated here.

[0061] Step 404: Determine the priority order of at least one task evaluation dimension.

[0062] In this embodiment, the execution subject of the information generation method (for example Figure 1 The server 105 shown in FIG. 105 ) uses the large model to evaluate the priorities of multiple task evaluation dimensions, thereby generating a priority order of the multiple task evaluation dimensions. When the evaluation model outputs the results, the priorities of different evaluation dimensions are often different.

[0063] For example, when evaluating a translation task, if some words are not translated, the score for the current translation will be very low. On the other hand, if all words are translated but some are not translated accurately, the score will be slightly higher. In other words, in translation tasks, the evaluation dimension of "whether all words are translated" takes precedence over the evaluation dimension of "whether words are translated accurately."

[0064] For example, the task evaluation dimension corresponding to the target task can be expressed as 、 , the priority order corresponding to these multiple task evaluation dimensions is: > .

[0065] Step 405 , traverse at least one task evaluation dimension according to the priority order, compare the evaluation result of the current task evaluation dimension with a preset threshold, and generate target information of the target task according to the comparison result.

[0066] In this embodiment, after determining the priority order, the execution subject will traverse from high priority to low priority in accordance with the priority order, that is, traverse multiple task evaluation dimensions in order from high to low. , obtain the evaluation result corresponding to the current task evaluation dimension, and compare it with the preset threshold, so as to determine the target information according to the comparison result, and the target information includes reward information.

[0067] Since the evaluation result is used to characterize whether the evaluation of the current dimension is met, the preset threshold is generally set to 0. That is, if the evaluation result is greater than 0, it means that the evaluation of the current dimension is met or partially met; if the evaluation result is equal to 0, it means that the evaluation of the current dimension is not met.

[0068] Furthermore, if the evaluation of the current dimension is not satisfied, the information of the previous task evaluation dimension is directly used as the final information, that is, the target information. If the evaluation of the current dimension is satisfied or partially satisfied, the current information of the current evaluation dimension is calculated based on the information of the previous task evaluation dimension and the evaluation score corresponding to the current evaluation dimension. The traversal continues until all task evaluation dimensions have been traversed or until the target information is generated, and the target information is output.

[0069] Therefore, by determining the priority order of the evaluation dimensions and calculating the target information according to the priority order, the information is satisfied layer by layer from high priority to low priority, thereby improving the comprehensiveness and accuracy of the information results, and thus improving the user experience.

[0070] In some optional implementations of this embodiment, step 405 includes: in response to determining that the evaluation result of the task evaluation dimension is equal to a preset threshold, determining that the target information is information of a previous task evaluation dimension, wherein the previous task evaluation dimension is a task evaluation dimension previous to the current task evaluation dimension.

[0071] In this implementation, since the execution subject will traverse all task evaluation dimensions in descending order of priority, if the execution subject determines the current task evaluation dimension Evaluation results If the value is equal to the preset threshold (for example, equal to 0), it means that the evaluation of the current dimension has not been satisfied. At this time, the reward information of the previous task evaluation dimension of the current task evaluation dimension will be directly used. Output as target information (target reward), that is, target reward information .

[0072] Through the step-by-step reward fusion calculation strategy based on the priority of the evaluation dimension, the evaluation results of each evaluation dimension are judged separately, making the final target reward information more accurate and targeted.

[0073] from Figure 4 It can be seen that Figure 3 Compared with the corresponding embodiments, the information generation method in this embodiment highlights the steps of calculating target information based on task evaluation dimensions and evaluation results, thereby judging the evaluation results of each evaluation dimension separately through a step-by-step reward fusion calculation strategy based on the priority of the evaluation dimension, so that the target information finally generated is more accurate and targeted.

[0074] Continue to refer Figure 5 , Figure 5The fourth embodiment of the information generation method according to the present disclosure is shown in process 500. The information generation method includes the following steps: Step 501: Determine the task type of the target task.

[0075] Step 502: Determine the task evaluation dimension corresponding to the target task according to the task prompt word and task type of the target task.

[0076] Step 503: Generate an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result.

[0077] Step 504: Determine the priority order of at least one task evaluation dimension.

[0078] Steps 501-504 are basically the same as steps 401-404 of the aforementioned embodiment. For specific implementation methods, reference can be made to the aforementioned description of steps 401-404, which will not be repeated here.

[0079] Step 505: Input the task prompt word and at least one task evaluation dimension into the large language model, and output an evaluation score corresponding to the at least one task evaluation dimension.

[0080] In this embodiment, the execution subject of the information generation method (for example Figure 1 The server 105 shown in the figure will input the task prompt words and all task evaluation dimensions into the big model, so that the big model can score the importance of each task evaluation dimension, and then output the evaluation score corresponding to each task evaluation dimension, that is, the evaluation score is used to represent the importance of the task evaluation dimension. The evaluation score is expressed as .

[0081] Step 506: traverse at least one task evaluation dimension in order of priority. For the current task evaluation dimension, in response to determining that the evaluation result of the task evaluation dimension is greater than a preset threshold, calculate the current information of the current evaluation dimension based on the information of the previous task evaluation dimension and the evaluation score corresponding to the current evaluation dimension.

[0082] In this embodiment, the execution subject will traverse all task evaluation dimensions in the order of priority of the task evaluation dimensions, for example, traverse in order from high to low priority, and if it is determined that the current task evaluation dimension Evaluation results Greater than a preset threshold (e.g. >0), indicating that the evaluation of the current task evaluation dimension is satisfied or partially satisfied, at which point the reward value (i.e., information) of the current task evaluation dimension is calculated. Specifically, it can be calculated according to the following formula : ; in, is the reward value of the evaluation dimension for the previous task, The evaluation score corresponding to the current evaluation dimension.

[0083] Step 507 : until the evaluation result of the task evaluation dimension is equal to the preset threshold, the information of the task evaluation dimension whose evaluation result is equal to the preset threshold is determined as the target information.

[0084] In this embodiment, the above-mentioned execution entity will continue to traverse, that is, traverse the next task evaluation dimension of the current task evaluation dimension according to the priority order of the task evaluation dimension, and judge the reward value of the next task evaluation dimension with the preset threshold, generate the target reward value of the target task according to the judgment result, and output the final generated target reward value.

[0085] Therefore, when the evaluation result of the current task evaluation dimension is greater than the preset threshold, the current reward value of the current evaluation dimension is calculated according to the reward value of the previous task evaluation dimension and the evaluation score corresponding to the current evaluation dimension, and the traversal is continued until the target reward value is obtained. Thus, the evaluation results of each evaluation dimension are judged separately through the step-by-step reward fusion calculation strategy based on the priority of the evaluation dimension, so that the final generated target reward value is more accurate and targeted.

[0086] In some optional implementations of this embodiment, it also includes: in response to determining that all task evaluation dimensions have been traversed and the evaluation results corresponding to each task evaluation dimension are not equal to a preset threshold, determining the information of the last traversed task evaluation dimension in at least one task evaluation dimension as the target information.

[0087] In this implementation, if all task evaluation dimensions have been traversed and the evaluation results for each dimension are greater than the preset threshold (i.e., all evaluation dimensions are satisfied), the information (reward information) of the last traversed task evaluation dimension will be output as the target information (target reward). This ensures that the target reward value can be generated in this case.

[0088] from Figure 5 It can be seen that Figure 4Compared with the corresponding embodiments, the information generation method in this embodiment highlights the step of calculating the target reward value based on the task evaluation dimension and the evaluation result. Therefore, when the evaluation result of the current task evaluation dimension is greater than the preset threshold, the current reward value of the current evaluation dimension is calculated based on the reward value of the previous task evaluation dimension and the evaluation score corresponding to the current evaluation dimension, and the traversal is continued until the target reward value is obtained. Therefore, through the step-by-step reward fusion calculation strategy based on the priority of the evaluation dimension, the evaluation results of each evaluation dimension are judged separately, so that the final generated target reward value is more accurate and targeted.

[0089] Continue to refer Figure 6 , Figure 6 The fifth embodiment of the information generation method according to the present disclosure is shown in process 600. The information generation method includes the following steps: Step 601: Determine the task type of the target task.

[0090] Step 602: Determine the task evaluation dimension corresponding to the target task according to the task prompt word and task type of the target task.

[0091] Step 603: Generate an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result.

[0092] Step 604: Determine target information of the target task based on the task evaluation dimension and the evaluation result.

[0093] Steps 601-604 are basically the same as steps 201-204 of the aforementioned embodiment. For specific implementation methods, reference can be made to the aforementioned description of steps 201-204, which will not be repeated here.

[0094] Step 605 : Adjust the parameters of the large language model according to the target information to obtain an adjusted large language model.

[0095] In this embodiment, the execution subject of the information generation method (for example Figure 1 The server 105 shown adjusts the parameters of the large model based on the generated target information. This means that the large model can be optimized based on the target reward information. Here, the large model is considered an intelligent agent, and the reward information is used as environmental feedback, driving the iterative model update.

[0096] from Figure 6 It can be seen that Figure 2Compared with the corresponding embodiments, the information generation method in this embodiment highlights the step of adjusting the parameters of the large model according to target information (such as target reward information). This embodiment generates target information by adopting an asynchronous mechanism, thereby improving the resource utilization of the GPU and avoiding waiting on the training end; and by using the target information to adjust the parameters of the large model to achieve optimization of the large model, thereby improving the training efficiency and optimization efficiency of the model, and this method supports large-scale training tasks.

[0097] Furthermore, in some application scenarios, a reward system for large-scale learning is also provided. This reward system adopts an asynchronous, batch, and highly scalable design and provides the following capabilities: 1) A plug-in Verifier design and a multi-Verifier combination mechanism replace the traditional single reward model approach, adapting to different types of RL tasks and making reward calculation more accurate and explainable. Here, each Verifier corresponds to a task type.

[0098] 2) Asynchronous reward calculation mechanism to maximize GPU resource utilization and avoid waiting on the training end.

[0099] 3) An independent Verifier evaluation system prevents the model from "learning" the reward calculation logic through training strategies.

[0100] 4) Highly concurrent architecture supports large-scale training tasks, ensuring that reward calculation does not become a training bottleneck.

[0101] As an integrated system, this reward system does not pursue a unified reward model. Instead, it can assign a more targeted Verifier or reward model (normalized to 0-1) to each query. This makes the rewards returned by the reward system sufficiently accurate and effectively guides model optimization.

[0102] Building on this system, we further provide a method for calculating reward information for different tasks. This reward information generation method uses a sample-inspired multi-level reward fusion method, combining multiple evaluation dimensions to calculate a more accurate and reasonable unified reward score for each query. This method involves sample-inspired multi-dimensional reward weight calculation and a tiered reward fusion strategy based on evaluation priority.

[0103] Specifically, the sample-inspired multi-dimensional reward weight calculation process is: For a given prompt, the big model is first required to analyze the main evaluation dimensions covered by the user needs behind the prompt. 、 .

[0104] Afterwards, the prompt and the extracted evaluation dimensions are fed back into the large model, and the model is asked to score the importance of these evaluation dimensions. 、 , and use the score as the weight for subsequent multi-dimensional reward fusion.

[0105] The tiered reward fusion strategy based on evaluation priority includes: When evaluating model output, different dimensions are prioritized differently. For example, when evaluating a translation task, if some words are not translated, the translation score will be low. On the other hand, if all words are translated but some are not accurately translated, the score will be slightly higher. In other words, for translation tasks, whether all words are translated takes precedence over whether the translation is accurate.

[0106] Therefore, a step-by-step reward fusion method based on evaluation priority is provided. Specifically: (1) For a given prompt, use the large model to evaluate the different dimensions of the current prompt 、 The priority order is: > .

[0107] (2) According to the obtained order, traverse from high priority to low priority step by step:

[0108] A. For the i-th dimension , and the evaluation result is , the dimension importance assessment score is ; 1) If =0, indicating that the evaluation of the current dimension is not satisfied, and the final reward is directly returned as r= ; 2) If >0, indicating that the evaluation of the current dimension is satisfied or partially satisfied. In this case, the reward of the current level is calculated as .

[0109] B. Repeat the above process until the final reward r is obtained.

[0110] The sample-inspired evaluation strategy aligns evaluation closely with current user requests, making coverage more targeted and alleviating the over-optimization and reward hacking issues associated with fixed evaluation dimensions. Furthermore, a reward fusion method based on evaluation priorities caters to human preferences and can mitigate redlining issues in the model to a certain extent. It can satisfy requests from high to low priorities in a tiered manner, thereby improving the user experience.

[0111] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an information generating device, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0112] like Figure 7 As shown, the information generating device 700 of this embodiment includes: a task type determination module 701, a task dimension determination module 702, an evaluation result determination module 703, and an information determination module 704. The task type determination module 701 is configured to determine the task type of the target task; the task dimension determination module 702 is configured to determine the task evaluation dimension corresponding to the target task based on the task prompt word and the task type of the target task; the evaluation result determination module 703 is configured to generate an evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result, wherein the task result is generated by the large language model based on the target task and the task prompt word; and the information determination module 704 is configured to determine the target information of the target task based on the task evaluation dimension and the evaluation result.

[0113] In this embodiment, the specific processing of the task type determination module 701, the task dimension determination module 702, the evaluation result determination module 703 and the information determination module 704 and the technical effects thereof can be referred to in the respective Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiment are not repeated here.

[0114] In some optional implementations of this embodiment, the task dimension determination module 702 is further configured to: determine the task evaluation strategy corresponding to the task type; extract at least one task keyword in the task prompt word; match at least one task keyword with the evaluation dimension keyword corresponding to the task evaluation strategy, and determine at least one task evaluation dimension corresponding to the target task based on the matching result.

[0115] In some optional implementations of this embodiment, the evaluation result determination module 703 is further configured to: input the task evaluation dimension and the task result into the large language model, and output the evaluation result corresponding to the task evaluation dimension.

[0116] In some optional implementations of this embodiment, the information determination module 704 includes: a priority determination submodule, configured to determine the priority order of at least one task evaluation dimension; an information calculation submodule, configured to traverse at least one task evaluation dimension according to the priority order, and for the current task evaluation dimension, compare the evaluation result of the task evaluation dimension with a preset threshold, and generate target information of the target task based on the comparison result, wherein the target information includes reward information.

[0117] In some optional implementations of this embodiment, the reward calculation submodule includes: a first reward calculation unit, configured to determine that the target information is information of the previous task evaluation dimension in response to determining that the evaluation result of the task evaluation dimension is equal to a preset threshold, wherein the previous task evaluation dimension is the task evaluation dimension previous to the current task evaluation dimension.

[0118] In some optional implementations of this embodiment, the above-mentioned information generating device 700 also includes: an evaluation score calculation module, configured to input the task prompt word and at least one task evaluation dimension into the large language model, and output an evaluation score corresponding to at least one task evaluation dimension, wherein the evaluation score is used to characterize the importance of the task evaluation dimension; and the information calculation submodule also includes: a second reward calculation unit, configured to respond to determining that the evaluation result of the task evaluation dimension is greater than a preset threshold, calculate the current information of the current evaluation dimension based on the information of the previous task evaluation dimension and the evaluation score corresponding to the current evaluation dimension; until the evaluation result of the task evaluation dimension is equal to the preset threshold, the information of the task evaluation dimension whose evaluation result is equal to the preset threshold is determined as the target information.

[0119] In some optional implementations of this embodiment, the above-mentioned information generating device 700 also includes: a judgment module, which is configured to determine the information of the last traversed task evaluation dimension in at least one task evaluation dimension as the target information in response to determining that all task evaluation dimensions have been traversed and the evaluation results corresponding to each task evaluation dimension are not equal to a preset threshold.

[0120] In some optional implementations of this embodiment, the information generating apparatus 700 further includes: an updating module configured to adjust parameters of the large language model according to the target information to obtain an adjusted large language model.

[0121] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0122] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0123] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0124] Multiple components in device 800 are connected to I / O interface 805, including: an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0125] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the information generation method. For example, in some embodiments, the information generation method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the information generation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the information generation method by any other suitable means (e.g., via firmware).

[0126] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0127] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. Such program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0128] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0130] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0131] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0132] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0133] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating information, comprising: Determine the task type of the target task; Determining a task evaluation dimension corresponding to the target task according to the task prompt word of the target task and the task type; Generating an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result, wherein the task result is generated by a large language model according to the target task and the task prompt word; Determine target information of the target task based on the task evaluation dimension and the evaluation result.

2. The method according to claim 1, wherein Determining the task evaluation dimension corresponding to the target task according to the task prompt word and the task type of the target task includes: Determine a task evaluation strategy corresponding to the task type; extracting at least one task keyword from the task prompt words; The at least one task keyword is matched with an evaluation dimension keyword corresponding to the task evaluation strategy, and at least one task evaluation dimension corresponding to the target task is determined based on the matching result.

3. The method according to claim 1, wherein Generating an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result includes: The task evaluation dimension and the task result are input into the large language model, and the evaluation result corresponding to the task evaluation dimension is output.

4. The method according to claim 2, wherein: Determining target information of the target task according to the task evaluation dimension and the evaluation result includes: determining a priority order of the at least one task evaluation dimension; Traverse the at least one task evaluation dimension according to the priority order, compare the evaluation result of the current task evaluation dimension with a preset threshold, and generate target information of the target task according to the comparison result, wherein the target information includes reward information.

5. The method according to claim 4, wherein Generating target information of the target task according to the comparison result includes: In response to determining that the evaluation result of the task evaluation dimension is equal to the preset threshold, the target information is determined to be information of a previous task evaluation dimension, wherein the previous task evaluation dimension is a task evaluation dimension previous to the current task evaluation dimension.

6. The method according to claim 5, further comprising: Inputting the task prompt word and the at least one task evaluation dimension into a large language model, and outputting an evaluation score corresponding to the at least one task evaluation dimension, wherein the evaluation score is used to represent the importance of the task evaluation dimension; and Generating target information of the target task according to the comparison result further includes: In response to determining that the evaluation result of the task evaluation dimension is greater than the preset threshold, calculating current information of the current evaluation dimension according to the information of the previous task evaluation dimension and the evaluation score corresponding to the current evaluation dimension; Until the evaluation result of the task evaluation dimension is equal to the preset threshold, the information of the task evaluation dimension whose evaluation result is equal to the preset threshold is determined as the target information.

7. The method according to claim 6, further comprising: In response to determining that all task evaluation dimensions have been traversed and the evaluation results corresponding to each task evaluation dimension are not equal to the preset threshold, information of the last traversed task evaluation dimension in the at least one task evaluation dimension is determined as the target information.

8. The method according to any one of claims 1 to 7, further comprising: Parameters of the large language model are adjusted according to the target information to obtain an adjusted large language model.

9. An information generating device comprising: a task type determination module, configured to determine a task type of a target task; A task dimension determination module is configured to determine a task evaluation dimension corresponding to the target task according to a task prompt word of the target task and the task type; an evaluation result determination module configured to generate an evaluation result corresponding to the task evaluation dimension based on the task evaluation dimension and the task result, wherein the task result is generated by the large language model based on the target task and the task prompt word; The information determination module is configured to determine the target information of the target task according to the task evaluation dimension and the evaluation result.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause the computer to execute the method according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.