Model training method, device, medium, and program product
Patent Information
- Application Number
- CN202611252558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,如果结构化的工具调用指令的工具名称、参数键和参数值等信息不准确,会直接导致工具调用失败或返回错误结果
[0035]在本申请实施例中,第二方面至第四方面在实现上述第一方面中的任意一种方法时,可以达到与上述第一方面中任意一种方法相同或相似的技术效果。
Smart Images

Figure CN122819352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method, device, medium, and program product. Background Technology
[0002] With the development of large language models (LLMs), these models not only need to understand users' natural language input, but also need to automatically plan and invoke external tools or services to complete user commands, such as content generation tools, translation tools, schedule management tools, and lifestyle service tools. In these scenarios, the output of a large language model is usually not plain text, but rather structured tool invocation instructions, including tool names, parameter keys, parameter values, and other information.
[0003] However, if the tool name, parameter keys, and parameter values in the structured tool invocation command are inaccurate, it will directly lead to tool invocation failure or return of incorrect results. For example, an incorrect tool name will result in the inability to match the correct service interface, an incorrect parameter key will cause the tool to fail to parse the required parameters correctly, and an incorrect parameter value will cause the tool to be invoked but the execution result to deviate from the user's expectations. Summary of the Invention
[0004] The purpose of this application is to provide a model training method, device, medium, and program product to improve the accuracy of structured instructions generated by the model.
[0005] The first aspect of this application provides a model training method applied to an electronic device. The method includes: acquiring training data, wherein the training data includes input information and corresponding standard invocation instructions; performing N rounds of iterative training on the model based on the training data to obtain a target model, where N is an integer greater than 0; wherein the i-th round of training in the N rounds of iterative training includes: processing the input information in the training data through the current model to generate a predicted invocation instruction, wherein both the predicted invocation instruction and the standard invocation instruction include M layers of elements, and the levels of the M layers of elements decrease sequentially, where M is an integer greater than or equal to 2; performing a layer-by-layer comparison between the predicted invocation instruction and the standard invocation instruction to obtain a reward value for the predicted invocation instruction, wherein during the layer-by-layer comparison, if the comparison of the j-th layer element fails, the reward value of the j-th layer element and the reward value of elements with a lower level than the j-th layer element are set to 0, where j is a positive integer less than or equal to M; updating model parameters based on the reward value, where i is a positive integer less than or equal to N.
[0006] In this embodiment, the predicted invocation instruction is compared layer by layer with the standard invocation instruction. If any layer fails the comparison, the reward value for the current layer and lower layers is set to 0. This ensures that the reward value strictly conforms to the logical dependency of the tool invocation: if a high-level element (e.g., tool name) is incorrect, even if a low-level element (e.g., parameter) matches by chance, it will not receive a reward. This fundamentally eliminates the "reward leakage" problem caused by accidental parameter matching, ensuring that the model receives the correct optimization signal during reinforcement learning. It also prevents the model from misjudging and reinforcing incorrect tool invocation behavior as valid behavior, thereby improving the accuracy of the structured instructions generated by the trained target model.
[0007] In one possible implementation of the first aspect described above, the M-level elements include: tool name, parameter key, and parameter value; wherein the tool name has a higher level than the parameter key, and the parameter key has a higher level than the parameter value.
[0008] In this embodiment, the reward calculation can be performed step by step according to the logical dependency order from tool name to parameter key to parameter value, ensuring that high-level errors will block the transmission of low-level rewards, avoiding the artificially high rewards caused by accidental correctness at low levels, and improving the accuracy and reliability of reward signals.
[0009] In one possible implementation of the first aspect above, the predicted invocation instruction and the standard invocation instruction are compared layer by layer, including: comparing the tool name in the predicted invocation instruction with the tool name in the standard invocation instruction; if the tool name in the predicted invocation instruction is different from the tool name in the standard invocation instruction, the tool name comparison is determined to be unsuccessful, and the reward value of the tool name, the reward value of the parameter key, and the reward value of the parameter value are set to 0.
[0010] In this embodiment, when the tool name matching fails, the reward values for the tool name, parameter key, and parameter value are all set to 0, ensuring that the entire call does not generate any reward when the tool name is incorrect. This reflects the hard dependency constraint of tool calls: the tool name is the entry point for the call, and an incorrect tool name means that the call cannot be executed at all in the actual system. By setting the reward to zero, the model is prevented from receiving only partial rewards due to parameter matching when the tool name is incorrect, allowing the model to prioritize learning the correct tool name recognition ability in the early stages of training.
[0011] In one possible implementation of the first aspect above, the step-by-step comparison between the predicted invocation instruction and the standard invocation instruction further includes: if the tool name in the predicted invocation instruction is the same as the tool name in the standard invocation instruction, the tool name comparison is determined to be successful, the reward value of the tool name is set to the first preset score, and the comparison between the parameter key in the predicted invocation instruction and the parameter key in the standard invocation instruction continues; if the parameter key in the predicted invocation instruction is different from the parameter key in the standard invocation instruction, the parameter key comparison is determined to be unsuccessful, and the reward value of the parameter key and the reward value of the parameter value are set to 0.
[0012] In this embodiment, a first preset score is awarded when the tool name match is successful, and the parameter key match continues; when the parameter key match fails, the rewards for both the parameter key and parameter value are set to 0. This allows the model to receive positive reinforcement signals when the tool name is correct, encouraging the model to prioritize learning the correct tool selection. Simultaneously, it ensures that even if the parameter value is correct, no reward is given when the parameter key is incorrect, effectively preventing the model from receiving false rewards due to an incorrect parameter key but a coincidentally matched parameter value.
[0013] In one possible implementation of the first aspect described above, the step-by-step comparison between the predicted invocation instruction and the standard invocation instruction further includes: if the parameter key in the predicted invocation instruction is the same as the parameter key in the standard invocation instruction, the parameter key comparison is deemed successful, the reward value of the parameter key is set to a first preset score, and the semantic constraint type of the parameter value is determined. Based on the semantic constraint type of the parameter value, the parameter value in the predicted invocation instruction is compared with the parameter value in the standard invocation instruction, wherein the semantic constraint type includes strict parameters and semantic parameters; if the semantic constraint type is a strict parameter, and the parameter value in the predicted invocation instruction is exactly the same as the parameter value in the standard invocation instruction, the reward value of the parameter value is set to a third preset score; if the semantic constraint type is a strict parameter, and the parameter value in the predicted invocation instruction is not exactly the same as the parameter value in the standard invocation instruction, the reward value of the parameter value is set to 0; if the semantic constraint type is a semantic parameter, the semantic similarity between the parameter value in the predicted invocation instruction and the parameter value in the standard invocation instruction is calculated, and the reward value of the parameter value is determined based on the semantic similarity, wherein the magnitude of the reward value of the parameter value is positively correlated with the magnitude of the semantic similarity.
[0014] In this embodiment, a first preset score reward is given when the parameter key matching passes, thus positively reinforcing the model. Furthermore, a distinction is made between strict parameters and semantic parameters: strict parameters are judged using a perfect match, while semantic parameters are judged using semantic similarity, with the reward value positively correlated with similarity. This differentiated processing solves the reward sparsity problem caused by uniformly applying strict string matching to all parameters. For semantic parameters such as text descriptions and keywords, a reward can still be obtained even if the model's generated expression is not completely identical to the reference answer but is semantically correct. This alleviates the problem of the model receiving zero rewards for a long time in the early stages of training due to inaccurate generation, while also ensuring the accuracy requirements of key parameters such as identifiers and numbers, achieving refined reward modeling in a heterogeneous parameter space.
[0015] In one possible implementation of the first aspect above, the strict parameters include at least one of the following: identifier parameters and number parameters; the semantic parameters include at least one of the following: text description parameters and keyword parameters.
[0016] In one possible implementation of the first aspect above, obtaining the reward value of the predicted invocation instruction includes: weighted summation of the reward values of the M-layer elements to obtain the invocation correctness reward value of the predicted invocation instruction; obtaining the format validity reward value of the predicted invocation instruction, the format validity reward value being used to indicate whether the format of the predicted invocation instruction conforms to a preset valid tool invocation format; and determining the reward value of the predicted invocation instruction based on the invocation correctness reward value and the format validity reward value.
[0017] In this embodiment, a call correctness reward value is obtained by weighted summation of the reward values of the M-layer elements. Simultaneously, a format validity reward value is introduced, decoupling the evaluation of the content correctness and format standardization of the tool call and fusing them into the final reward value. This subjectes the model to dual constraints during training: pursuing the correctness of tool names, parameter keys, and parameter values while ensuring the standardization of the output format, preventing the model from generating call formats that cannot be parsed by the system. Compared to schemes that only evaluate content correctness, this scheme comprehensively covers the validity requirements of tool call instructions, improving the model's usability in actual deployment.
[0018] In one possible implementation of the first aspect above, obtaining the format validity reward value of the predicted invocation instruction includes: if the format of the predicted invocation instruction conforms to a preset valid tool invocation format, setting the format validity reward value to a preset positive value; if the format of the predicted invocation instruction does not conform to a valid tool invocation format, setting the format validity reward value to a preset negative value.
[0019] In this embodiment, predicted invocation commands that conform to the format specifications are given a preset positive reward, while predicted invocation commands that do not conform to the format specifications are given a preset negative penalty. This clear positive and negative reward signal guides the model to quickly learn the correct tool invocation format. The introduction of negative rewards effectively suppresses the model from outputting incorrectly formatted or unparseable invocation commands, enabling the model to form format validity constraints early in training, significantly reducing the generation of invalid outputs, and improving training efficiency and the usability of model outputs. Compared to schemes that only give positive rewards for correct formats, the negative penalty mechanism provides stronger constraints.
[0020] In one possible implementation of the first aspect above, the number of standard invocation instructions and the number of predicted invocation instructions are both greater than 1; and the predicted invocation instructions are compared with the standard invocation instructions layer by layer, including: matching the corresponding predicted invocation instructions for each standard invocation instruction in a preset order; wherein, in the process of matching the corresponding predicted invocation instructions for each standard invocation instruction, the predicted invocation instructions that are the same as the first-level elements of the standard invocation instructions are selected from the predicted invocation instructions that have never been matched for matching.
[0021] In this embodiment, when multiple standard calling instructions and multiple predicted calling instructions exist, the first-level element (tool name) is selected from the predicted calling instructions that have never been matched before for each standard calling instruction to be matched, and the matched predicted calling instructions are marked as non-reusable. This greedy matching mechanism solves the problem of uncertain correspondence between predicted calling instructions and standard calling instructions in parallel multi-tool calling scenarios, and avoids multiple standard calling instructions repeatedly matching the same predicted calling instruction, which would lead to inflated rewards. At the same time, the constraint of "non-reusable matching" ensures the fairness of reward evaluation: each predicted calling instruction can only be included in the matching score once, so that the reward value truly reflects the comprehensive calling capability of the model in multi-tool parallel scenarios. This greedy matching mechanism reduces the computational complexity of matching while ensuring the accuracy of evaluation, and is suitable for efficient evaluation of large-scale parallel calls in actual training scenarios.
[0022] In one possible implementation of the first aspect above, the number of training data is greater than 1; and the method further includes: determining the difficulty value of each training data based on the reward value corresponding to each training data in the i-th round of training; and determining the training data used in the (i+1)-th round of training based on the difficulty value of each training data.
[0023] In this embodiment, by converting the reward value of each training data point in the i-th training round into a difficulty value, and redetermining the training data used in the (i+1)-th training round based on the difficulty value, the dynamic adjustment of the training data distribution as the model's capabilities change is achieved. This allows the training data composition to dynamically "self-evolve" with the model's capabilities: as the model's capabilities improve, samples that were originally highly difficult (training data) may become of moderate difficulty and be included in the training, ensuring the model always learns around its current capability boundaries. This avoids the problem of insufficient high-value samples in the later stages of training with a fixed training set, improving overall training efficiency and sample utilization.
[0024] In one possible implementation of the first aspect above, the difficulty value of each training data is determined based on the reward value corresponding to each training data in the i-th round of training, including: determining the success score of each training data based on the reward value corresponding to each training data in the i-th round of training; and determining the difficulty value of each training data based on the success score of each training data, wherein the higher the success score, the lower the difficulty value.
[0025] In this embodiment, a success score is determined by the reward values of multiple predicted invocation instructions, and the difficulty value is inferred from the success score (the higher the success score, the lower the difficulty value), establishing a quantitative mapping relationship from reward value to difficulty value. This scheme utilizes the model's own reinforcement learning reward feedback to dynamically evaluate the difficulty, enabling the difficulty value to be automatically updated as the model's capabilities improve, without requiring additional manual costs. Furthermore, the evaluation results are highly consistent with the model's actual performance, providing an accurate data foundation for subsequent capability boundary-aware sampling.
[0026] In one possible implementation of the first aspect above, determining the training data used in the (i+1)th round of training based on the difficulty value of each training data includes: determining the capability boundary of the current model based on the difficulty value of each training data; selecting target training data from the training data whose difference between the difficulty value and the capability boundary is less than or equal to a preset boundary width; and determining the training data used in the (i+1)th round of training based on the target training data.
[0027] In this embodiment, by determining the capability boundary of the current model (i.e., the region at the boundary in the difficulty value distribution) and selecting samples (training data) with difficulty values close to the capability boundary as target training data, the main training samples of the model are always concentrated in the high-value region with difficulty values close to the capability boundary. This avoids the convergence problem caused by encountering overly difficult samples in the early stages of training, and also avoids the waste of resources caused by repeatedly learning overly simple samples in the later stages of training, significantly improving training efficiency and model convergence speed.
[0028] In one possible implementation of the first aspect above, the training data used in the (i+1)th round of training is determined based on the target training data, including: selecting a preset number of training data with the highest difficulty value from the training data whose difficulty value is greater than the capability boundary, as difficult training data; and using the target training data and the difficult training data as the training data used in the (i+1)th round of training.
[0029] In this embodiment, in addition to the boundary samples, a predetermined number of the most difficult samples (those with difficulty values higher than the capability boundary) are introduced as challenging training data. Both types of samples constitute the next round of training data. This combined sampling strategy of "primarily boundary samples, supplemented by challenging samples" allows the model to focus on learning samples near the capability boundary while continuously exposing itself to a small number of highly challenging samples. This avoids the risk of the model getting trapped in local optima due to learning only boundary samples for an extended period. Furthermore, the model can continuously attempt tasks of increasing difficulty, thereby gradually expanding its capability boundary and achieving a balance between training stability and exploratory learning.
[0030] In one possible implementation of the first aspect above, determining the difficulty value of each training data point based on the reward value corresponding to each training data point in the i-th training round further includes: obtaining the reward value corresponding to each training data point in the (i-1)-th training round; normalizing the reward value corresponding to each training data point in the (i-1)-th training round to obtain the normalized reward value of each training data point in the (i-1)-th training round; obtaining the updated cumulative reward value of each training data point in the (i-1)-th training round based on the normalized reward value of each training data point in the (i-1)-th training round and the cumulative reward value of each training data point in the (i-2)-th training round; and determining the difficulty value of each training data point in the i-th round based on the updated cumulative reward value of each training data point in the (i-1)-th training round; where i is an integer greater than 2.
[0031] In this embodiment, the cumulative reward value of round i-1 is obtained by normalizing the reward value and then performing an exponential moving average with the cumulative reward value of round i-2, thereby determining the difficulty value of round i. This scheme ensures that the difficulty value of the current round is not solely determined by the reward value of the previous round, but rather incorporates historical cumulative information into the difficulty estimation of the current round through an exponential moving average. The historical cumulative reward value reflects the stable performance level of the model during long-term training. By weighting and fusing it with the reward value of the current round, the interference of single-round reward fluctuations on the difficulty assessment is effectively smoothed. This avoids drastic fluctuations in the difficulty value caused by single-round sampling noise or model randomness, allowing the sampling distribution of training data to evolve steadily with the model's capabilities, thus enhancing the robustness of the training loop.
[0032] A second aspect of this application provides an electronic device, comprising: a memory for storing instructions executable by one or more processors of the electronic device; and a processor, one of the processors of the electronic device, for executing the instructions stored in the memory to implement any of the methods of the first aspect described above.
[0033] A third aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a device, cause the device to implement any of the methods described in the first aspect.
[0034] The fourth aspect of this application provides a computer program product including instructions that, when executed on a device, cause the device to implement any of the methods described in the first aspect.
[0035] In the embodiments of this application, when the second to fourth aspects implement any one of the methods in the first aspect, they can achieve the same or similar technical effects as any one of the methods in the first aspect. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0037] Figure 1 The image shown is an interface diagram of the Smart Assistant application corresponding to the scenario where a user uses the Smart Assistant in the mobile phone 10 to specify a travel plan;
[0038] Figure 2 The image shown depicts the interface of the Smart Assistant app, representing a scenario where a user uses the Smart Assistant on a mobile phone to translate text.
[0039] Figure 3 The image shown is an interface diagram of the Smart Assistant application corresponding to the scenario where a user uses the Smart Assistant in the mobile phone 10 to practice speaking.
[0040] Figure 4 A flowchart of a model training method is shown according to an embodiment of this application;
[0041] Figure 5 A schematic diagram of a hierarchical gating and greedy matching mechanism is shown according to an embodiment of this application;
[0042] Figure 6According to an embodiment of this application, the logic and schematic diagram of an asymmetric semantic reward mechanism based on parametric semantic attributes (i.e., the semantic constraint type of parameter values) are shown.
[0043] Figure 7 According to an embodiment of this application, a logical schematic diagram of a reward-driven training set self-evolution and reinforcement learning closed loop is shown.
[0044] Figure 8 A schematic diagram comparing a fixed threshold selection range with the adaptive selection range of this application is shown according to an embodiment of the present application;
[0045] Figure 9 An architecture diagram of a model training method based on capability boundary awareness and hierarchical rewards is shown according to an embodiment of this application;
[0046] Figure 10 A schematic diagram of the hardware structure of an electronic device 100 is shown according to an embodiment of this application. Detailed Implementation
[0047] The illustrative embodiments of this application include, but are not limited to, a model training method, apparatus, medium, and program product.
[0048] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described clearly and in detail below with reference to the accompanying drawings.
[0049] In scenarios where a large language model needs to automatically plan and call external tools or services to complete user commands, such as calling content generation tools, translation tools, schedule management tools, and life service tools, the output of the large language model is usually a structured tool call command, including information such as tool name, parameter key, and parameter value.
[0050] For example, refer to Figure 1 The image shown is an interface diagram of the Smart Assistant application corresponding to the scenario where a user uses the Smart Assistant on their mobile phone to specify a travel plan.
[0051] like Figure 1 As shown, the user interface 10a includes a session area and a model output area. The session area 101 displays the user-inputted text message "Please create a detailed plan for my trip through the water towns of Jiangnan." The model output area 102 displays the response content generated by the model based on the tool's call results, "A two-day in-depth plan for a trip through the water towns of Jiangnan...".
[0052] For example, refer to Figure 2 The image shown depicts the interface of the Smart Assistant app, representing a scenario where a user uses the Smart Assistant on their mobile phone to translate text.
[0053] like Figure 2As shown, the user interface 20a includes a session area 201 and a model output area 202. The session area 201 displays the user-inputted text message "Translate this text into English". The model output area 202 displays the response content generated by the model based on the tool's call result, i.e., the translated English text.
[0054] For example, refer to Figure 3 The image shown depicts the interface of the Smart Assistant app, representing a scenario where a user uses the Smart Assistant on their mobile phone to practice speaking.
[0055] like Figure 3 As shown, the user interface 30a includes a conversation area 301 and a model output area 302. The conversation area 301 displays the user-inputted text message "Practice English speaking with me". The model output area 302 displays the response content generated by the model based on the tool's call results, i.e., the interactive guidance text for speaking practice.
[0056] In practice, tool invocation commands are low-level commands executed in the background. Their structure includes three levels: tool name, parameter key, and parameter value, which are invisible to the user. The user-visible response content is natural language text generated by the model based on the tool invocation result. Only when the tool name, parameter key, and parameter value are all correct can the intelligent assistant correctly invoke the corresponding tool service and generate a response content that meets the user's expectations.
[0057] Understandable. Figures 1 to 3 The application scenarios shown are merely illustrative and do not constitute a limitation on scenarios where large language models need to automatically plan and call external tools or services to complete user commands.
[0058] However, as mentioned earlier, if the tool name, parameter keys, and parameter values of a structured tool invocation command are inaccurate, it will directly lead to tool invocation failure or return of incorrect results. For example, an incorrect tool name will prevent the correct service interface from being matched, an incorrect parameter key will prevent the tool from correctly parsing the required parameters, and an incorrect parameter value will cause the tool to be invoked but the execution result to deviate from the user's expectations.
[0059] For example, when a user requests "Please create a detailed itinerary for a leisurely trip through the water towns of Jiangnan," the agent needs to invoke the tourism content generation capability. The model should output a tool invocation command similar to the following: `aigc(type=travel guide)`. Here, `aigc` is the tool name, `type` is the parameter key, and `travel guide` is the parameter value. It's understandable that the electronic device can only correctly invoke the service and generate results if the tool name, parameter key, and parameter value are all correct.
[0060] Currently, training large language models for tool invocation typically optimizes their tool invocation capabilities through supervised learning or reinforcement learning. Reward calculation in the reinforcement learning phase often employs a rigid matching mechanism: for each predicted invocation command output by the model, it checks whether the tool name, parameter key, and parameter value match the standard invocation command. Corresponding reward values are assigned to the matching results of the tool name, parameter key, and parameter value. Finally, the reward values for the tool name, parameter key, and parameter value are weighted and summed to obtain the reward value for the predicted invocation command.
[0061] For example, given the user input text "Please create a detailed travel plan for the Jiangnan water towns," the standard command is `aigc(type=travel guide)`. If the model outputs a predicted command `weather(city=Jiangnan water towns)`, since the tool name `weather` for the predicted command is different from the tool name `aigc` for the standard command, but the parameter values "Jiangnan water towns" and "travel guide" for the predicted command are both text types and have a textual association with "Jiangnan water towns" in the user's request, the reinforcement learning phase may still award a partial reward to the parameter value. This reduces the accuracy and effectiveness of the reward signal and may cause the trained model to generate incorrect commands during actual use, affecting the user experience.
[0062] In view of this, embodiments of this application provide a model training method applied to an electronic device. The method includes: acquiring training data, wherein the training data includes: input information and standard calling instructions corresponding to the input information; performing N rounds of iterative training on the model based on the training data to obtain a target model, where N is an integer greater than 0; wherein the i-th round of training in the N rounds of iterative training includes: processing the input information in the training data through the current model to generate a predicted calling instruction, wherein both the predicted calling instruction and the standard calling instruction include M layers of elements, and the level of the M layers of elements decreases sequentially, where M is an integer greater than or equal to 2; performing layer-by-layer comparison between the predicted calling instruction and the standard calling instruction to obtain a reward value for the predicted calling instruction, wherein during the layer-by-layer comparison, if the comparison of the j-th layer element fails, the reward value of the j-th layer element and the reward value of the element with a lower level than the j-th layer element are set to 0, where j is a positive integer less than or equal to M; updating the model parameters based on the reward value; wherein i is a positive integer less than or equal to N.
[0063] It is understandable that by comparing the predicted invocation instructions with the standard invocation instructions layer by layer, and setting the reward value of the current level and lower levels to 0 when the comparison fails at any level, the reward value strictly conforms to the logical dependency relationship of the tool invocation: when a high-level element (e.g., tool name) is incorrect, a low-level element (e.g., parameters) will not receive a reward even if it matches by chance. In this way, the "reward leakage" problem caused by accidental parameter matching can be fundamentally eliminated, ensuring that the model receives the correct optimization signal during reinforcement learning, preventing the model from misjudging and reinforcing erroneous tool invocation behaviors as valid behaviors, thereby improving the accuracy of the structured instructions generated by the trained target model.
[0064] The model training method provided in the embodiments of this application is described in detail below with reference to the accompanying drawings.
[0065] For example, Figure 4 A flowchart illustrating a model training method is shown according to an embodiment of this application. It can be understood that... Figure 4 The execution subject of the flowchart shown is electronic device 100. For ease of description, the following will refer to... Figure 4 When describing the flowchart shown, the execution subject of the flowchart will not be repeated.
[0066] like Figure 4 As shown, this process includes, but is not limited to:
[0067] S401: Obtain training data, which includes: input information and the standard calling instructions corresponding to the input information.
[0068] In some embodiments, the input information may be natural language instructions, and the standard invocation instructions are the correct structured tool invocation instructions corresponding to the input information.
[0069] For example, for Figure 1 The travel planning scenario shown has the input message "Please create a detailed travel plan for the water towns of Jiangnan," and the corresponding standard command is aigc(type=travel guide). Figure 2 The translation scenario shown has the input message "Translate this text into English", and the corresponding standard command is translate(text=text to be translated, target_language=English). Figure 3 The spoken English practice scenario shown has the input message "Practice spoken English with me", and the corresponding standard command is spoken_practice(language=English).
[0070] It is understood that the number of training data can be 1 (i.e., 1 sample data) or greater than 1 (i.e., multiple sample data), and this application does not impose any restrictions on this.
[0071] S402: Perform N rounds of iterative training on the model based on the training data to obtain the target model, where N is an integer greater than 0.
[0072] In some embodiments, the i-th round of training in N rounds of iterative training includes: processing the input information in the training data through the current model to generate a predicted invocation instruction, wherein both the predicted invocation instruction and the standard invocation instruction include M layers of elements, and the levels of the M layers of elements decrease sequentially, where M is an integer greater than or equal to 2; performing a layer-by-layer comparison between the predicted invocation instruction and the standard invocation instruction to obtain the reward value of the predicted invocation instruction, wherein, during the layer-by-layer comparison, if the comparison of the j-th layer element fails, the reward value of the j-th layer element and the reward value of the element with a lower level than the j-th layer element are set to 0, where j is a positive integer less than or equal to M; updating the model parameters based on the reward value; where i is a positive integer less than or equal to N.
[0073] It is understandable that by comparing the predicted invocation instructions with the standard invocation instructions layer by layer, and setting the reward value of the current level and lower levels to 0 when the comparison fails at any level, the reward value strictly conforms to the logical dependency relationship of the tool invocation: when a high-level element (e.g., tool name) is incorrect, a low-level element (e.g., parameters) will not receive a reward even if it matches by chance. In this way, the "reward leakage" problem caused by accidental parameter matching can be fundamentally eliminated, ensuring that the model receives the correct optimization signal during reinforcement learning, preventing the model from misjudging and reinforcing erroneous tool invocation behaviors as valid behaviors, thereby improving the accuracy of the structured instructions generated by the trained target model.
[0074] In some embodiments, the M-layer element includes: tool name, parameter key, and parameter value; wherein the tool name has a higher level than the parameter key, and the parameter key has a higher level than the parameter value.
[0075] It is understandable that reward calculation can proceed step by step according to the logical dependency order from tool name to parameter key to parameter value, ensuring that high-level errors will block the transmission of low-level rewards, avoiding inflated rewards caused by accidental correctness at low levels, and improving the accuracy and reliability of reward signals.
[0076] In some embodiments, the electronic device 100 performs a layer-by-layer comparison between the predicted call instruction and the standard call instruction by comparing the tool name in the predicted call instruction with the tool name in the standard call instruction. If the tool name in the predicted call instruction is different from the tool name in the standard call instruction, it is determined that the tool name comparison has failed, and the reward value of the tool name, the reward value of the parameter key, and the reward value of the parameter value are set to 0.
[0077] If the tool name in the predicted invocation command is the same as the tool name in the standard invocation command, the tool name comparison is deemed successful, the reward value of the tool name is set to the first preset score, and the comparison of the parameter key in the predicted invocation command with the parameter key in the standard invocation command continues; if the parameter key in the predicted invocation command is different from the parameter key in the standard invocation command, the parameter key comparison is deemed unsuccessful, and the reward value of the parameter key and the reward value of the parameter value are set to 0.
[0078] If the parameter key in the predicted invocation instruction is the same as the parameter key in the standard invocation instruction, the parameter key comparison is considered successful. The reward value of the parameter key is set to the first preset score, and the semantic constraint type of the parameter value is determined. Based on the semantic constraint type of the parameter value, the parameter value in the predicted invocation instruction is compared with the parameter value in the standard invocation instruction. The semantic constraint type includes strict parameters and semantic parameters. If the semantic constraint type is a strict parameter and the parameter value in the predicted invocation instruction is exactly the same as the parameter value in the standard invocation instruction, the reward value of the parameter value is set to the third preset score. If the semantic constraint type is a strict parameter and the parameter value in the predicted invocation instruction is not exactly the same as the parameter value in the standard invocation instruction, the reward value of the parameter value is set to 0. If the semantic constraint type is a semantic parameter, the semantic similarity between the parameter value in the predicted invocation instruction and the parameter value in the standard invocation instruction is calculated, and the reward value of the parameter value is determined based on the semantic similarity. The magnitude of the reward value of the parameter value is positively correlated with the magnitude of the semantic similarity.
[0079] Semantic similarity can be obtained through methods such as text vector cosine similarity, edit distance normalization, and lightweight semantic discrimination models. This application does not limit the specific methods used.
[0080] For example, for the user-input text "Please create a detailed travel plan for the Jiangnan water towns," the standard calling instruction is aigc(type=travel guide). If the model outputs a predicted calling instruction of weather(city=Jiangnan water towns), since the tool name weather of the predicted calling instruction is different from the tool name aigc of the standard calling instruction, this embodiment directly determines that the tool name comparison has failed and sets the reward value of the tool name, the reward value of the parameter key, and the reward value of the parameter value to 0. Even if the parameter value "Jiangnan water towns" of the predicted calling instruction and the parameter value "travel guide" of the standard calling instruction are both text types and have a textual association with "Jiangnan water towns" in the user request, no reward will be generated, thus avoiding the reward leakage problem caused by accidental parameter matching.
[0081] For example, if the model outputs a prediction invocation command of aigc (theme=Jiangnan Water Town Travel Guide), the tool name aigc is correct, but the parameter key theme is different from the standard parameter key type, this embodiment determines that the parameter key comparison has failed and sets both the reward value of the parameter key and the reward value of the parameter value to 0. The model will not receive any reward for the parameter value being semantically close to the standard value, thereby guiding the model to prioritize learning the correct parameter key structure.
[0082] For example, if the model outputs a prediction call instruction as aigc (type=Jiangnan Water Town Travel Guide), the tool name aigc is correct, the parameter key type is correct, and the parameter value "Jiangnan Water Town Travel Guide" is not completely consistent with the standard parameter value "Travel Guide" in literal form. However, since "Travel Guide" is a semantic parameter, this application embodiment calculates the semantic similarity between "Jiangnan Water Town Travel Guide" and "Travel Guide", and gives continuous rewards based on semantic similarity, so that the model can still get rewards when the parameter value expression is not completely consistent with the reference answer but the semantics are correct.
[0083] It's understandable that when a tool name match fails, the reward values for the tool name, parameter key, and parameter value are all set to 0, ensuring that the entire call generates no reward when the tool name is incorrect. This reflects the hard dependency constraint of tool calls: the tool name is the entry point for the call, and an incorrect tool name means that the call cannot be executed at all in the actual system. By setting the rewards to zero, the model avoids receiving partial rewards only for parameter matching when the tool name is incorrect, allowing the model to prioritize learning the correct tool name recognition ability in the early stages of training.
[0084] Furthermore, when the tool name match is successful, a first preset score is awarded, and the parameter key match continues; when the parameter key match fails, the rewards for both the parameter key and parameter value are set to 0. This ensures the model receives positive reinforcement when the tool name is correct, encouraging it to prioritize learning the correct tool selection. Simultaneously, it ensures that even if the parameter value is correct, an incorrect parameter key will not result in a reward, effectively preventing the model from receiving false rewards due to an incorrect parameter key but a coincidentally matching parameter value.
[0085] Furthermore, a first-preset score reward is given when the parameter key matching passes, providing positive reinforcement to the model. Building upon this, a distinction is made between strict parameters and semantic parameters: strict parameters are judged using exact matching, while semantic parameters are judged using semantic similarity, with the reward value positively correlated with similarity. This differentiated approach addresses the reward sparsity problem caused by uniformly applying strict string matching to all parameters. For semantic parameters such as text descriptions and keywords, a reward is still given even if the model's generated expression is not completely identical to the reference answer but is semantically correct. This alleviates the problem of the model receiving zero rewards for a long time in the early stages of training due to inaccurate generation, while also ensuring the accuracy requirements of key parameters such as identifiers and numbers, achieving refined reward modeling in a heterogeneous parameter space.
[0086] In some embodiments, strict parameters include at least one of the following: identifier parameters and number parameters; semantic parameters include at least one of the following: text description parameters and keyword parameters.
[0087] For example, parameters such as user identifiers (e.g., user_id), order numbers (e.g., order_id), and timestamps (e.g., timestamps) are strict parameters. These parameters must exactly match their standard values (parameter values in the standard calling command) during tool calls; otherwise, the call will fail or return incorrect results. Parameters such as search keywords (e.g., query), rewritten question text (e.g., rewritten_text), and summary prompts (e.g., summary_prompt) are semantic parameters. These parameters allow for semantically approximate expressions. For example, although "tourist attractions in city A" and "places to visit in city A" are not exactly the same in literal form, their high semantic similarity allows the model to still receive a reward.
[0088] In some embodiments, the reward value of the parameter value layer can be determined with reference to formula (1).
[0089] Formula (1)
[0090] in, This represents the reward value of the parameter value layer between the s-th standard invocation instruction and the t-th predicted invocation instruction; This indicates a gated function with parameter type, based on parameter value. The semantic constraint type determines whether reward calculation is enabled; This represents the reward function for the corresponding parameter type, used to calculate the specific reward value for that parameter type. This represents the parameter value of the s-th standard call instruction; This represents the parameter value of the t-th predicted call instruction.
[0091] In other embodiments, the reward value of the parameter value layer can be determined with reference to formula (2).
[0092] Formula (2)
[0093] in, This represents the reward value of the parameter value layer between the s-th standard invocation instruction and the t-th predicted invocation instruction; For exponential functions, if the semantic constraint type of the parameter value is strict, the value is 1; otherwise, it is 0. For exponential functions, if the semantic constraint type of the parameter value is a semantic parameter, the value is 1; otherwise, it is 0. The reward function represents a strict parameter (e.g., a reward of 1 when the predicted parameter value of the calling instruction is exactly the same as the parameter value of the standard calling instruction, and 0 otherwise). The reward function represents the semantic parameters and is used to calculate continuous rewards based on semantic similarity.
[0094] In some embodiments, for semantic parameters, the reward function You can refer to formula (3) to determine it.
[0095] Formula (3)
[0096] in, Represents the reward value for semantic parameters; This represents the semantic similarity between the parameter values in the predicted invocation instruction and the parameter values in the standard invocation instruction (the value range is usually from 0 to 1). This represents the preset semantic reward threshold (the value range is usually from 0 to 1).
[0097] In some embodiments, the electronic device 100 obtains the reward value of the predicted call instruction by: weighted summing of the reward values of the M-layer elements to obtain the call correctness reward value of the predicted call instruction; obtaining the format validity reward value of the predicted call instruction, which indicates whether the format of the predicted call instruction conforms to a preset valid tool call format; and determining the reward value of the predicted call instruction based on the call correctness reward value and the format validity reward value.
[0098] It's understandable that a weighted sum of the reward values from the M-layer elements yields a call correctness reward value, while a format validity reward value is introduced. This decouples the evaluation of the tool call's content correctness and format compliance, fusing them into the final reward value. This subjectes the model to dual constraints during training: pursuing the correctness of tool names, parameter keys, and parameter values while ensuring the standardization of the output format, preventing the model from generating call formats that the system cannot parse. Compared to schemes that only evaluate content correctness, this approach comprehensively covers the validity requirements of tool call commands, improving the model's usability in practical deployments.
[0099] In some embodiments, the electronic device 100 obtains the format validity reward value of the predicted call instruction by the following method: if the format of the predicted call instruction conforms to a preset valid tool call format, the format validity reward value is set to a preset positive value; if the format of the predicted call instruction does not conform to a valid tool call format, the format validity reward value is set to a preset negative value.
[0100] For example, if the predicted invocation instruction output by the model is missing a necessary tool invocation structure block (such as a missing tool name or parameter key), then the predicted invocation instruction does not conform to the valid tool invocation format, and the format validity reward value is a preset negative value; if the predicted invocation instruction output by the model contains a complete tool name, parameter key, and parameter value, then the predicted invocation instruction conforms to the valid tool invocation format, and the format validity reward value is a preset positive value.
[0101] It's understandable that predictive invocation commands that conform to the format specifications are given a pre-set positive reward, while those that don't are given a pre-set negative penalty. This clear distinction between positive and negative reward signals guides the model to quickly learn the correct tool invocation format. The introduction of negative rewards effectively suppresses the model from outputting incorrectly formatted or unparseable invocation commands, allowing the model to form format validity constraints early in training, significantly reducing invalid outputs and improving training efficiency and the usability of model outputs. Compared to schemes that only reward correct formats, the negative penalty mechanism provides stronger constraints.
[0102] In some embodiments, the reward value of the predicted invocation instruction can be determined by the following formula (4).
[0103] Formula (4)
[0104] in, This represents the reward value predicted for the invocation instruction. The preset format reward weight is used to adjust the proportion of the reward value for format validity. The reward value represents the validity of the format. This indicates the correctness reward value for the call.
[0105] In some embodiments, the reward value for the correctness of the predicted invocation instruction can be determined by the following formula (5).
[0106] Formula (5)
[0107] in, This represents the s-th standard call instruction. With the t-th prediction call instruction Reward value for correct calls between them; This represents the tool name of the s-th standard call instruction. This represents the tool name of the t-th predicted invocation instruction. This represents the reward value obtained when the tool name comparison passes. This represents the reward value of the parameter key layer. This represents the reward value for the parameter value layer.
[0108] Among them, the reward value of the parameter key layer It can be determined by the following formula (6).
[0109] Formula (6)
[0110] in, This represents the reward value of the parameter key layer. This represents the reward value obtained when the parameter key layer alignment passes. This represents the parameter key of the s-th standard call instruction. This represents the parameter key of the t-th predicted call instruction.
[0111] Among them, the reward value of the parameter key layer It can be determined by the following formula (7).
[0112] Formula (7)
[0113] in, This represents the reward value of the parameter value layer. This represents the reward value obtained when the parameter value layer comparison passes. This represents the parameter key of the s-th standard call instruction. This represents the parameter key of the t-th predicted call instruction. This represents the parameter value of the s-th standard call instruction. This represents the parameter value of the t-th predicted call instruction.
[0114] Specifically, if the tool name in the invocation command is predicted... The tool name in the standard calling command If they are the same, the overall reward is the reward value of the tool name layer. Reward value of parameter key layer and the reward value of the parameter value layer The sum. If the tool name in the invocation command is predicted. The tool name in the standard calling command If they are different, the gate is closed, and the s-th standard call instruction is executed. With the t-th prediction call instruction The correctness bonus for calls between them is 0.
[0115] It's understandable that the tool name is the entry point for tool invocation. An incorrect tool name means the invocation cannot be executed at all in the actual system. Therefore, even if a lower-level element matches by chance, it should not receive any reward. This formula implements a hard dependency constraint: high-level errors directly block the transmission of lower-level rewards.
[0116] If the parameter key in the call instruction is predicted Parameter keys in standard calling instructions If they are completely identical, the parameter bond layer receives a positive reward. If the parameter key in the call instruction is predicted. Parameter keys in standard calling instructions If they are not completely consistent, the reward for the parameter key layer is 0, and the reward for the parameter value layer is... It was also set to 0.
[0117] It is understandable that parameter keys are identifiers for parameter values. An incorrect parameter key means that the tool cannot correctly parse the required parameter. Even if the parameter value itself is correct, it cannot be correctly identified and used by the tool. Therefore, the parameter value layer should not receive any reward when the parameter key is incorrect.
[0118] It is understandable that when both the number of standard call instructions and the number of predicted call instructions are greater than 1, the call correctness reward value of the predicted call instructions can be determined by the following formula (8).
[0119] Formula (8)
[0120] in, This indicates the reward value for correct call execution; y represents the total number of standard invocation instructions, x represents the model's input data (i.e., training data), and y represents the model's output data (i.e., predicted invocation instructions). This represents the s-th standard call instruction. The reward value for correct call execution.
[0121] In some embodiments, the format validity reward value can be determined by the following formula (9).
[0122] Formula (9)
[0123] in, The format validity reward value is represented by x, which represents the model's input data (i.e., training data), and y represents the model's output data (i.e., the prediction call instruction). This indicates the criteria for determining whether the model output y conforms to the preset valid tool call format. If true (i.e., the format is correct), the format validity bonus is 1. If true (i.e., the format is incorrect), the format validity bonus is -1.
[0124] It is understood that 1 (preset positive value) and -1 (preset negative value) are merely illustrative examples. In other embodiments, the preset positive value and the preset negative value may be other values, and this application does not limit them.
[0125] In some embodiments, the number of standard invocation instructions and the number of predicted invocation instructions are both greater than 1 (e.g., multiple parallel invocation instructions exist to complete different tool invocations for input information). The electronic device 100 can perform a layer-by-layer comparison between predicted invocation instructions and standard invocation instructions by the following method: matching the corresponding predicted invocation instructions for each standard invocation instruction in a preset order; wherein, in the process of matching the corresponding predicted invocation instructions for each standard invocation instruction, the predicted invocation instructions that are the same as the first-level element of the standard invocation instructions are selected from the predicted invocation instructions that have never been matched for matching.
[0126] In other words, for scenarios involving the parallel invocation of multiple tools, a greedy matching process is performed sequentially on each standard invocation instruction: among the unused predicted invocation instructions with the same tool name, the one with the highest score is selected for matching; if there is no set of predicted invocation instructions with the same tool name as the standard invocation instruction, the matching score corresponding to that standard invocation instruction is 0. Here, the set of predicted invocation instructions with the same tool name is used to limit the matching range, and the matching score is used to measure the similarity between the predicted invocation instruction and the standard invocation instruction.
[0127] The method for matching each standard calling instruction with the corresponding predicted calling instruction can refer to the method of comparing the predicted calling instruction with the standard calling instruction layer by layer, as described above, and will not be elaborated here.
[0128] For example, suppose the standard invocation instructions include standard invocation instruction A and standard invocation instruction B. The tool name for standard invocation instruction A is get_weather, and the tool name for standard invocation instruction B is send_email. The predicted invocation instructions include predicted invocation instruction 1, predicted invocation instruction 2, and predicted invocation instruction 3. The tool name for predicted invocation instruction 1 is get_weather, the tool name for predicted invocation instruction 2 is get_weather, and the tool name for predicted invocation instruction 3 is send_email.
[0129] Following the greedy matching logic: First, match the standard call instruction A. From all the unmatched predicted call instructions, find predicted call instructions 1 and 2 with the same tool name, get_weather. Select the one with the highest matching score (e.g., predicted call instruction 1) and mark predicted call instructions 1 and 2 as "matched". Then, match the standard call instruction B. From the remaining unmatched predicted call instructions (predicted call instruction 3), find the predicted call instruction with the same tool name, send_email, and match it with the standard call instruction B.
[0130] In some embodiments, the above greedy matching logic can be implemented by formula (10).
[0131] Formula (10)
[0132] in, This represents the s-th standard call instruction. The corresponding call correctness reward value, This indicates the s-th standard call instruction. The set of all predicted invocation instructions with the same tool name. This represents the s-th standard call instruction. With the t-th prediction call instruction The reward value for the correctness of the calls between them.
[0133] It is understandable that when multiple standard invocation instructions and multiple predicted invocation instructions exist, the system sequentially selects the same Level 1 element (tool name) from the predicted invocation instructions that have never been matched before for each standard invocation instruction, and marks the matched predicted invocation instructions as non-reusable. This greedy matching mechanism solves the problem of uncertain correspondence between predicted invocation instructions and standard invocation instructions in parallel multi-tool invocation scenarios, avoiding multiple standard invocation instructions repeatedly matching the same predicted invocation instruction, which would lead to inflated rewards. At the same time, the constraint of "non-reusable matching" ensures the fairness of reward evaluation: each predicted invocation instruction can only be included in the matching score once, so that the reward value truly reflects the model's comprehensive invocation capability in multi-tool parallel scenarios. This greedy matching mechanism reduces the computational complexity of matching while ensuring evaluation accuracy, and is suitable for efficient evaluation of large-scale parallel invocation in actual training scenarios.
[0134] In some embodiments, the number of training data is greater than 1 (i.e., batch training is performed using data from the sample set). The electronic device 100 can also determine the difficulty value of each training data based on the reward value corresponding to each training data in the i-th round of training; and determine the training data used in the (i+1)-th round of training based on the difficulty value of each training data.
[0135] It is understandable that by converting the reward values of each training data point in the i-th training round into difficulty values, and then redetermining the training data used in the (i+1)-th training round based on these difficulty values, the distribution of training data is dynamically adjusted according to changes in model capability. This allows the composition of training data to dynamically "self-evolve" with model capability: as model capability improves, samples that were originally highly difficult (training data) may become of moderate difficulty and be included in training, ensuring the model always learns around the current capability boundary. This avoids the problem of insufficient high-value samples in the later stages of training with a fixed training set, thus improving overall training efficiency and sample utilization.
[0136] In some embodiments, the electronic device 100 may determine the difficulty value of each training data by: determining the success score of each training data based on the reward value corresponding to each training data in the i-th round of training; and determining the difficulty value of each training data according to the success score of each training data, wherein the higher the success score, the lower the difficulty value.
[0137] In some embodiments, based on the success score, the difficulty value of the training data is defined as 1 minus the success score, and the difficulty value of the training data ranges from 0 to 1. Specifically, when the difficulty value of the training data is close to 0, it indicates that the model can stably complete the task; when the difficulty value of the training data is close to 1, it indicates that the current model still has difficulty completing the task correctly; the higher the difficulty value of the training data, the more challenging the task is for the current model.
[0138] For example, for a given training data, the model trained in the i-th round generates multiple prediction commands (e.g., G candidate response trajectories) from that training data. The model obtains the reward value for each predicted instruction and uses the average of these reward values as the success score for the training data. If the success score of a training data is 0.9, its difficulty value is 0.1, indicating that the model can stably complete the task; if the success score of a training data is 0.2, its difficulty value is 0.8, indicating that the model still has difficulty completing the task correctly.
[0139] In some embodiments, the success score of the training data can be determined by formula (11).
[0140] Formula (11)
[0141] in, Indicates the first The success score for the s-th training data in the training round; G represents the number of candidate response trajectories generated by the model for the s-th training data; This represents the reward value of the g-th candidate response trajectory in the s-th training data. This represents the g-th prediction call instruction generated by the model for the s-th training data.
[0142] Understandingly, this approach establishes a quantitative mapping from reward value to difficulty value by determining the success score based on the reward values of multiple prediction commands, and then using the success score to infer the difficulty value (a higher success score corresponds to a lower difficulty value). This solution utilizes the model's own reinforcement learning reward feedback to dynamically evaluate the difficulty, enabling the difficulty value to automatically update as the model's capabilities improve. This requires no additional manual effort, and the evaluation results are highly consistent with the model's actual performance, providing an accurate data foundation for subsequent capability boundary-aware sampling.
[0143] In some embodiments, the electronic device 100 may determine the training data used in the (i+1)th round of training by the following method: determining the capability boundary of the current model based on the difficulty value of each training data; selecting target training data from the training data whose difference between the difficulty value and the capability boundary is less than or equal to a preset boundary width; and determining the training data used in the (i+1)th round of training based on the target training data.
[0144] For example, suppose there are 5 training data points with difficulty values of 0.1, 0.3, 0.5, 0.7, and 0.9, respectively, then the capability boundary is 0.5. If the preset boundary width is 0.2, then the training data with difficulty values between 0.3 and 0.7 (the training data with difficulty values of 0.3, 0.5, and 0.7) will be selected as the target training data.
[0145] It is understandable that by determining the capability boundary of the current model (i.e., the region at the boundary in the difficulty value distribution) and selecting samples (training data) with difficulty values close to the capability boundary as target training data, the main training samples of the model are always concentrated in the high-value region with difficulty values close to the capability boundary. In this way, the non-convergence problem caused by exposure to overly difficult samples in the early stage of training is avoided, as well as the waste of resources caused by repeatedly learning overly simple samples in the later stage of training, which significantly improves training efficiency and model convergence speed.
[0146] In some embodiments, the capability boundary of the current model can be determined by formula (12).
[0147] Formula (12)
[0148] in, Indicates the first The capability boundary of the current model in the i-th training round is the quantified value of the overall capability level of the model in the i-th training round. This represents the training set (the collection of all training data). This represents the total number of training data in the training set; This represents the difficulty value of the s-th training data in the i-th round of training.
[0149] In some embodiments, the target training data can be determined by formula (13).
[0150] Formula (13)
[0151] in, This represents the target training data (i.e., the target training data set) in the i-th round of training. This represents the s-th training data; This represents the training set (the collection of all training data). This represents the difficulty value of the s-th training data in the i-th round of training; This represents the capability boundary of the current model in the i-th round of training; Indicates the preset boundary width.
[0152] In some embodiments, the electronic device 100 may determine the training data to be used in the (i+1)th round of training based on the target training data by the following method: selecting a preset number of training data with the highest difficulty value among the training data with difficulty values greater than the capability boundary from the training data, and using them as difficult training data; and using the target training data and the difficult training data as the training data to be used in the (i+1)th round of training.
[0153] For example, suppose there are 5 training data points with difficulty values of 0.1, 0.3, 0.5, 0.7, and 0.9, respectively, then the capability boundary is 0.5. If the preset boundary width is 0.2 and the preset quantity is 1, then the training data point with the highest difficulty value (i.e., the training data with a difficulty value of 0.9) is selected from the training data points with a difficulty value greater than the capability boundary of 0.5 (the training data points with difficulty values of 0.7 and 0.9) as the difficult training data. This difficult training data, together with the target training data points (the training data points with difficulty values of 0.3, 0.5, and 0.7), constitutes the training data used in the (i+1)th round of training.
[0154] Understandably, in addition to the boundary samples, a predetermined number of the most difficult samples from the training data with difficulty values exceeding the capability boundary are introduced as challenging training data. Both types of samples constitute the next round of training data. This combined sampling strategy of "primarily boundary samples, supplemented by challenging samples" allows the model to focus on learning samples near its capability boundary while continuously exposing itself to a small number of highly challenging samples. This avoids the risk of the model getting trapped in local optima due to learning only boundary samples for an extended period. Furthermore, the model can continuously attempt tasks of increasing difficulty, thereby gradually expanding its capability boundary and achieving a balance between training stability and exploratory learning.
[0155] In some embodiments, difficult training data can be determined by formula (14).
[0156] Formula (14)
[0157] in, This represents the difficult training data in the i-th round of training (i.e., the set of highly difficult training data selected). This represents the training data in the i-th round of training where the difficulty value is greater than the ability boundary; This represents the difficulty value of the s-th training data in the i-th round of training; Indicates by difficulty value Sort from highest to lowest and take the top K (K is the preset number).
[0158] In some embodiments, the training data used in the (i+1)th round of training can be determined by formula (15).
[0159] Formula (15)
[0160] in, This represents the training data used in the (i+1)th round of training; This represents the target training data in the i-th round of training; This represents the difficult training data in the i-th round of training.
[0161] In some embodiments, the electronic device 100 may further determine the difficulty value of each training data by: obtaining the reward value corresponding to each training data in the (i-1)th round of training; normalizing the reward value corresponding to each training data in the (i-1)th round of training to obtain the normalized reward value of each training data in the (i-1)th round of training; obtaining the updated cumulative reward value of each training data in the (i-1)th round based on the normalized reward value of each training data in the (i-1)th round of training and the cumulative reward value of each training data in the (i-2)th round; and determining the difficulty value of each training data in the i-th round based on the updated cumulative reward value of each training data in the (i-1)th round; where i is an integer greater than 2.
[0162] It is understandable that the difficulty value of round i is determined by normalizing the reward value of round i-1 and then performing an exponential moving average with the cumulative reward value of round i-2. This yields the cumulative reward value of round i-1, thus determining the difficulty value of round i. This approach ensures that the difficulty value of the current round is not solely determined by the reward value of the previous round, but rather incorporates historical cumulative information into the difficulty estimation of the current round through an exponential moving average. Historical cumulative reward values reflect the stable performance level of the model during long-term training. By weighting and fusing them with the reward value of the current round, the interference of single-round reward fluctuations on difficulty assessment is effectively smoothed. This avoids drastic fluctuations in the difficulty value caused by single-round sampling noise or model randomness, allowing the sampling distribution of training data to evolve steadily with the model's capabilities, thereby enhancing the robustness of the training loop.
[0163] In some embodiments, the reward value corresponding to each training data point during the (i-1)th training round can be determined by the following method: after the (i-1)th training round, the training data is processed using the model obtained after training. ( The reward values corresponding to the Q candidate response trajectories are used as the corresponding training data. The reward value corresponding to the (i-1)th round of training.
[0164] For example, training data During the (i-1)th training round, the corresponding reward value can generate Q candidate response trajectories (predicted call instructions), and the corresponding reward value is calculated to be determined by formula (16).
[0165] ,in Formula (16)
[0166] in, Representing training data The reward value corresponding to the (i-1)th training round; Representing training data The reward value corresponding to the generated Q candidate response trajectories (predicted call instructions); Representing training data The reward value corresponding to the first candidate response trajectory among the Q generated candidate response trajectories; Representing training data The reward value corresponding to the Qth candidate response trajectory among the generated Q candidate response trajectories.
[0167] In some embodiments, training data The normalized reward value during the (i-1)th training round can be determined by formula (17).
[0168] Formula (17)
[0169] in, Representing training data The normalized reward value during the (i-1)th training round; Representing training data The reward value corresponding to the (i-1)th training round; This represents the minimum reward value corresponding to all training data during the (i-1)th training round. This represents the maximum reward value for all training data during the (i-1)th training round; This indicates a preset small constant (used to prevent division by zero, such as...). ).
[0170] In some embodiments, the updated training data The cumulative reward value in the (i-1)th round can be determined by formula (18).
[0171] Formula (18)
[0172] in, This indicates the updated training data. The cumulative reward value in round i-1; This indicates the updated training data. The cumulative reward value in round i-2; Representing training data The normalized reward value during the (i-1)th training round; This represents the difficulty update coefficient (e.g., a hyperparameter between 0 and 1), used to control the weight of the cumulative reward value in round i-2 and the normalized reward value during round i-1 training.
[0173] In some embodiments, training data The difficulty value in the i-th round can be determined by formula (19).
[0174] Formula (19)
[0175] in, Representing training data The difficulty value in the i-th round; This indicates the updated training data. The cumulative reward value in round i-1.
[0176] Understandable. Figure 4 The data processing procedure shown is only one example; in other embodiments, Figure 4 The model training process shown may also include the deployment and inference of the target model. Those skilled in the art can also provide further details based on actual needs. Figure 4 The application does not restrict the addition, deletion, or merging of other content in the training process shown.
[0177] In some embodiments, the electronic device 100 may also employ a group relative policy optimization (GRPO) algorithm to update model parameters.
[0178] For example, for each training cycle, first, start with the training set at the current moment. Obtain training data For each training data Generate G candidate response trajectories Where G is an integer greater than 1, such as 4, 8, etc., this application does not impose any restrictions on it. Then, using Figure 4 The method shown calculates the reward value for each candidate response trajectory. .
[0179] Since the reward scale may differ between different training data, the electronic device 100 can use within-group normalization to calculate the relative advantage of the group in order to improve training stability.
[0180] For example, for training data The relative advantage of the g-th candidate response trajectory is calculated as the sum of the within-group reward value minus the within-group average reward, divided by the within-group reward standard deviation and a preset constant. Here, the within-group average reward is the average of the reward values of all candidate responses within the same sample, and the within-group reward standard deviation measures the dispersion of the within-group reward. Within-group normalization can eliminate differences in reward scale between different samples, improving training stability.
[0181] For example, training data The relative advantage of the group of the g-th candidate response trajectory can be determined by formula (20).
[0182] Formula (20)
[0183] in, Representing training data The relative advantage of the group of the g-th candidate response trajectory Representing training data The reward value of the g-th candidate response trajectory, This represents the average reward within the group (i.e., the average reward of the G candidate response trajectories in the training data). This represents the standard deviation of rewards within a group (used to measure the dispersion of rewards within a group). This indicates a preset small constant (used to prevent division by zero, such as...). ).
[0184] For example, the average reward within the group It can be determined by formula (21).
[0185] Formula (21)
[0186] in, This represents the average reward within the group. Representing training data The reward value of the g-th candidate response trajectory.
[0187] For example, the standard deviation of group rewards It can be determined by formula (22).
[0188] Formula (22)
[0189] in, Indicates the standard deviation of the group's rewards. Representing training data The reward value of the g-th candidate response trajectory, This indicates the average reward within the group.
[0190] It's understandable, training data The relative advantage of the group of the g-th candidate response trajectory Used to indicate training data The reward of the g-th candidate response trajectory deviates from the average value by the standard deviation within the group. If A positive number indicates that the reward for that trajectory is higher than the group average and needs to be strengthened. If... A negative number indicates that the reward for that trajectory is below the group average and needs to be suppressed.
[0191] It is understandable that by normalizing within groups, the difference in reward scale between different samples is eliminated, and the advantage value is no longer affected by the size of the absolute reward value of the sample, thus improving training stability.
[0192] Finally, the model parameters are updated using the pruning policy objective function, which is the smaller of the ratio of the new and old policy probabilities multiplied by the dominance function and the ratio of the pruned probabilities multiplied by the dominance function. The ratio of the new and old policy probabilities measures the degree of probability change between the current and old policies when generating the response, while the pruning threshold limits the magnitude of policy updates to prevent drastic changes in the model during a single update.
[0193] For example, the objective function of the pruning strategy can be determined by formula (23).
[0194] Formula (23)
[0195] in, Represent the objective function of the pruning strategy; This represents the probability ratio between the old and new strategies, used to measure the degree of change in the probability of generating the response between the current and old strategies; The advantage value measures how well the response is compared to the average level. This represents the pruning threshold, used to limit the magnitude of policy updates and prevent drastic changes in the model during a single update.
[0196] Among them, the probability ratio of the new strategy to the old strategy It can be determined by formula (24).
[0197] Formula (24)
[0198] in, This represents the ratio of the probabilities of the new and old strategies. This indicates the probability that the current policy (the model being trained) will generate this response. This indicates the probability that the old strategy (the model before the update) would generate this response.
[0199] For example, for input Before training, the model had a probability of generating a correct call to aigc (type=travel guide) of 0.3. After training, the probability of generating a correct call to aigc (type=travel guide) increased to 0.6. It is 0.6. It is 0.3. The value is 2.
[0200] In other embodiments, to prevent the policy from shifting excessively during reinforcement learning, the electronic device 100 may introduce a reference policy constraint term to limit the distance between the current policy (the model being trained) and the reference policy (the model before the update or a reference model with fixed parameters).
[0201] For example, KL divergence (KL divergence) can be used to limit the distance between the current policy and the reference policy. KL divergence measures the difference in probability distributions between the current and reference policies, while KL constraint weights control the strength of the constraint. This constraint allows the model to progressively enhance complex tool-calling behavior while maintaining generative stability.
[0202] For example, the objective function of the pruning strategy after introducing the reference strategy constraint can be determined by formula (25).
[0203] Formula (25)
[0204] Where L represents the objective function of the pruning strategy after introducing the reference strategy constraint term; Represent the objective function of the pruning strategy; This represents the KL constraint weights (hyperparameters) used to control the strength of the constraints; This represents the KL divergence between the current strategy and the reference strategy; Indicates the current policy (the model being trained); This indicates a reference strategy (the model before the update or a reference model with fixed parameters).
[0205] It is understandable that, through the above mechanism, the model can gradually enhance the behavior of invoking complex tools while maintaining the stability of generation.
[0206] This application embodiment uses a hierarchical gating mechanism to decompose tool invocation commands into three levels, comparing and evaluating them level by level according to the tool name, parameter key, and parameter value. To more clearly illustrate the specific implementation of the above-mentioned hierarchical comparison and reward calculation, the following describes... Figure 5 Provide a detailed description.
[0207] For example, Figure 5 A schematic diagram of a hierarchical gating and greedy matching mechanism is shown according to an embodiment of this application. For example... Figure 5 As shown, the logic diagram includes four parts: an input layer, a hierarchical gating reward structured evaluation layer, a greedy matching pattern layer, and an overall reward aggregation structure layer.
[0208] At the input layer, standard invocation instructions (i.e., reference invocation) are obtained. ) and the prediction invocation instructions output by the model (i.e., prediction invocation) ).
[0209] In the hierarchical gating reward structured evaluation layer, comparisons are performed level by level in the order of tool name layer, parameter key layer, and parameter value layer.
[0210] At the tool name level, compare reference calls. Tool Name With predictive invocation Tool Name Are they consistent? If so... and If they are the same, proceed to the parameter key layer for further evaluation and overall scoring. ;like and If they are different, the gate is closed, and the overall score is adjusted. (Refer to the relevant description of formula (5) above for details).
[0211] At the parameter key level, evaluation continues only if the tool name is correct, comparing reference calls. parameter key With predictive invocation parameter key Are they consistent? If so... and If they are consistent, the parameter key layer receives a positive reward. ,Right now ;like and If inconsistent, then (Refer to the relevant description of formula (6) above for details).
[0212] At the parameter value level, evaluation continues only if both the tool name and parameter key are correct, comparing with the reference call. parameter values With predictive invocation parameter values Are they consistent? If so... and If the values are consistent, the parameter value layer receives a positive reward. ,Right now .like and If inconsistent, then (Refer to the relevant description of formula (7) above for details).
[0213] The above layer-by-layer gating logic can be summarized as follows: if the tool name is incorrect, the reward value is 0; the parameter key is evaluated only if the tool name is correct; the parameter value is evaluated only if the parameter key is correct. This layer-by-layer gating reward design ensures that the reward value strictly conforms to the logical dependencies of the tool call, thereby avoiding the "reward leakage" problem caused by accidental parameter matching.
[0214] In the greedy matching pattern layer, when there are multiple standard invocation instructions and multiple predicted invocation instructions, it is necessary to match the corresponding predicted invocation instruction for each standard invocation instruction. Specifically, for the s-th standard invocation instruction... (i.e., reference call) The electronic device searches for a match among the predicted call instructions that have not yet been matched. All predicted invocation instructions with the same tool name constitute the s-th standard invocation instruction. A set of all predictive invocation instructions with the same tool name .like If it is not an empty set, then from Select with The highest-scoring predicted instruction is matched against the given instruction. ;like It is an empty set, that is, it does not exist with If a predictive invocation instruction has the same tool name, then the standard invocation instruction will not be matched with any predictive invocation instruction. (Refer to the relevant description of formula (10) above for details). The predicted call instruction that has been matched will be marked as "matched" and will not participate in the matching process of subsequent reference calls.
[0215] In the overall reward aggregation structure layer, the overall reward is determined based on the matching score of each standard invocation instruction. The overall reward includes a format validity reward. Rewards for correct tool usage Two parts. Format legality reward. The evaluation factor is used to determine whether the format of the predicted tool invocation instruction conforms to the preset valid tool invocation format. A value of 1 is assigned if it conforms, and a value of -1 if it does not. A fixed negative reward is given when the model output lacks necessary tool invocation structure blocks (such as missing tool names or parameter keys). Tool invocation correctness reward. The average score for matching all standard call instructions. The overall reward is a weighted sum of the format validity reward and the tool call correctness reward.
[0216] It is understandable that, through the aforementioned hierarchical gating mechanism, the model can learn the ability to recognize tool names, generate parameter structures, and predict parameter values layer by layer during training, thereby gradually improving the accuracy of tool invocation. The greedy matching mechanism solves the problem of uncertain correspondence between predicted invocations and standard invocations in parallel multi-tool invocation scenarios, avoiding multiple standard invocation instructions repeatedly matching the same predicted invocation instruction, which would lead to inflated rewards.
[0217] for Figure 5 The specific comparison method of the parameter value layer in the layered gating mechanism shown below, namely the differentiated reward mechanism between strict parameters and semantic parameters, will be discussed in conjunction with the following. Figure 6 Please provide a detailed explanation.
[0218] For example, Figure 6 According to an embodiment of this application, the logic and schematic diagram of an asymmetric semantic reward mechanism based on parameter semantic attributes (i.e., the semantic constraint type of parameter values) are shown.
[0219] like Figure 6 As shown, the logic of the asymmetric semantic reward mechanism includes four parts: the initial stage, the processing logic of strict parameters, the processing logic of semantic parameters, and the total reward output.
[0220] In the initial phase, the tool parameter pattern is extended by defining a parameter semantic type (i.e., semantic constraint type) for each parameter value. Parameter semantic types include two types: strict parameters and semantic parameters. Strict parameters describe parameters that must be matched exactly, such as user identifiers, order numbers, and timestamps. These parameters must be completely consistent with the standard values (parameter values in the standard call instructions) in the tool call; otherwise, the call will fail or return incorrect results. Semantic parameters describe parameters that allow semantic approximation, such as search keywords, question rewrite text, and summary prompts. These parameters allow for semantically approximate expressions.
[0221] For strict parameters, it is determined whether the parameter value in the predicted call instruction (predicted value) is completely consistent with the parameter value in the standard call instruction (reference value). If they are completely consistent, a reward is given (Reward=1); if they are not completely consistent, the reward is zero (Reward=0). This perfect matching mechanism ensures the accuracy requirements of the critical parameters.
[0222] For semantic parameters, the semantic similarity between the parameter value in the predicted invocation instruction and the reference parameter value in the standard invocation instruction is calculated. Semantic similarity can be obtained through methods such as text vector cosine similarity, edit distance normalization, and lightweight semantic discrimination models; this application does not limit the specific method used. After obtaining the semantic similarity, a continuous soft reward is applied based on this similarity: when the semantic similarity is low, the reward value is low or even zero; when the semantic similarity is high, the reward value increases accordingly.
[0223] During the reward calculation process, different reward functions are dynamically selected and combined based on the parameter type to output the total reward. The parameter value reward function can be expressed as follows: (Refer to the relevant description of formula (2) above for details).
[0224] It is understandable that, through the above mechanism, even if the parameter text generated by the model is not exactly the same as the reference answer in literal form, as long as their semantic expression is similar, they can still obtain continuous rewards. This effectively alleviates the reward sparsity problem caused by the traditional strict matching mechanism, enabling the model to obtain denser effective reward signals in the early stages of training and improving the exploration efficiency of reinforcement learning.
[0225] The above text combined Figure 5 and Figure 6 This paper describes in detail the hierarchical gating reward mechanism and the parameter semantic reward mechanism of the embodiments of this application. Based on this, the embodiments of this application also propose a reward-driven training set self-evolution mechanism, which will be discussed below. Figure 7 Please provide a detailed explanation.
[0226] For example, Figure 7 According to an embodiment of this application, a logical schematic diagram of a reward-driven training set self-evolution and reinforcement learning closed loop is shown.
[0227] like Figure 7 As shown, the closed-loop logic includes policy optimization, hierarchical heterogeneous reward evaluation, current model capability evaluation, sample resampling, and retraining.
[0228] In the strategy optimization stage, the electronic device 100 can sample input samples (training data) from the current training set, generate multiple candidate response trajectories for each training data using the current model, and calculate the reward value for each candidate response trajectory based on the aforementioned hierarchical gating reward mechanism. Subsequently, the electronic device 100 uses the GRPO algorithm to update the model parameters, thereby increasing the probability of generating high-reward responses and decreasing the probability of generating low-reward responses. Specifically, for details on updating model parameters using the GRPO algorithm, please refer to the above description of updating model parameters using the GRPO algorithm, which will not be repeated here.
[0229] In the hierarchical heterogeneous reward stage, the electronic device 100 uses the updated model to regenerate multiple candidate response trajectories from the training data, and calculates the reward value for each candidate response trajectory using the aforementioned hierarchical gating reward mechanism and an asymmetric semantic reward mechanism based on parametric semantic attributes. For details, please refer to the above. Figure 4 The relevant descriptions in the document are not repeated here.
[0230] In the current model capability evaluation stage, electronic device 100 calculates a success score and difficulty value for each training data point based on the reward value from the previous hierarchical heterogeneous reward stage. Specifically, for each training data point, electronic device 100 generates multiple candidate response trajectories using the current model and calculates the reward value for each trajectory. The average of these reward values is used as the success score for that training data point; a higher success score indicates better model performance on that sample. Subsequently, electronic device 100 defines the difficulty value of the training data as 1 minus the success score. A difficulty value closer to 0 indicates that the model can stably complete the task, while a difficulty value closer to 1 indicates that the model still struggles to complete the task correctly. For details on the calculation methods of success score and dynamic difficulty, please refer to the above. Figure 4 The relevant descriptions in the document are not repeated here.
[0231] In the sample resampling stage, the electronic device 100 first determines the capability boundary of the current model based on the difficulty values of all training data. Then, the electronic device 100 constructs a capability boundary band based on this boundary, selecting training data whose difficulty value differs from the capability boundary by a preset boundary width as target training data. Simultaneously, to prevent training set degradation and increase the model's exploration capability, the electronic device 100 also selects a preset number of training data with the highest difficulty from the training data whose difficulty value exceeds the capability boundary as difficult training data. Finally, the electronic device 100 merges the boundary samples and difficult samples to obtain the training data used in the next round of training. For details on the capability boundary, target training data, and the specific calculation methods in the hierarchical heterogeneous reward stage, please refer to the above. Figure 4 The relevant descriptions in the document are not repeated here.
[0232] In the retraining phase, the electronic device 100 uses the updated training data obtained from the sample resampling phase as the input data for the next round of iterative training, and re-executes the policy optimization phase. This process is repeated until the model converges or reaches the preset number of training rounds.
[0233] It is understandable that, through the aforementioned closed-loop logic, the composition of training data dynamically "self-evolves" with the model's capabilities: as the model's capabilities improve, previously challenging training data may become of moderate difficulty and be included in the training, ensuring that the model always learns around its current capability boundaries. This avoids the problem of insufficient high-value samples in the later stages of training with a fixed training set, thus improving overall training efficiency and sample utilization. Simultaneously, by progressively expanding with challenging samples, the model can continuously challenge itself with increasingly difficult tasks, gradually expanding its capability boundaries and achieving a balance between training stability and exploratory nature.
[0234] Understandable. Figure 7The method for determining the capability boundary in the "Model Current Capability Assessment" stage of the closed loop shown differs significantly from the fixed threshold method. To illustrate this difference more intuitively, the following section combines... Figure 8 A comparative explanation will be provided.
[0235] For example, Figure 8 A schematic diagram comparing a fixed threshold selection range with the adaptive selection range of this application is shown according to an embodiment of the present application.
[0236] like Figure 8 As shown in the figure, the horizontal axis represents the training rounds (from round 0 to round 14), and the vertical axis represents the sample difficulty value (ranging from 0 to 1). The curves and labeled areas in the figure are used to compare the differences between the fixed threshold selection range and the adaptive selection range of this application.
[0237] The fixed threshold selection range is represented by a horizontal dashed line in the graph. This threshold remains constant at approximately 0.6 throughout the entire training process. As can be seen from the graph, regardless of the training epoch, the fixed threshold line remains at the same level, meaning that samples with a difficulty value below 0.6 are always selected as training data. This static selection method can filter out some moderately difficult samples in the early stages of training (e.g., epochs 0 to 4), but in the later stages (e.g., after epoch 10), due to the significant improvement in model capabilities, many samples with a difficulty value below 0.6 become too easy for the model, resulting in low training value and wasted computational resources. Simultaneously, in the early stages of training, some samples with higher difficulty values cannot be selected because they exceed the fixed threshold of 0.6, preventing the model from encountering more challenging samples in the early stages.
[0238] The adaptive selection range in this application is represented by a slanted solid line in the figure, which dynamically shifts downwards with the increase of training epochs. In the early stages of training (e.g., epochs 0 to 2), the adaptive selection range is relatively high, around 0.4 to 0.8, mainly focusing on relatively difficult samples. This reflects the limited model capability in the early stages of training, but selecting moderately difficult samples can accelerate the initial learning of the model. As the number of training epochs increases (epochs 3 to 8), the adaptive selection range gradually shifts downwards to around 0.4 to 0.2, indicating that as the model capability gradually improves, the originally difficult samples are gradually mastered, the capability boundary shifts downwards, and the model begins to focus on a new medium-difficulty region. In the later stages of training (epochs 9 to 14), the adaptive selection range further shifts downwards to around 0.1 to 0. At this point, the model capability has been greatly improved, the originally high-difficulty samples have become medium-difficulty, and the adaptive selection range is adjusted to the new capability boundary region.
[0239] The image shows three marked areas for comparing the filtering results of the two methods:
[0240] The samples in the "both selected" region (top left area of the figure) were selected as training data by both the fixed threshold method and the adaptive method of this application. These samples are typically of moderate difficulty and will be included in training regardless of the strategy used, representing a common training sample set for both methods. As can be seen from the figure, the samples in this region are mainly distributed in the early stages of the training rounds (rounds 1 to 4), indicating a significant overlap in the selection range of the two methods in the early stages of training.
[0241] The samples in the "Samples Selected by Fixed Threshold" area (lower right corner of the figure) were selected only by the fixed threshold method, while the adaptive method of this application did not select them. These samples had low difficulty values (0.15 to 0.00) in the later stages of training (rounds 9 to 14), and were considered simple samples for a model whose capabilities had been significantly improved. The adaptive method of this application dynamically adjusts the selection range to exclude these simple samples from the training set, thereby concentrating training resources on high-value samples with greater learning potential.
[0242] The samples in the "Not Selected" area (the bottom right area in the image) were not selected in either method. These samples had extremely low difficulty values (close to 0 or even negative) in the later stages of training (rounds 11 to 14), making them too simple for both methods and not valuable for training.
[0243] It is worth noting that in the 14th round, the adaptive selection range of this application dropped to around 0.05, indicating that the model's capabilities have been sufficiently improved and that the model has mastered almost all positive difficulty samples. New and more difficult samples need to be introduced to continue to drive the model's progress.
[0244] It is understandable that the fixed threshold method uses a horizontally unchanged threshold line for selection throughout the training process, which cannot adapt to the dynamic changes in model capabilities: a high threshold in the early stages of training results in some medium-difficulty samples not being selected, while a low threshold in the later stages of training results in a large number of simple samples being included in the training set, causing a waste of resources. In contrast, the adaptive method in this application dynamically adjusts the selection range as the training rounds increase, always focusing training resources on high-value sample regions near the current capability boundary of the model, thereby maintaining high sample utilization and training efficiency throughout the entire training process.
[0245] It is understood that, unless otherwise specified, in the embodiments of this application, "sample", "sample data", "training data", "training sample" and "training sample data" all refer to the input data used to train the model, and this application does not limit them.
[0246] In conjunction with the above Figures 5 to 8 After introducing the hierarchical gating mechanism, parameter semantic reward mechanism, and training set self-evolution closed-loop logic, the model training method of this application will be explained from the overall architecture level below.
[0247] For example, Figure 9 An architecture diagram of a model training method based on capability boundary awareness and hierarchical rewards is shown according to an embodiment of this application.
[0248] like Figure 9 As shown, the architecture is divided into two main modules: the left module is the "hierarchical gating and heterogeneous parameter semantic reward" module, and the right module is the "reward-driven training set self-evolution and reinforcement learning closed loop" module. Both modules use a unified tool to call the reward signal. Interact with each other.
[0249] The "Hierarchical Gating and Heterogeneous Parameter Semantic Reward" module is the reward calculation part. It includes four levels in a bottom-up order: tool name matching, parameter key matching, parameter value reward, and format validity reward. Finally, it outputs a unified tool call reward. At the tool name matching level, it checks whether the tool name in the predicted invocation command matches the tool name in the standard invocation command. If they don't match (are not exactly the same), the invocation is deemed invalid, and the reward is 0. If they match (are exactly the same), it proceeds to the parameter key matching level. At the parameter key matching level, only if the tool name is correct, it continues to evaluate whether the parameter keys in the predicted invocation command match the parameter keys in the standard invocation command. If they match (are exactly the same), it proceeds to the parameter value reward level. At the parameter value reward level, it distinguishes between strict parameters and semantic parameters: for strict parameters, a perfect match reward is given, meaning a reward is given when the parameter value in the predicted invocation command matches the parameter value in the standard invocation command exactly; otherwise, the reward is 0. For semantic parameters, a semantic similarity reward is given, meaning the semantic similarity between the parameter value in the predicted invocation command and the parameter value in the standard invocation command is calculated, and a continuous reward is given based on the similarity.
[0250] Furthermore, the format validity reward is used as an independent dimension to evaluate the format compliance of the predicted invocation command, determining whether it conforms to the preset valid tool invocation format. Finally, the rewards from all the above levels are aggregated to obtain a unified tool invocation reward. .in, For specific determination methods, please refer to the above. Figure 4 In The relevant descriptions are not repeated here.
[0251] The "Reward-Driven Training Set Self-Evolution and Reinforcement Learning Closed Loop" module is the training optimization part, using the current i-th round of training data. Starting with step 1, a complete reinforcement learning closed-loop logic is formed by following the order of steps 1 to 8. Specifically:
[0252] Step 2 generates multiple candidate response trajectories for the GRPO strategy, that is, using the current model to generate multiple candidate response trajectories for each training data.
[0253] Step 3 is the hierarchical reward evaluation, which uses the hierarchical gating and heterogeneous parameter semantic reward mechanism of the "hierarchical gating and heterogeneous parameter semantic reward" module to calculate the reward value for each of the multiple trajectories generated in step 2.
[0254] Step 4 is to calculate the relative group advantage, that is, to normalize the intra-group reward of each training data to eliminate the difference in reward scale between different training data, and obtain the relative group advantage of each trajectory.
[0255] Step 5 is to update the strategy model, that is, based on the group relative advantage calculated in step 4, the model parameters are updated using the pruning strategy objective function, so that high reward responses are strengthened and low reward responses are suppressed.
[0256] To prevent the policy from shifting excessively during reinforcement learning, KL divergence constraints are also introduced.
[0257] Step 6 is online difficulty estimation, which involves regenerating trajectories on the training data using the updated policy model and calculating reward values. Based on the reward values, the success score and difficulty value of each training data point are determined.
[0258] Step 7 is to model the capability boundary, that is, to determine the capability boundary of the current model based on the difficulty values of all training data, and to construct the capability boundary band based on the capability boundary.
[0259] Step 8 involves resampling the target training data and the difficult training data. Specifically, target training data is selected from the training data whose difficulty value is less than or equal to the preset boundary width, and a preset number of training data with the highest difficulty value are selected from the samples with difficulty values higher than the capability boundary as difficult training data. The two are then combined as the training data used in the next round of training. Then, the process returns to step 1 and enters the next iteration.
[0260] Furthermore, the capability boundary evolves dynamically within the closed-loop logic of reinforcement learning. That is, as the number of training rounds increases, the model's capability continuously improves, and the capability boundary continues to migrate towards higher difficulty levels.
[0261] Understandable. Figure 9The overall architecture shown fully illustrates the closed-loop logic of reinforcement learning, including reward evaluation, policy optimization, difficulty update, and sample resampling. The prediction call instructions generated by the model first pass through the "hierarchical gating and heterogeneous parameter semantic reward" module to obtain a reward value. This reward value is used on the one hand to update the model parameters in the GRPO policy optimization (steps 2 to 5) of the "reward-driven training set self-evolution and reinforcement learning closed loop" module, and on the other hand to generate the difficulty value of the training data through online difficulty estimation (step 6). Then, the next round of training data is reconstructed through capability boundary modeling (step 7) and sample resampling (step 8), realizing "reward-driven self-evolution". That is, the training data evolves dynamically with the model's capability, and the training always revolves around the current capability boundary of the model.
[0262] For example, Table 1 shows a comparison of the tool call test results of a model training method provided in this application embodiment and other existing methods, according to an embodiment of this application.
[0263] Table 1. Comparison of tool call test results between the model training method provided in this application embodiment and other existing methods.
[0264] As shown in Table 1, the model training method provided in this application and other existing methods have scores (or accuracy) on the API-Bank and BFCL V3 benchmark tests, two mainstream tools.
[0265] The API-Bank test includes evaluations across four dimensions: Overall Score, Basic API Calling Capability (API-Bank L1), API Retrieval and Calling Capability (API-Bank L2), and Complex Task Planning and Calling Capability (API-Bank L3).
[0266] The BFCL V3 test includes evaluation in five dimensions: accuracy of abstract syntax tree matching in non-real-time environment (BFCL V3 Non-LiveAST), execution accuracy in multi-tool parallel invocation scenario (BFCL V3 Multi-Live), function call capability in multi-turn dialogue (BFCL V3 Multi-Turn), relevance judgment capability (BFCL V3 Relevance), and irrelevance judgment capability (BFCL V3 Irrelevant).
[0267] The models compared include ToolACE-8B, GPT-4-turbo-2024-04-09, GPT-4-multi-2024-07-18, GPT-3.5-turbo-0125, LLaMA-3.1-8B-Instruct, and multiple variants of Qwen2.5-7B-Instruct (Raw pedestal model, SFT400 supervised fine-tuning model, SFT4k supervised fine-tuning model, SFT400+GRPO reinforcement learning model), Tool-N1, ToolRL, ToolSample, and the target model trained by the method of this application.
[0268] As shown in Table 1, for API-Bank Overall, the proposed method scores 72.2 points, higher than all comparable methods. For API-Bank L1, the proposed method scores 76.4 points, higher than all comparable methods. For API-Bank L2, both the proposed method and ToolSample score 65.7 points, significantly higher than ToolRL's 61.2 points and ToolACE-8B's 47.4 points. For API-Bank L3, the proposed method scores 62.6 points, far exceeding ToolSample's 39.7 points and ToolRL's 35.1 points, demonstrating a significant improvement. This indicates that the proposed method has a clear advantage in complex scenarios requiring multi-step planning and invocation.
[0269] In the BFCL V3 test, the proposed method performed excellently across five dimensions: Non-Live AST (62.9 points), Multi-Live (88.6 points), Multi-Turn (77.4 points), Relevance (83.3 points), and Irrelevant (80.4 points). For example, for Multi-Live, the proposed method scored 88.6 points, significantly outperforming ToolSample's 86.4 points, ToolRL's 86.7 points, and Qwen2.5-7B-Instruct (SFT400+GRPO)'s 80.7 points, demonstrating a clear advantage over similarly sized open-source models.
[0270] It is understandable that the above experimental results demonstrate that the model training method provided in this application achieves excellent performance across multiple dimensions in the API-Bank and BFCL V3 benchmark tests, two mainstream tools, through its hierarchical gating reward mechanism, asymmetric semantic reward mechanism based on parametric semantic attributes, and reward-driven training set self-evolution mechanism. In particular, the method shows the greatest improvement in API-Bank L3 complex task planning capabilities, verifying the effectiveness of this application's method in enhancing the model's complex inference and multi-step planning capabilities.
[0271] Meanwhile, in practical application scenarios such as multi-live parallel invocation of multiple tools and multi-turn dialogue, the method in this application has also achieved a performance level comparable to top-tier commercial models, demonstrating good practical value.
[0272] It is understandable that the above experimental results show that the model training method provided in this application has achieved significant performance improvements in multiple dimensions of the API-Bank and BFCL V3 benchmark tests, two mainstream tools, through the hierarchical gating reward mechanism, the asymmetric semantic reward mechanism based on parameter semantic attributes, and the reward-driven training set self-evolution mechanism. In particular, it has performed outstandingly in the API-Bank L3 complex task planning capability and the BFCL V3 Multi-Live multi-tool parallel calling capability, which verifies the effectiveness and generalization ability of the method in this application.
[0273] It is understood that the target model obtained by the model training method provided in this application embodiment can be applied to various scenarios that require structured tool calls or function call generation. This target model can be deployed in electronic device 100 to receive natural language commands input by the user, understand the user's intent, and automatically plan and call external tools or services to complete the user's commands.
[0274] For example, this target model can be applied to the following scenarios: intelligent agent multi-task tool invocation, i.e., the model can autonomously plan and invoke multiple tools to complete multi-step tasks in complex task scenarios; intelligent assistant automatic service invocation, i.e., the model can understand the user's natural language requests and automatically invoke the corresponding system capabilities or third-party services; multi-step application programming interface (API) call generation, i.e., the model can generate a series of API call sequences according to user needs and coordinate their execution; code generation and function call optimization, i.e., the model can generate structured function call instructions that conform to specific syntax specifications. In all of the above scenarios, the structured call instructions output by the target model can be correctly parsed and executed by the electronic device 100, thereby accurately completing the user's instructions.
[0275] It is understood that the target model can also be deployed in other electronic devices besides electronic device 100, and this application does not impose any restrictions on this.
[0276] It is understood that the models in the embodiments of this application can be large language models based on the Transformer architecture, such as the Qwen series models, LLaMA series models, GPT series models, etc.; or they can be language models without the Transformer architecture, such as state space models (SSM). The embodiments of this application do not limit the specific architecture and number of parameters of the model. As long as the model needs to be trained to have the ability to generate structured tool calls, the model training method provided in the embodiments of this application can be used.
[0277] It is understood that electronic device 100 can be a cloud server. The cloud server can be a physical server or virtual server deployed in a public cloud, private cloud, or hybrid cloud environment, such as an elastic cloud server (ECS), a bare metal server, or a cloud virtual machine.
[0278] Electronic device 100 can also be a local server. This local server can be a physical server deployed in an enterprise's internal data center or a private cloud server.
[0279] Electronic device 100 can also be a terminal device. The terminal device can be a personal computer (PC), laptop, workstation, tablet, or smartphone used by the developer.
[0280] Electronic device 100 can also be a combination of the above forms, such as a distributed system composed of cloud servers and terminal devices, in which different functional modules can be deployed on different electronic devices and work together through network communication.
[0281] It is understood that the specific type of electronic device 100 is not limited in the embodiments of this application, as long as it can execute the model training method provided in the embodiments of this application.
[0282] To better understand the technical solutions of the embodiments of this application, the devices involved in the embodiments of this application will be described below with reference to the accompanying drawings.
[0283] For example, Figure 10 A schematic diagram of the hardware structure of an electronic device 100 is shown according to an embodiment of this application.
[0284] like Figure 10 As shown, the electronic device 100 includes one or more (only one is shown in the figure) processors 110, memory 120, communication interface 130, and bus 140. The processors 110, memory 120, and communication interface 130 are interconnected via the bus 140.
[0285] The processor 110 may include one or more processing units, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), an application-specific integrated circuit, etc., for executing related programs to achieve the functions required by the modules in the compilation apparatus of this application embodiment, or to execute the model training method of this application method embodiment.
[0286] The memory 120 may include one or more memories for storing data or one or more applications. The memory may be read-only memory (ROM), static storage device, dynamic storage device, random access memory (RAM), high-speed random access memory, double data rate synchronous dynamic RAM (DDR), high bandwidth memory (HBM), or non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0287] In some embodiments, if the processor is a CPU, the corresponding memory 120 is main memory. In some embodiments, if the processor 110 is a GPU, the corresponding memory 120 can be video memory.
[0288] The processor 110 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the data processing method of this application can be completed by software instructions in the processor 110. The aforementioned processor 110 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory 120. The processor 110 reads the information in the memory 120 and, in conjunction with its hardware, performs the functions required by the units included in the compilation apparatus of this application embodiment, or executes the model training method of the method embodiment of this application.
[0289] Communication interface 130 is used to enable communication between electronic device 100 and other devices or communication networks. Communication interface 130 may include wired or wireless communication interfaces, so that electronic device 100 can access the Internet via wired or wireless means, and obtain data from or send data to other devices based on the Internet.
[0290] Bus 140 is used to connect processor 110, memory 120, communication interface 130 and other possible modules or circuits.
[0291] In some embodiments, the electronic device 100 can be used to execute the model training method corresponding to the method embodiments described above. To avoid repetition, it will not be described again here.
[0292] It should be understood that Figure 10 The structure of the electronic device 100 shown is only an example. In other embodiments, the electronic device 100 may include more or fewer modules, which is not limited here.
[0293] This application also provides a computer-readable storage medium storing at least one computer program instruction, at least one program segment, code set, or instruction set, which is loaded and executed by a processing circuit to implement the model training method provided in this application.
[0294] This application also provides a program product that includes instructions that, when executed by an electronic device, enable the electronic device to implement the model training method provided in this application.
[0295] This application also provides a chip, including a processor coupled to a memory, for executing computer programs or instructions stored in the memory, so that the chip implements the model training method provided in this application.
[0296] Embodiments of this application can be implemented as computer programs or program code that execute on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0297] Various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or combinations of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0298] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.
[0299] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0300] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, optical discs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0301] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0302] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0303] It should be noted that in the examples and description of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0304] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.
Claims
1. A model training method, characterized in that, Applied to electronic devices, the method includes: Acquire training data, wherein the training data includes: input information and standard calling instructions corresponding to the input information; The model is trained iteratively for N rounds based on the training data to obtain the target model, where N is an integer greater than 0; where, The i-th round of training in the N rounds of iterative training includes: The current model processes the input information in the training data to generate a prediction invocation instruction. Both the prediction invocation instruction and the standard invocation instruction include M layers of elements, and the level of the M layers of elements decreases sequentially. M is an integer greater than or equal to 2. The predicted call instruction is compared with the standard call instruction layer by layer to obtain the reward value of the predicted call instruction. In the process of layer-by-layer comparison, if the comparison of the element at the j-th layer fails, the reward value of the element at the j-th layer and the reward value of the element at the lower level than the element at the j-th layer are set to 0, where j is a positive integer less than or equal to M. Update the model parameters based on the reward value; Where i is a positive integer less than or equal to N.
2. The method according to claim 1, characterized in that, The M-layer elements include: tool name, parameter key, and parameter value; wherein... The tool name is at a higher level than the parameter key, and the parameter key is at a higher level than the parameter value.
3. The method according to claim 2, characterized in that, The step-by-step comparison of the predicted invocation instruction with the standard invocation instruction includes: The tool name in the predicted invocation instruction is compared with the tool name in the standard invocation instruction. If the tool name in the predicted invocation instruction is different from the tool name in the standard invocation instruction, the tool name comparison is determined to be unsuccessful, and the reward value of the tool name, the reward value of the parameter key, and the reward value of the parameter value are set to 0.
4. The method according to claim 3, characterized in that, The step-by-step comparison of the predicted invocation instruction with the standard invocation instruction further includes: If the tool name in the predicted call instruction is the same as the tool name in the standard call instruction, the tool name comparison is deemed successful, the reward value of the tool name is set to the first preset score, and the parameter key in the predicted call instruction is compared with the parameter key in the standard call instruction. If the parameter key in the predicted invocation instruction is different from the parameter key in the standard invocation instruction, the parameter key comparison is determined to have failed, and the reward value of the parameter key and the reward value of the parameter value are set to 0.
5. The method according to claim 4, characterized in that, The step-by-step comparison of the predicted invocation instruction with the standard invocation instruction further includes: If the parameter key in the predicted invocation instruction is the same as the parameter key in the standard invocation instruction, the parameter key comparison is determined to be successful. The reward value of the parameter key is set to the first preset score, and the semantic constraint type of the parameter value is determined. The parameter value in the predicted invocation instruction is compared with the parameter value in the standard invocation instruction based on the semantic constraint type of the parameter value. The semantic constraint type includes strict parameters and semantic parameters. If the semantic constraint type is a strict parameter, and the parameter value in the predicted invocation instruction is exactly the same as the parameter value in the standard invocation instruction, the reward value of the parameter value is set to the third preset score; If the semantic constraint type is a strict parameter, and the parameter value in the predicted invocation instruction is not exactly the same as the parameter value in the standard invocation instruction, the reward value of the parameter value is set to 0; If the semantic constraint type is a semantic parameter, calculate the semantic similarity between the parameter value in the predicted invocation instruction and the parameter value in the standard invocation instruction, and determine the reward value of the parameter value based on the semantic similarity, wherein the magnitude of the reward value of the parameter value is positively correlated with the magnitude of the semantic similarity.
6. The method according to claim 5, characterized in that, The strict type parameters include at least one of the following: identification parameters and numbering parameters; The semantic parameters include at least one of the following: text description parameters and keyword parameters.
7. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the reward value of the predicted invocation instruction includes: The reward values of the M-layer elements are weighted and summed to obtain the correctness reward value of the predicted invocation instruction. Obtain the format validity reward value of the predicted invocation instruction, which is used to indicate whether the format of the predicted invocation instruction conforms to a preset valid tool invocation format; The reward value for the predicted call instruction is determined based on the correctness reward value and the format validity reward value.
8. The method according to claim 7, characterized in that, The step of obtaining the format validity reward value of the predicted invocation instruction includes: If the format of the predicted invocation instruction conforms to the preset legal tool invocation format, the format legality reward value is set to a preset positive value; If the format of the predicted invocation instruction does not conform to the legal tool invocation format, the format legality reward value is set to a preset negative value.
9. The method according to claim 1, characterized in that, The number of standard invocation instructions and the number of predicted invocation instructions are both greater than 1; and... The step-by-step comparison of the predicted invocation instruction with the standard invocation instruction includes: The predicted invocation instruction is matched with each standard invocation instruction in a preset order; among which... In the process of matching the corresponding predicted call instruction for each standard call instruction, the predicted call instruction that is the same as the first-level element of the standard call instruction is selected from the predicted call instructions that have never been matched.
10. The method according to claim 1, characterized in that, The number of training data is greater than 1; and, The method further includes: Based on the reward value corresponding to each training data in the i-th round of training, the difficulty value of each training data is determined. Based on the difficulty value of each training data point, determine the training data used in the (i+1)th round of training.
11. The method according to claim 10, characterized in that, The step of determining the difficulty value of each training data point based on the reward value corresponding to each training data point during the i-th round of training includes: Based on the reward value corresponding to each training data point in the i-th training round, determine the success score of each training data point. Based on the success scores of each training data point, the difficulty value of each training data point is determined, wherein the higher the success score, the lower the difficulty value.
12. The method according to claim 10 or 11, characterized in that, The determination of the training data used in the (i+1)th round of training based on the difficulty value of each training data point includes: Determine the capability boundaries of the current model based on the difficulty values of each training data point; Select target training data from the training data whose difficulty value is less than or equal to the capability boundary width. Based on the target training data, determine the training data used in the (i+1)th round of training.
13. The method according to claim 12, characterized in that, The step of determining the training data used in the (i+1)th round of training based on the target training data includes: From the training data, select a preset number of training data with the highest difficulty value among the training data whose difficulty value is greater than the ability boundary, and use them as difficult training data; The target training data and the difficult training data are used as the training data for the (i+1)th round of training.
14. The method according to claim 10, characterized in that, The step of determining the difficulty value of each training data point based on the reward value corresponding to each training data point in the i-th round of training also includes: Obtain the reward value for each training data point during the (i-1)th training round; The reward values corresponding to each training data point in the (i-1)th training round are normalized to obtain the normalized reward values of each training data point in the (i-1)th training round. Based on the normalized reward value of each training data in the (i-1)th training round and the cumulative reward value of each training data in the (i-2)th round, the updated cumulative reward value of each training data in the (i-1)th round is obtained. Based on the updated cumulative reward value of each training data point in round i-1, the difficulty value of each training data point in round i is determined; where... i is an integer greater than 2.
15. An electronic device, characterized in that, include: A memory for storing instructions executed by one or more processors of the electronic device; as well as The processor is one of the processors of the electronic device, used to execute instructions stored in the memory to implement the method of any one of claims 1 to 14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on the device, cause the device to perform the method of any one of claims 1 to 14.
17. A computer program product, characterized in that, The computer program product includes instructions that, when executed on the device, cause the device to perform the method of any one of claims 1 to 14.