Reinforcement learning training method and device for multi-modal model
By introducing tool usage tips and reference model inference content into the training of multimodal models, the reinforcement learning training process of multimodal models is optimized, solving the problem of low training efficiency and realizing a rapid improvement in the model's tool calling capability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI XIYU JIZHI TECH CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-06-02
Smart Images

Figure CN122133742A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a reinforcement learning training method and apparatus for a multimodal model. Background Technology
[0002] Multimodal models can solve problems that cannot be solved without using tools, thus expanding the model's capabilities. Currently, the reinforcement learning training process for enabling tool invocation in multimodal models generally involves: constructing multimodal training data, where each training data point includes a prompt for the multimodal model to invoke at least one tool; the multimodal model inferring from each training data point and outputting a prediction result; comparing the model's prediction result with the standard result to determine the reward value; and then adjusting the parameters of the multimodal model based on the reward value until reinforcement learning training is completed, resulting in a multimodal model with tool invocation capabilities.
[0003] However, the above training method, although it tells the model that it can call at least one tool in each training data, the model is initially unable to know when to call it and how to use the tool to solve the problem correctly. Therefore, in the early stage of traditional reinforcement learning training, the reward ratio is extremely low. The model needs to conduct extensive and long-term exploration to acquire the ability to call the tool, and then it needs to undergo a long period of training to acquire the ability to use the tool to solve the problem correctly. The whole training process has the problem of low model training efficiency. Summary of the Invention
[0004] This invention provides a reinforcement learning training method and apparatus for multimodal models to solve the problem of low training efficiency in the entire training process, which aims to enable the model to correctly solve problems using tools.
[0005] According to one aspect of the present invention, a reinforcement learning training method for a multimodal model is provided, the method comprising: Acquire multiple sets of multimodal training data; wherein each set of multimodal training data includes tool usage prompt text content and image content; The multimodal training data is input into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes the prediction result and the first inference content; At least one reference model is obtained, and the multimodal training data is input into the reference model to obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has the function of calling tools; Based on the predicted data and the second inference content, the target reward data corresponding to the predicted data is determined, and the parameters of the model to be trained are updated based on the target reward data to obtain the target model; wherein, the target reward data includes information from two dimensions: the prediction result and the first inference content.
[0006] According to another aspect of the present invention, a reinforcement learning training apparatus for a multimodal model is provided, the apparatus comprising: A multimodal training data acquisition module is used to acquire multiple pieces of multimodal training data; wherein each piece of training data includes tool usage prompt text content and image content; The prediction module is used to input the multimodal training data into the model to be trained and obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes the prediction result and the first inference content; The data acquisition module is used to acquire at least one reference model, input the multimodal training data into the reference model, and obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has tool calling function; An update module is used to determine the target reward data corresponding to the prediction data based on the prediction data and the second inference content, and update the parameters of the model to be trained based on the target reward data to obtain the target model; wherein, the target reward data includes information from two dimensions: prediction result and the first inference content.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the reinforcement learning training method for the multimodal model according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the reinforcement learning training method for a multimodal model according to any embodiment of the present invention.
[0009] The technical solution of this invention involves acquiring multiple sets of multimodal training data. Each set of multimodal training data includes tool usage prompt text and image content. The tool usage prompt text can guide the model to determine whether to use a tool, and when the model determines that a tool needs to be used, it can output the correct tool call format. This allows for more accurate acquisition of at least one prediction data corresponding to each set of multimodal training data after inputting the multimodal training data into the model. Simultaneously, the multimodal training data is input into a reference model to obtain the second inference content corresponding to each set of multimodal training data. This process only requires outputting the model's thinking process and does not require outputting the result, thereby saving computational power, shortening inference time, and improving training efficiency. Because the prediction data includes the prediction result and the first inference content, the target reward data corresponding to the prediction data determined based on the prediction data and the second inference content can better reflect the correctness of the model in the intermediate inference process and the correctness of the prediction result. Therefore, the parameters of the model to be trained are updated based on the target reward data to obtain the target model. This accelerates the reinforcement learning training of the model's ability to call tools, solving the problem of low model training efficiency in the entire training process of training the model to have the ability to correctly solve problems using tools.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a reinforcement learning training method for a multimodal model according to an embodiment of the present invention; Figure 2 This is a flowchart of another reinforcement learning training method for a multimodal model provided according to an embodiment of the present invention; Figure 3 This is a flowchart of another reinforcement learning training method for a multimodal model provided according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of another reinforcement learning training device for a multimodal model according to an embodiment of the present invention; Figure 5This is a schematic diagram of the structure of an electronic device that implements the reinforcement learning training method for a multimodal model according to an embodiment of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0015] Figure 1 This is a flowchart illustrating a reinforcement learning training method for a multimodal model according to an embodiment of the present invention. This embodiment is applicable to situations where reinforcement learning training is performed on a multimodal model to enable the model to correctly invoke tools. This method can be executed by a reinforcement learning training device for the multimodal model, which can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities. Figure 1 As shown, the reinforcement learning training method for the multimodal model of the present invention may include: S110. Obtain multiple multimodal training data sets; each multimodal training data set includes tool usage prompt text content and image content.
[0016] Multimodal training data, which can be data that integrates multiple modalities such as text, images, audio, and video and achieves semantic alignment between modalities, is the core foundation for training large multimodal models and directly determines the model's understanding ability, generation quality, and generalization ability. Tool usage prompts describe the tools that the model to be trained can use, including the tool's name / identifier, function, and usage instructions. This helps the model determine whether to use a tool and outputs the correct tool call format when a tool is needed.
[0017] Specifically, obtaining multiple pieces of multimodal training data may include: obtaining multiple pieces of multimodal training data from a preset multimodal dataset, each piece of multimodal training data containing information from at least two different modalities, which may include text and at least one of images, audio, and video.
[0018] S120. Input the multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes the prediction result and the first inference content.
[0019] The prediction result can be understood as the final output of the model to be trained after performing forward inference calculations on the input multimodal training data. The first inference content can be understood as the content of the inference process of the model to be trained performing forward inference on the input multimodal training data.
[0020] Specifically, if the model to be trained executes one sampling inference logic during one inference training process, then based on each multimodal training data and the model to be trained, one prediction data corresponding to each multimodal training data is obtained; if the model to be trained executes parallel sampling inference logic for a preset number of times during one inference training process, then based on each multimodal training data and the model to be trained, a preset number of prediction data corresponding to each multimodal training data is obtained; the preset number of prediction data is equal to the preset number of times, and the preset number of times is greater than or equal to 2.
[0021] S130. Obtain at least one reference model, input multimodal training data into the reference model, and obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has the function of calling tools.
[0022] The second inference content can be understood as the content of the inference process during which the reference model performs forward inference on the input multimodal training data. For example, the second inference content may include the inference content regarding whether the reference model calls a tool and / or the specific tool called.
[0023] Specifically, multimodal training data is input into a reference model. The reference model performs feature extraction and semantic understanding on the input multimodal training data, and analyzes and judges the task intent, execution conditions, and required resources corresponding to each piece of multimodal training data based on its internal inference logic and tool invocation decision rules, generating second inference content corresponding to that multimodal training data. The second inference content may include whether the reference model needs to invoke external tools, and if so, the specific tool type or tool identifier to be invoked.
[0024] In other words, the reference model of this invention does not need to perform a complete inference process to output the final result. It only needs to infer whether the reference model calls the tool and / or the content of the specific tool called, and output the corresponding inference content, i.e. the second inference content, so as to save computing power consumption and shorten the inference time of the model, thereby improving the training efficiency of the reinforcement learning training of the entire multimodal model.
[0025] S140. Based on the prediction data and the second inference content, determine the target reward data corresponding to the prediction data, update the parameters of the model to be trained based on the target reward data, and obtain the target model; wherein, the target reward data includes information from two dimensions: prediction results and the first inference content.
[0026] Specifically, the target reward data corresponding to the predicted data can include: if the model to be trained uses a parallel sampling method to output multiple predicted data, then for each predicted data corresponding to each multimodal training data, a corresponding target reward data can be determined, that is, each multimodal training data corresponds to multiple target reward data; if the model to be trained uses a single sampling method to output one predicted data, then for each predicted data corresponding to each multimodal training data, only one target reward data is determined, that is, each multimodal training data corresponds to one target reward data.
[0027] Specifically, determining the target reward data corresponding to the predicted data based on the predicted data and the second inference content can include: determining the first reward data for the predicted data based on the prediction results in the predicted data and the reference results of the training data corresponding to the predicted data; the first reward data can be data that rewards the correctness of the prediction results of the model to be trained; determining the second reward data for the predicted data based on the first inference content in the predicted data and the second inference content of the reference model; the second reward data can be data that rewards the correctness of the first inference content of the model to be trained; determining the target reward data corresponding to the predicted data based on the first reward data and the second reward data, and then updating the parameters of the model to be trained based on the target reward data to obtain the target model.
[0028] Furthermore, based on the first inference content in the predicted data and the second inference content of the reference model, determining the second reward data of the predicted data may include: determining first reference information based on the first inference content in the predicted data; the first reference information may include first tool information on whether the model to be trained calls a tool and / or the tool called by the model to be trained; determining second reference information based on the second inference content; the second reference information may include second tool information on whether the reference model calls a tool and / or the tool called by the reference model; and determining the second reward data of the predicted data based on the degree of matching between the first reference information and the second reference information.
[0029] Accordingly, based on the degree of matching between the first reference information and the second reference information, the second reward data for the predicted data can be determined as follows: there is a preset correlation between the degree of matching between the first reference information and the second reference information and the second reward data, and the higher the degree of matching, the higher the second reward data; then after obtaining the degree of matching between the first reference information and the second reference information, the second reward data of the corresponding predicted data can be matched from the preset correlation.
[0030] Furthermore, determining the target reward data corresponding to the predicted data based on the first reward data and the second reward data may include: setting corresponding first reward weights and second reward weights for the first reward data and the second reward data respectively, and determining the target reward data corresponding to the predicted data by weighted summation of the first reward data, the first reward weight, the second reward data, and the second reward weight.
[0031] Optionally, updating the parameters of the model to be trained based on the target reward data to obtain the target model may include: determining the gradient value based on the prediction data and target reward data corresponding to each multimodal training data, updating the parameters of the model to be trained based on the gradient value, and obtaining the target model.
[0032] Furthermore, determining the gradient value based on the prediction data and target reward data corresponding to each multimodal training data may also include: if each multimodal training data corresponds to multiple prediction data and multiple target reward data, then based on each prediction data and corresponding target reward data corresponding to each multimodal training data, determine the normalized objective function value corresponding to each multimodal training data; and determine the gradient value based on the normalized objective function value.
[0033] The technical solution of this invention involves acquiring multiple sets of multimodal training data. Each set of multimodal training data includes tool usage prompt text and image content. The tool usage prompt text can guide the model to determine whether to use a tool, and when the model determines that a tool needs to be used, it can output the correct tool call format. This allows for more accurate acquisition of at least one prediction data corresponding to each set of multimodal training data after inputting the multimodal training data into the model. Simultaneously, the multimodal training data is input into a reference model to obtain the second inference content corresponding to each set of multimodal training data. This process only requires outputting the model's thinking process and does not require outputting the result, thereby saving computational power, shortening inference time, and improving training efficiency. Because the prediction data includes the prediction result and the first inference content, the target reward data corresponding to the prediction data determined based on the prediction data and the second inference content can better reflect the correctness of the model in the intermediate inference process and the correctness of the prediction result. Therefore, the parameters of the model to be trained are updated based on the target reward data to obtain the target model. This accelerates the reinforcement learning training of the model's ability to call tools, solving the problem of low model training efficiency in the entire training process of training the model to have the ability to correctly solve problems using tools.
[0034] Figure 2 This is a flowchart of another reinforcement learning training method for a multimodal model provided by an embodiment of the present invention. The technical solution of this embodiment further optimizes the process of S140 in the aforementioned embodiments based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the reinforcement learning training method for the multimodal model of the present invention includes: S210. Obtain multiple sets of multimodal training data; each set of multimodal training data includes tool usage prompt text content and image content.
[0035] S220. Input the multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes the prediction result and the first inference content.
[0036] S230. Obtain at least one reference model, input multimodal training data into the reference model, and obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has the function of calling tools.
[0037] S240. Based on the prediction results in the prediction data and the reference results corresponding to the multimodal training data, determine the first alignment information; wherein, the first alignment information is used to describe whether the prediction results in the prediction data are correct.
[0038] If each piece of multimodal training data corresponds to at least one prediction data, the standard result corresponding to each piece of multimodal training data can be determined as the reference result. Alternatively, if each piece of multimodal training data corresponds to multiple prediction data, the average result of the multiple prediction results of the multiple prediction data can be determined as the reference result.
[0039] Specifically, the prediction results in the prediction data are aligned with the reference results corresponding to the multimodal training data to obtain the aligned prediction results and reference results. The aligned prediction results and reference results are then compared to obtain the comparison results, and the first comparison information is determined based on the comparison results.
[0040] Aligning the prediction results in the prediction data with the reference results corresponding to the multimodal training data can include: aligning the prediction results and reference results in the same dimension, the same format, and the same semantic space, based on the data type.
[0041] Comparing the aligned prediction results with the reference results to obtain the alignment results may include: determining the corresponding alignment rules according to the task type, comparing the aligned prediction results with the reference results based on the alignment rules, and obtaining the alignment results; wherein, the task type may include at least one of classification tasks, text sequence tasks, and object detection tasks; classification tasks can be understood as tasks that directly compare whether the prediction results are consistent with the reference results; text sequence tasks can be understood as tasks that compare whether semantics, keywords, and sequence structure match; object detection tasks can be understood as tasks that determine whether the deviation between the prediction results and the reference results is within the allowable range based on threshold indicators such as intersection-union ratio.
[0042] S250. Based on the first inference content in the prediction data, determine the first analysis result; wherein, the first analysis result includes whether the tool is called during the inference process of the model to be trained.
[0043] The first inference content can be understood as the inference content of the model to be trained during the inference process, which considers whether to call tools and / or call specific tools. This inference content consists of one or more tokens.
[0044] Specifically, text parsing is performed on the first inference content to identify semantic features such as intent keywords, task types, data dependencies, and operation instructions, thus determining the first parsed content. The first parsed content is then matched against preset tool invocation judgment conditions to determine the first matching result, and the first analysis result is determined based on the first matching result. The preset tool invocation judgment conditions can be understood as determining whether the parsed content contains conditions for the training model to invoke tools during inference. These conditions can include at least one of a first judgment condition and a second judgment condition. The first judgment condition can be whether it contains special encoding for invoking tools; the second judgment condition can be whether it requires external computation, retrieval, format conversion, data editing, interface calls, or other operations that the model itself cannot perform.
[0045] The process of determining the first analysis result based on the first matching result may include: if any tool invocation determination condition is met in the first matching result, then the tool that needs to be invoked during the inference process of the model to be trained is determined as the first analysis result; if none of the tool invocation determination conditions are met in the first matching result, then the tool that does not need to be invoked during the inference process of the model to be trained is determined as the first analysis result.
[0046] S260. Based on the second reasoning content, determine the second analysis result; wherein, the second analysis result includes whether the tool was called during the reference model reasoning process.
[0047] The second inference content can be understood as the inference content that considers whether to call tools and / or call specific tools during the inference process of the reference model. This inference content consists of one or more tokens.
[0048] Specifically, the second inference content is parsed to identify semantic features such as intent keywords, task types, data dependencies, and operation instructions, thus determining the second parsed content. The second parsed content is then matched against preset tool invocation criteria to determine the second matching result, and the second analysis result is determined based on this result. The preset tool invocation criteria can be understood as determining whether the parsed content contains conditions for the training model to invoke tools during inference. These criteria can include at least one of a first judgment condition and a second judgment condition. The first judgment condition may be whether it contains special encoding for invoking tools; the second judgment condition may be whether it requires external computation, retrieval, format conversion, data editing, interface calls, or other operations that the model itself cannot perform.
[0049] The process of determining the second analysis result based on the second matching result may include: if any two tool invocation conditions are met in the second matching result, then the tool that needs to be invoked during the inference process of the model to be trained is determined as the second analysis result; if none of the tool invocation conditions are met in the second matching result, then the tool that does not need to be invoked during the inference process of the model to be trained is determined as the second analysis result.
[0050] S270. Based on the first comparison information, the first analysis result, and the second analysis result, determine the target reward data corresponding to the predicted data.
[0051] The first comparison information may include a first comparison parameter and / or a second comparison parameter; the first comparison parameter is used to describe whether the prediction result in the prediction data is correct; the second comparison parameter is used to describe the degree of deviation between the prediction result in the prediction data and the reference result.
[0052] Specifically, a first preset correspondence exists between the first comparison parameter and the reward data; a second preset correspondence exists between the second comparison parameter and the reward data; a first reference reward data is determined based on the first comparison parameter and the first preset correspondence; a second reference reward data is determined based on the second comparison parameter and the second preset correspondence; and a first target reward data is determined based on the first reference reward data and / or the second reference reward data. Further, the first analysis result and the second analysis result are compared, and the second target reward data is determined based on the comparison result. The more similar the first analysis result and the second analysis result, the higher the second target reward data. Finally, based on the first target reward data and the second target reward data, the target reward data corresponding to the predicted data is determined. For example, corresponding first reward weights and second reward weights can be set for the first target reward data and the second target reward data, respectively. The target reward data corresponding to the predicted data is determined by weighted summation of the first target reward data, the first reward weight, the second target reward data, and the second reward weight. Other preset rules can also be used. For example, if both the first and second target reward data are positive, they are added together to obtain the target reward data; if the first target reward data is positive and the second target reward data is zero, the first target reward data is used as the target reward data or the target reward data is set to 0; if the first target reward data is 0 and the second target reward data is positive, the target reward data is set to 0 or set to negative as a penalty. This more accurately incentivizes the training model to correctly call tools and solve problems, thereby improving training efficiency.
[0053] Although this invention prompts the model to call tools if necessary for each multimodal training data point, the model itself decides whether to call tools, especially in the early stages of reinforcement learning training when it lacks the ability to call tools. Therefore, it does not call tools for most multimodal training data that should be addressed. Thus, by introducing a reference model, a second analysis result is provided for each multimodal training data point, indicating whether and / or which tool to call. This rewards the model's decision-making process and results for predicting each multimodal training data point, rather than simply rewarding the prediction results as in existing technologies, which neglect to guide and enhance the model's thought process. Therefore, compared to existing technologies, this invention achieves the same training effect with fewer training steps, improving training efficiency.
[0054] Optionally, in this embodiment of the invention, determining the target reward data corresponding to the predicted data based on the first comparison information, the first analysis result, and the second analysis result may include: acquiring multiple reference models and generating multiple second inference entries; determining the confidence weight of each reference model; the confidence weight is used to reflect the tool-calling performance of the reference model; determining the target second analysis result corresponding to the predicted data based on the second analysis result of each second inference entry and the confidence weight of each reference model; and determining the target reward data corresponding to the predicted data based on the first comparison information, the first analysis result, and the target second analysis result. In this embodiment of the invention, by setting a corresponding confidence weight according to the tool-calling performance of each reference model, deviations in the final target reward data caused by the output error of a single reference model can be avoided.
[0055] Optionally, determining the target reward data corresponding to the predicted data based on the first comparison information, the first analysis result, and the second analysis result may further include: acquiring a reference model and generating multiple second inference contents using a parallel sampling method; performing statistical analysis on each second inference content to determine the second analysis result of each second inference content and the frequency of occurrence of each second analysis result; determining the target second analysis result corresponding to the predicted data based on the frequency of occurrence of each second analysis result; and determining the target reward data corresponding to the predicted data based on the first comparison information, the first analysis result, and the target second analysis result. In this embodiment of the invention, the statistical results obtained from multiple inferences using the reference model are used as the second target analysis results, thereby avoiding deviations in the final target reward data caused by output errors from a single inference.
[0056] In an embodiment of the present invention, optionally, determining the target reward data corresponding to the predicted data based on the first comparison information, the first analysis result, and the second analysis result may include steps A1-A2: Step A1: Based on the first comparison information and the first analysis result, determine the prediction type of each prediction data; the prediction type is one of the following: Type 1, Type 2, Type 3, and Type 4; Type 1 is used to describe the model to be trained calling the tool and the prediction result is correct; Type 2 is used to describe the model to be trained calling the tool and the prediction result is incorrect; Type 3 is used to describe the model to be trained not calling the tool and the prediction result is correct; Type 4 is used to describe the model to be trained not calling the tool and the prediction result is incorrect.
[0057] Specifically, the first analysis result includes whether the training model calls a tool during inference, i.e., the first analysis result can include whether the training model calls a tool during inference or not; the first comparison information is used to describe whether the prediction result in the prediction data is correct, i.e., the first comparison information can include whether the prediction result in the prediction data is correct or incorrect; then, based on the combination of the first comparison information and the first analysis result, the prediction data can be divided into four prediction types: the first type describes the training model calling a tool and the prediction result is correct; the second type describes the training model calling a tool and the prediction result is incorrect; the third type describes the training model not calling a tool and the prediction result is correct; and the fourth type describes the training model not calling a tool and the prediction result is incorrect.
[0058] Step A2: Based on the prediction type and the results of the second analysis, determine the target reward data corresponding to the prediction data.
[0059] Specifically, the second analysis result includes whether the reference model inference process calls tools, i.e., whether the reference model inference process calls tools or not. Based on the second analysis result and the prediction type, different combination relationships between the reference model inference process calling tools or calling tools during the reference model inference process and each prediction type are determined. There is a first preset reward relationship between different combination relationships and the target reward data. Then, after determining the target combination relationship based on the prediction type and the second analysis result, the target reward data corresponding to the prediction data is determined according to the target combination relationship and the first preset reward relationship.
[0060] For example, if the second analysis result is a tool called during the inference process of the reference model, then the maximum reward value can be applied to the first type of predicted data, no reward or reduced reward value can be applied to the third type of predicted data, no reward or smaller penalty can be applied to the fourth type of predicted data, and the maximum penalty can be applied to the second type of predicted data.
[0061] If the second analysis result indicates that the tool was not invoked during the reference model inference process, then the maximum reward value can be applied to the third type of predicted data, no reward or a reduced reward value can be applied to the first type of predicted data, no reward or a smaller penalty can be applied to the second type of predicted data, and the maximum penalty can be applied to the fourth type of predicted data. Of course, the first preset reward relationship can also be other correspondences; this is only an example and is not limited.
[0062] In this embodiment of the invention, based on first comparison information and first analysis results, the prediction type of each predicted data is determined. This enables precise analysis of both the final prediction result and the inference process of the model to be trained, thereby accurately classifying the predicted data. This facilitates a more accurate identification of factors influencing the adjustment of the parameters of the model to be trained. Furthermore, based on the prediction type and second analysis results, the target reward data corresponding to the predicted data is determined. The inference content of the reference model is added as a reference item to determine the factors influencing the parameter updates of the model to be trained, ensuring the accuracy of the target reward data corresponding to the predicted data. This allows for accurate parameter updates of the model to be trained based on the target reward data, improving model training efficiency.
[0063] In an embodiment of the present invention, optionally, if the prediction type is a first type or a second type, the first analysis result further includes a first calling tool identifier, and the second analysis result further includes a second calling tool identifier. Determining the target reward data corresponding to the prediction data based on the prediction type and the second analysis result may include steps B1-B3: Step B1: Determine the second alignment information based on the first and second invocation tool identifiers; the second alignment information is used to describe whether the first invocation tool identifier invoked by the model to be trained is consistent with the second invocation tool identifier invoked in the reference model.
[0064] The second comparison information may include whether the first calling tool identifier and the second calling tool identifier are the same or not.
[0065] Step B2: Based on the second alignment information and the prediction type, determine the prediction subtype; the prediction subtype is one of the first subtype, the second subtype, the third subtype, and the fourth subtype; the first subtype is used to describe the model to be trained correctly calling the tool and the prediction result is correct; the second subtype is used to describe the model to be trained correctly calling the tool and the prediction result is incorrect; the third subtype is used to describe the model to be trained not correctly calling the tool and the prediction result is correct; the fourth subtype is used to describe the model to be trained not correctly calling the tool and the prediction result is incorrect.
[0066] Specifically, based on the combination of the second alignment information and the prediction type, the prediction type can be divided into four prediction subtypes: the first subtype describes the model under training correctly calling the tool and the prediction result being correct; the second subtype describes the model under training correctly calling the tool and the prediction result being incorrect; the third subtype describes the model under training not correctly calling the tool and the prediction result being correct; and the fourth subtype describes the model under training not correctly calling the tool and the prediction result being incorrect.
[0067] Step B3: Based on the prediction subtype, determine the target reward data corresponding to the prediction data.
[0068] Specifically, given that there is a second preset reward relationship between each prediction subtype and the reward data, after determining the prediction subtype of the prediction data based on the second comparison information and the prediction type, the target reward data corresponding to the prediction data is determined according to the prediction subtype and the second preset reward relationship.
[0069] For example, the maximum reward value is applied to the predicted data of the first subtype, no reward or a reduced reward value is applied to the predicted data of the third subtype, no reward or a smaller penalty is applied to the predicted data of the fourth subtype, and the maximum penalty is applied to the predicted data of the second subtype. Of course, the second preset reward relationship can also be other correspondences; this is only an example and is not limited.
[0070] In this embodiment of the invention, based on the first and second tool identifiers, second comparison information is determined to further subdivide the prediction types of the first or second type based on the second comparison information into prediction subtypes. This allows for a more precise distinction between whether the model to be trained correctly invokes the tool when it is invoked, thereby more accurately determining the target reward data corresponding to the prediction data.
[0071] The technical solution of this invention involves acquiring multiple sets of multimodal training data; inputting the multimodal training data into a model to be trained to obtain at least one prediction data point corresponding to each set of multimodal training data; acquiring at least one reference model, inputting the multimodal training data into the reference model to obtain second inference content corresponding to each set of multimodal training data; determining first comparison information based on the prediction results in the prediction data and the reference results corresponding to the multimodal training data, thereby accurately determining whether the prediction results output by the model to be trained are correct, facilitating the subsequent application of rewards to the prediction results of the model to be trained. Then, based on the first inference content in the prediction data, determining a first analysis result, thereby accurately locating whether tools are invoked during the inference process of the model to be trained, facilitating the subsequent application of rewards to the parameters of the inference process of the model to be trained. Simultaneously, based on the second inference content, the second analysis result is determined, and the content regarding whether the tool is called during the inference process of the reference model is introduced to determine the accuracy of the first analysis result. The entire process ensures the accuracy of determining the target reward data corresponding to the predicted data based on the first comparison information, the first analysis result, and the second analysis result. Furthermore, this invention combines the intermediate inference process of the model to be trained and the prediction result to determine the target reward data that affects the adjustment of the parameters of the model to be trained, thereby achieving accurate updating of the parameters of the model to be trained based on the target reward data and obtaining the target model. This solves the problem of low model training efficiency in the entire training process where the model has the ability to correctly solve problems using tools.
[0072] Figure 3 This is a flowchart illustrating another reinforcement learning training method for a multimodal model provided by an embodiment of the present invention. The technical solution of this embodiment further optimizes the process S120 in the aforementioned embodiments based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 3 As shown, the reinforcement learning training method for the multimodal model of the present invention includes: S310. Obtain multiple sets of multimodal training data; each set of multimodal training data includes tool usage prompt text content and image content; the tool usage prompt text content includes at least one tool processing component from the following: at least one callable tool identifier, callable toolset, and interface information of the tool encoding model.
[0073] Each tool in the callable toolset includes a tool name / identifier and a description of its usage / function. This allows the model to obtain the tool's function description and instructions on how to correctly use the specified tool based on the tool name / identifier in the multimodal training data, thus enabling the model to output the correct calling format when it determines that a tool needs to be used.
[0074] S320. Input multimodal training data into the model to be trained and output the first reference inference content; the first reference inference content includes whether to call the tool processing component and / or the data processing requirements for calling the tool processing component.
[0075] The data processing requirements may include inference information output during the inference process of the model to be trained, such as which target tool processing components are called and how they are processed.
[0076] S330. If it is determined based on the first reference inference content that the model to be trained needs to call the target tool processing component, then the first data processing requirement is determined according to the data processing requirements of the target tool processing component; the first data processing requirement is used to describe the processing logic of calling the target tool processing component to process the multimodal training data.
[0077] Specifically, the text of the data processing requirements of the target tool processing component is parsed to identify the processing logic that describes calling the target tool processing component to process multimodal training data, thereby obtaining the first data processing requirements.
[0078] S340. Call the target tool processing component to process the multimodal training data according to the first data processing requirements, and obtain at least one first intermediate result.
[0079] Specifically, the target tool processing component is invoked so that it processes the multimodal training data according to the first data processing requirements to obtain at least one first intermediate result.
[0080] S350. Based on all first intermediate results and the model to be trained, obtain at least one prediction data corresponding to each multimodal training data; the prediction data includes the prediction result and the first inference content.
[0081] Specifically, all the first intermediate results, the multimodal training data corresponding to the first intermediate results, and the prediction output data corresponding to the first reference inference content are input into the model to be trained for re-inference to obtain at least one prediction data corresponding to each multimodal training data.
[0082] In an embodiment of the present invention, optionally, the tool usage prompt text includes prompts for calling the tool and a description of the tool's usage / functions; the prompts for calling the tool include: at least one identifier of a callable tool and / or the content of a callable toolset; inputting multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data may include steps C1-C4: Step C1: Input multimodal training data into the model to be trained for the first pre-filling to obtain the second reference inference content; the second reference inference content includes whether to call at least one target tool for data processing and / or the data processing requirements for calling at least one target tool for processing.
[0083] Step C2: If it is determined based on the second reference inference content that the model to be trained needs to call at least one target tool, then determine the second data processing requirement according to the usage / function description and data processing requirements of the target tool, and output the first call code; the second data processing requirement is used to describe the processing logic of calling the target tool to process the multimodal training data.
[0084] Specifically, the data processing requirements of the target tools are parsed to identify the processing logic that describes which target tools are invoked and how they process the multimodal training data, thus obtaining the second data processing requirements.
[0085] Step C3: Based on the first call encoding, stop the model to be trained from continuing to predict the output state, and call the target tool to process the multimodal training data according to the data processing requirements of the target tool, and output at least one second intermediate result.
[0086] Specifically, the target tool is invoked to process the multimodal training data according to the second data processing requirements, so as to obtain at least one second intermediate result.
[0087] Step C4: Input at least one second intermediate result, the multimodal training data corresponding to the second intermediate result, and the prediction output data corresponding to the first pre-filling into the model to be trained for the second pre-filling, and continue to perform prediction output to obtain at least one prediction data corresponding to each multimodal training data.
[0088] For example, taking the image cropping tool as an example, a more detailed description of the process of inputting multimodal training data into the model to be trained and obtaining at least one prediction data corresponding to each multimodal training data can include: (1) In the first pre-filling stage, the prompts for calling the tools, the usage / function introduction of the corresponding tools, and the corresponding images and requirements will be pre-filled for the first time, and the predicted output will be made.
[0089] (2) During the prediction output process, if the model to be trained decides to call the image cropping tool, it outputs the second data processing requirement and the first call code according to the usage / function introduction and data processing requirements of the image cropping tool; wherein, the second data processing requirement can be image cropping based on a preset number of coordinate points in the image.
[0090] (3) When the model to be trained receives the first call code, it stops predicting the output and calls the image cropping tool to crop the input image according to the preset number of coordinate points on the image and outputs the cropped partial image.
[0091] (4) Re-input the cropped local image, the corresponding multimodal training data, and the prediction output data before termination into the model for a second pre-filling; (5) After the second pre-filling is completed, continue to make prediction output, complete the inference process, and obtain at least one prediction data corresponding to each multimodal training data.
[0092] In this embodiment of the invention, for the case of at least one callable tool identifier and at least one tool processing component in the callable tool set, the process of inputting multimodal training data into the model to be trained and obtaining at least one prediction data corresponding to each multimodal training data is described in detail, so as to achieve accurate analysis of how the model to be trained outputs at least one prediction data corresponding to each multimodal training data in this case.
[0093] In an embodiment of the present invention, optionally, the tool usage prompt text includes interface information of the tool's encoded model. Inputting multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data may include steps D1-D5: Step D1: Input the multimodal training data into the model to be trained for the first pre-filling to obtain the third reference inference content; the third reference inference content includes whether to call the tool encoding model for processing and / or the data processing requirements for calling the tool encoding model for processing.
[0094] Step D2: If it is determined based on the third reference inference content that the model to be trained needs to call the target tool encoding model, then obtain the target interface information of the target tool encoding model and output the second call code.
[0095] Step D3: Based on the second call encoding, stop the model to be trained from continuing to predict the output state, and call the target tool encoding model based on the target interface information. According to the target tool encoding model and the data processing requirements of the target tool encoding model, generate tool code and third data processing requirements; the third data processing requirements are used to describe the processing logic of processing multimodal training data based on the tool code.
[0096] The tool code is used as the specified tool to be invoked.
[0097] Specifically, the data processing requirements of the target tool encoding model are input into the target tool encoding model to generate tool code; the data processing requirements of the target tool encoding model are parsed to identify the processing logic that describes the processing of multimodal training data based on the tool code, thus obtaining the third data processing requirements.
[0098] Step D4: Based on the tool code, process the multimodal training data according to the third data processing requirements, and output at least one third intermediate result.
[0099] Specifically, the tool code is obtained, and the tool code is compiled and run to generate the specified tool to be called, so that the specified tool processes the multimodal training data according to the third data processing requirements and outputs at least one third intermediate result.
[0100] Step D5: Input at least one third intermediate result, the multimodal training data corresponding to the third intermediate result, and the prediction output data corresponding to the first pre-filling into the model to be trained for the second pre-filling, and continue to perform prediction output to obtain at least one prediction data corresponding to each multimodal training data.
[0101] This invention, in the case of a tool-encoded model interface information tool processing component, provides a detailed description of the process of inputting multimodal training data into the model to be trained and obtaining at least one prediction data corresponding to each multimodal training data point. This enables precise analysis of how the model to be trained outputs at least one prediction data corresponding to each multimodal training data point in this scenario. By setting a tool-encoded model, greater flexibility can be achieved compared to directly setting a specified tool or toolset, allowing the model to be trained to obtain greater data processing capabilities, thereby improving the performance of the target model.
[0102] S360. Obtain at least one reference model, input multimodal training data into the reference model, and obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has the function of calling tools.
[0103] S370. Based on the prediction data and the second inference content, determine the target reward data corresponding to the prediction data, update the parameters of the model to be trained based on the target reward data, and obtain the target model; wherein, the target reward data includes information from two dimensions: prediction results and the first inference content.
[0104] The technical solution of this invention involves acquiring multiple sets of multimodal training data. Each set of multimodal training data includes tool usage prompt text and image content. The tool usage prompt text includes at least one tool processing component from the interface information of at least one callable tool identifier, a callable toolset, and a tool encoding model. The multimodal training data is input into the model to be trained, and a first reference inference content is output. The first reference inference content includes whether to call the tool processing component and / or the data processing requirements for calling the tool processing component. Further, if it is determined based on the first reference inference content that the model to be trained needs to call a target tool processing component, a first data processing requirement is determined according to the data processing requirements of the target tool processing component. Since the first data processing requirement describes the processing logic of calling the target tool processing component to process the multimodal training data, the target tool processing component is called so that it can accurately process the multimodal training data according to the first data processing requirement, obtaining at least one accurate first intermediate result. Then, based on all the first intermediate results and the model to be trained, at least one prediction data corresponding to each set of multimodal training data is accurately obtained. Simultaneously, at least one reference model is obtained, and multimodal training data is input into the reference model to obtain the second inference content corresponding to each piece of multimodal training data. Based on the prediction data and the second inference content, the target reward data corresponding to the prediction data is determined, and the parameters of the model to be trained are updated based on the target reward data to obtain the target model. This solves the problem of low model training efficiency in the entire training process of training the model to have the ability to correctly solve problems using tools.
[0105] Figure 4 This is a schematic diagram of another reinforcement learning training device for a multimodal model provided in an embodiment of the present invention. This embodiment is applicable to reinforcement learning training of multimodal models, enabling the multimodal model to correctly call tools. This reinforcement learning training device for multimodal models can be implemented in hardware and / or software, and can be configured in any electronic device with network communication capabilities. Figure 4 As shown, the reinforcement learning training device for the multimodal model of the present invention may include: The multimodal training data acquisition module 410 is used to acquire multiple pieces of multimodal training data; wherein each piece of training data includes tool usage prompt text content and image content; The prediction module 420 is used to input the multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes a prediction result and a first inference content; The data acquisition module 430 is used to acquire at least one reference model, input the multimodal training data into the reference model, and obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has the function of calling tools; The update module 440 is used to determine the target reward data corresponding to the prediction data based on the prediction data and the second inference content, and update the parameters of the model to be trained based on the target reward data to obtain the target model; wherein, the target reward data includes information in two dimensions: prediction result and the first inference content.
[0106] Based on the above embodiments, optionally, the update module includes: a comparison unit, a first analysis unit, a second analysis unit, and a reward data determination unit: the comparison unit is used to determine first comparison information based on the prediction results in the prediction data and the reference results corresponding to the multimodal training data; wherein, the first comparison information is used to describe whether the prediction results in the prediction data are correct; the first analysis unit is used to determine a first analysis result based on the first inference content in the prediction data; wherein, the first analysis result includes whether the tool is called during the inference process of the model to be trained; the second analysis unit is used to determine a second analysis result based on the second inference content; wherein, the second analysis result includes whether the tool is called during the inference process of the reference model; the reward data determination unit is used to determine the target reward data corresponding to the prediction data based on the first comparison information, the first analysis result, and the second analysis result.
[0107] Optionally, based on the above embodiments, the reward data determination unit is further configured to: determine the prediction type of each predicted data based on the first comparison information and the first analysis result; the prediction type is one of a first type, a second type, a third type, and a fourth type; the first type is used to describe that the model to be trained calls the tool and the prediction result is correct; the second type is used to describe that the model to be trained calls the tool and the prediction result is incorrect; the third type is used to describe that the model to be trained does not call the tool and the prediction result is correct; the fourth type is used to describe that the model to be trained does not call the tool and the prediction result is incorrect; and determine the target reward data corresponding to the predicted data based on the prediction type and the second analysis result.
[0108] Based on the above embodiments, optionally, if the prediction type is a first type or a second type, the first analysis result further includes a first calling tool identifier, the second analysis result further includes a second calling tool identifier, and the reward data determination unit further includes a reward data determination subunit. The reward data determination subunit is used to: determine second comparison information based on the first calling tool identifier and the second calling tool identifier; the second comparison information is used to describe whether the first calling tool identifier called by the model to be trained is consistent with the second calling tool identifier called in the reference model; determine a prediction subtype based on the second comparison information and the prediction type; the prediction subtype is one of a first subtype, a second subtype, a third subtype, and a fourth subtype; the first subtype is used to describe that the model to be trained correctly calls the tool and the prediction result is correct; the second subtype is used to describe that the model to be trained correctly calls the tool and the prediction result is incorrect; the third subtype is used to describe that the model to be trained does not correctly call the tool and the prediction result is correct; the fourth subtype is used to describe that the model to be trained does not correctly call the tool and the prediction result is incorrect; and determine the target reward data corresponding to the prediction data based on the prediction subtype.
[0109] Based on the above embodiments, optionally, the prediction module is used to: for each piece of multimodal training data, the model to be trained uses a parallel sampling method to output the corresponding multiple prediction data; Optionally, based on the above embodiments, the update module is further configured to: determine the corresponding target reward data for each prediction data corresponding to each multimodal training data, and obtain multiple target reward data; determine the gradient value based on each prediction data and target reward data corresponding to each multimodal training data, and update the parameters of the model to be trained based on the gradient value to obtain the target model.
[0110] Based on the above embodiments, optionally, the tool usage prompt text content includes at least one tool processing component from at least one callable tool identifier, callable toolset, and interface information of the tool encoding model; the prediction module includes a first prediction unit, which is used to: input the multimodal training data into the model to be trained, and output first reference inference content; the first reference inference content includes whether to call the tool processing component and / or the data processing requirements for calling the tool processing component for processing; if it is determined based on the first reference inference content that the model to be trained needs to call the target tool processing component, then determine a first data processing requirement according to the data processing requirements of the target tool processing component; the first data processing requirement is used to describe the processing logic of calling the target tool processing component to process the multimodal training data; call the target tool processing component, process the multimodal training data according to the first data processing requirement, and obtain at least one first intermediate result; based on all the first intermediate results and the model to be trained, obtain at least one prediction data corresponding to each multimodal training data.
[0111] Based on the above embodiments, optionally, the tool usage prompt text includes prompts for calling the tool and a description of the tool's usage / function; the prompts for calling the tool include: at least one callable tool identifier and / or the content of a callable toolset; the prediction module includes a second prediction unit, which is used to: input the multimodal training data into the model to be trained for a first pre-filling to obtain second reference inference content; the second reference inference content includes whether to call at least one target tool for data processing and / or the data processing requirements for calling at least one target tool for processing; if it is determined based on the second reference inference content that the model to be trained needs to call at least one target tool, then according to the description of the target tool's usage / function... Based on the data processing requirements, a second data processing requirement is determined, and a first invocation code is output. The second data processing requirement describes the processing logic of invoking the target tool to process the multimodal training data. Based on the first invocation code, the state of the prediction output of the model to be trained is stopped, and the target tool is invoked to process the multimodal training data according to the data processing requirements of the target tool, outputting at least one second intermediate result. The at least one second intermediate result, the multimodal training data corresponding to the second intermediate result, and the prediction output data corresponding to the first pre-filling are input into the model to be trained for a second pre-filling, and prediction output continues to be performed to obtain at least one prediction data corresponding to each multimodal training data.
[0112] Based on the above embodiments, optionally, the tool usage prompt text content includes interface information of the tool encoding model, and the prediction module includes a third prediction unit. The third prediction unit is used to: input the multimodal training data into the model to be trained for a first pre-filling to obtain third reference inference content; the third reference inference content includes whether to call the tool encoding model for processing and / or the data processing requirements for calling the tool encoding model for processing; if it is determined based on the third reference inference content that the model to be trained needs to call the target tool encoding model, then obtain the target interface information of the target tool encoding model and output a second call code; based on the second call code, stop the model to be trained from continuing to predict the output state, and based on the target interface... The information calls the target tool encoding model, and generates tool code and a third data processing requirement based on the target tool encoding model and its data processing requirements. The third data processing requirement describes the processing logic for processing the multimodal training data based on the tool code. Based on the tool code, the multimodal training data is processed according to the third data processing requirement, and at least one third intermediate result is output. The at least one third intermediate result, the multimodal training data corresponding to the third intermediate result, and the prediction output data corresponding to the first pre-filling are input into the model to be trained for a second pre-filling, and prediction output is continued to obtain at least one prediction data corresponding to each multimodal training data.
[0113] Based on the above embodiments, optionally, the reward data determination unit is further configured to: acquire multiple reference models and generate multiple second inference contents; determine the confidence weight of each reference model; the confidence weight is used to reflect the tool invocation performance of the reference model; determine the target second analysis result corresponding to the prediction data based on the second analysis result of each second inference content and the confidence weight of each reference model; and determine the target reward data corresponding to the prediction data based on the first comparison information, the first analysis result and the target second analysis result.
[0114] The reinforcement learning training device for multimodal models provided in this embodiment of the invention can execute the reinforcement learning training method for multimodal models provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0115] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0116] Figure 5A schematic diagram of an electronic device is shown, which can be used to implement the reinforcement learning training method for multimodal models according to embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0117] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0118] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0119] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as reinforcement learning training methods for multimodal models.
[0120] In some embodiments, the reinforcement learning training method for the multimodal model can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the reinforcement learning training method for the multimodal model described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the reinforcement learning training method for the multimodal model by any other suitable means (e.g., by means of firmware).
[0121] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0122] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0123] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0125] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0126] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0127] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0128] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A reinforcement learning training method for a multimodal model, characterized in that, The method includes: Acquire multiple sets of multimodal training data; wherein each set of multimodal training data includes tool usage prompt text content and image content; The multimodal training data is input into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes the prediction result and the first inference content; At least one reference model is obtained, and the multimodal training data is input into the reference model to obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has the function of calling tools; Based on the predicted data and the second inference content, the target reward data corresponding to the predicted data is determined, and the parameters of the model to be trained are updated based on the target reward data to obtain the target model; wherein, the target reward data includes information from two dimensions: the prediction result and the first inference content.
2. The method according to claim 1, characterized in that, Based on the predicted data and the second inference content, the target reward data corresponding to the predicted data is determined, including: Based on the prediction results in the prediction data and the reference results corresponding to the multimodal training data, first comparison information is determined; wherein, the first comparison information is used to describe whether the prediction results in the prediction data are correct. Based on the first inference content in the predicted data, a first analysis result is determined; wherein, the first analysis result includes whether the tool is called during the inference process of the model to be trained; Based on the second reasoning content, a second analysis result is determined; wherein, the second analysis result includes whether a tool was invoked during the reasoning process of the reference model; Based on the first comparison information, the first analysis result, and the second analysis result, the target reward data corresponding to the predicted data is determined.
3. The method according to claim 2, characterized in that, Based on the first comparison information, the first analysis result, and the second analysis result, the target reward data corresponding to the predicted data is determined, including: Based on the first comparison information and the first analysis result, the prediction type of each predicted data is determined; the prediction type is one of a first type, a second type, a third type, and a fourth type; the first type describes the model to be trained calling the tool and the prediction result is correct; the second type describes the model to be trained calling the tool and the prediction result is incorrect; the third type describes the model to be trained not calling the tool and the prediction result is correct; the fourth type describes the model to be trained not calling the tool and the prediction result is incorrect. Based on the prediction type and the second analysis result, the target reward data corresponding to the prediction data is determined.
4. The method according to claim 3, characterized in that, If the prediction type is a first type or a second type, the first analysis result further includes a first calling tool identifier, and the second analysis result further includes a second calling tool identifier. Based on the prediction type and the second analysis result, the target reward data corresponding to the prediction data is determined, including: Based on the first calling tool identifier and the second calling tool identifier, second comparison information is determined; the second comparison information is used to describe whether the first calling tool identifier called by the model to be trained is consistent with the second calling tool identifier called in the reference model. Based on the second comparison information and the prediction type, a prediction subtype is determined; the prediction subtype is one of a first subtype, a second subtype, a third subtype, and a fourth subtype; the first subtype describes that the model to be trained correctly calls the tool and the prediction result is correct; the second subtype describes that the model to be trained correctly calls the tool and the prediction result is incorrect; the third subtype describes that the model to be trained does not correctly call the tool and the prediction result is correct; the fourth subtype describes that the model to be trained does not correctly call the tool and the prediction result is incorrect. Based on the predicted subtype, the target reward data corresponding to the predicted data is determined.
5. The method according to any one of claims 1 to 4, characterized in that, The step of inputting the multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data includes: For each piece of multimodal training data, the model to be trained uses a parallel sampling method to output multiple corresponding prediction data; Accordingly, based on the predicted data and the second inference content, the target reward data corresponding to the predicted data is determined, and the parameters of the model to be trained are updated based on the target reward data to obtain the target model, including: For each prediction data corresponding to each multimodal training data, a corresponding target reward data is determined, resulting in multiple target reward data. The gradient value is determined based on each predicted data and target reward data corresponding to each multimodal training data, and the parameters of the model to be trained are updated based on the gradient value to obtain the target model.
6. The method according to claim 1, characterized in that, The tool usage prompt text includes at least one tool processing component from at least one callable tool identifier, callable toolset, and interface information of the tool encoding model; the step of inputting the multimodal training data into the model to be trained to obtain at least one prediction data corresponding to each multimodal training data includes: The multimodal training data is input into the model to be trained, and a first reference inference content is output; the first reference inference content includes whether to call the tool processing component and / or the data processing requirements for calling the tool processing component. If it is determined based on the first reference inference content that the model to be trained needs to call the target tool processing component, then a first data processing requirement is determined according to the data processing requirements of the target tool processing component; the first data processing requirement is used to describe the processing logic of calling the target tool processing component to process the multimodal training data; The target tool processing component is invoked to process the multimodal training data according to the first data processing requirements, thereby obtaining at least one first intermediate result; Based on all the first intermediate results and the model to be trained, at least one prediction data point is obtained for each multimodal training data point.
7. The method according to claim 6, characterized in that, The tool usage prompt text includes prompts for calling the tool and a description of its usage / functions; the prompts for calling the tool include: at least one callable tool identifier and / or the content of a callable toolset; inputting the multimodal training data into the model to be trained to obtain at least one prediction data point corresponding to each multimodal training data point includes: The multimodal training data is input into the model to be trained for the first pre-filling to obtain the second reference inference content; the second reference inference content includes whether to call at least one target tool for data processing and / or the data processing requirements for calling at least one target tool for processing; If, based on the second reference inference content, it is determined that the model to be trained needs to call at least one target tool, then, according to the usage / function description and data processing requirements of the target tool, a second data processing requirement is determined, and a first call code is output; the second data processing requirement is used to describe the processing logic of calling the target tool to process the multimodal training data; Based on the first invocation code, the model to be trained is stopped from predicting the output state, and the target tool is invoked to process the multimodal training data according to the data processing requirements of the target tool, and at least one second intermediate result is output. At least one second intermediate result, the multimodal training data corresponding to the second intermediate result, and the prediction output data corresponding to the first pre-filling are input into the model to be trained for a second pre-filling, and prediction output is continued to obtain at least one prediction data corresponding to each multimodal training data.
8. The method according to claim 6, characterized in that, The tool's usage prompt text includes interface information for the tool's encoded model. It inputs the multimodal training data into the model to be trained, obtaining at least one prediction data point corresponding to each multimodal training data point, including: The multimodal training data is input into the model to be trained for the first pre-filling to obtain the third reference inference content; the third reference inference content includes whether to call the tool encoding model for processing and / or the data processing requirements for calling the tool encoding model for processing; If it is determined based on the third reference inference content that the model to be trained needs to call the target tool encoding model, then the target interface information of the target tool encoding model is obtained, and the second call code is output; Based on the second invocation encoding, the training model is stopped from predicting the output state, and the target tool encoding model is invoked based on the target interface information. According to the target tool encoding model and its data processing requirements, tool code and a third data processing requirement are generated. The third data processing requirement describes the processing logic for processing the multimodal training data based on the tool code. Based on the tool code, the multimodal training data is processed according to the third data processing requirements, and at least one third intermediate result is output. At least one of the third intermediate results, the multimodal training data corresponding to the third intermediate results, and the prediction output data corresponding to the first pre-filling are input into the model to be trained for a second pre-filling, and prediction output is continued to obtain at least one prediction data corresponding to each multimodal training data.
9. The method according to claim 2, characterized in that, Based on the first comparison information, the first analysis result, and the second analysis result, the target reward data corresponding to the predicted data is determined, including: Obtain multiple reference models and generate multiple second-order inference statements; Determine the confidence weight for each reference model; the confidence weight is used to reflect the tool invocation performance of the reference model. Based on the second analysis result of each second inference content and the confidence weight of each reference model, the target second analysis result corresponding to the predicted data is determined; Based on the first comparison information, the first analysis result, and the target second analysis result, the target reward data corresponding to the predicted data is determined.
10. A reinforcement learning training device for a multimodal model, characterized in that, The device includes: A multimodal training data acquisition module is used to acquire multiple pieces of multimodal training data; wherein each piece of training data includes tool usage prompt text content and image content; The prediction module is used to input the multimodal training data into the model to be trained and obtain at least one prediction data corresponding to each multimodal training data; wherein, the prediction data includes the prediction result and the first inference content; The data acquisition module is used to acquire at least one reference model, input the multimodal training data into the reference model, and obtain the second inference content corresponding to each piece of multimodal training data; wherein, the reference model is a model of the same type as the model to be trained and has tool calling function; An update module is used to determine the target reward data corresponding to the prediction data based on the prediction data and the second inference content, and update the parameters of the model to be trained based on the target reward data to obtain the target model; wherein, the target reward data includes information from two dimensions: prediction result and the first inference content.