Inference model training method, device, server, storage medium and product

By using image and text datasets to label and encode the multimodal pre-training model during the training process, the problem of low accuracy of large-scale language models in multimodal thinking is solved, the reasoning ability of multimodal chain thinking is realized, and the accuracy of reasoning answers is improved.

CN120525015BActive Publication Date: 2025-10-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511007853.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-03
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

Current large-scale language models have significant bottlenecks in simulating human multimodal thinking, resulting in reduced accuracy of reasoning answers. The reason is that multimodal data is converted into a single modality for processing.

Method used

During the training process of the reasoning model, a data set containing images and text is collected, and the multimodal pre-training model is trained after labeling and encoding to generate reasoning capabilities with multimodal chain thinking, ensuring that visual information is considered in the reasoning process.

Benefits of technology

It improves the accuracy of reasoning answers, avoids result deviation caused by single-modal processing, and enhances the processing ability of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525015B_ABST
    Figure CN120525015B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, server, storage medium, and product for training an inference model, relating to the field of artificial intelligence technology. The method uses a data set containing an inference process, including images and text, for training. To improve the training effect, the data set is labeled. The inference process not only uses textual reasoning, but also generates images, incorporating visual information into the inference process. This enables the inference model to possess the inference capability of multimodal chain thinking, avoiding deviations in the results of reasoning problem processing, thereby improving the accuracy of the inference answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to training methods, devices, servers, storage media, and products for reasoning models. Background Art

[0002] Early large-scale language models focused on rapid decision-making problems, such as language generation and knowledge question answering. With the development of large-scale language models, their ability to handle logical problems has improved. However, current large-scale language models still face significant bottlenecks in simulating human multimodal thinking.

[0003] Currently, even when the problem to be inferred contains multimodal data, the inference process often involves converting the multimodal data into a single modality, namely, text. However, this approach is inherently monomodal, leading to biased results and reduced accuracy of the inference answer. Summary of the Invention

[0004] This application provides a training method, device, server, storage medium and product for an inference model to at least solve the problem of low accuracy of inference answers in related technologies.

[0005] This application provides a method for training an inference model, including:

[0006] Collecting a data set for training; each data in the data set includes reasoning questions, reasoning processes, and reasoning answers, and the reasoning process includes images and text;

[0007] According to the reasoning question, reasoning process and reasoning answer, the data set is labeled to obtain a labeled data set;

[0008] Encode the images and texts in the labeled dataset to obtain an encoded dataset;

[0009] The multimodal pre-training model is trained according to the encoded data set and the pre-set hyperparameters of the multimodal pre-training model to obtain an inference model; wherein the inference model is used to output the inference process and inference answer according to the problem to be inferred.

[0010] This application also provides a training device for an inference model, comprising:

[0011] The acquisition module is used to collect data sets for training; each data in the data set includes reasoning questions, reasoning processes, and reasoning answers, and the reasoning process includes images and text;

[0012] The labeling module is used to label the data set according to the reasoning question, reasoning process and reasoning answer to obtain the labeled data set;

[0013] An encoding module is used to encode the images and texts in the labeled dataset to obtain an encoded dataset;

[0014] The training module is used to train the multimodal pre-training model based on the encoded data set and the pre-set hyperparameters of the multimodal pre-training model to obtain an inference model; wherein the inference model is used to output the inference process and inference answer according to the problem to be inferred.

[0015] The present application also provides a server, comprising: a memory for storing a computer program; and a processor for implementing the steps of the training method of any of the above-mentioned inference models when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the training method of any of the above-mentioned inference models are implemented.

[0017] The present application also provides a computer program product, including a computer program, which implements the steps of the training method of any of the above-mentioned inference models when the computer program is executed by a processor.

[0018] This application uses a dataset containing reasoning processes, including both images and text, for training. To improve training effectiveness, the dataset is labeled. The reasoning process not only uses textual reasoning but also generates images, incorporating visual information into the process. This gives the reasoning model the ability to reason in a multimodal chain, avoiding deviations in the results of reasoning problem processing and thus improving the accuracy of the reasoning answers. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A schematic diagram of a scenario for the training method of the inference model provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of a method for training an inference model provided in an embodiment of the present application;

[0022] Figure 3 This is an architectural diagram of the multimodal pre-training model provided in an embodiment of the present application;

[0023] Figure 4A schematic diagram of the structure of a training device for an inference model provided in an embodiment of the present application;

[0024] Figure 5 A schematic diagram of the structure of the server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0027] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.

[0028] Early large-scale language models focused on rapid decision-making problems, such as language generation and knowledge question answering. With the development of large-scale language models, their ability to handle logical problems has improved. However, current large-scale language models still face significant bottlenecks in simulating human multimodal thinking. Currently, even if the problem to be inferred contains multimodal data, the inference process will first convert the multimodal data into a single modality, namely text modality. However, this method is still essentially single-modal, which leads to deviations in the results of reasoning problem processing and reduces the accuracy of the inference answer.

[0029] In order to solve the technical problems in the above-mentioned related technologies, this application proposes the following technical concept: Considering that multimodal data is first converted into a single modality during the reasoning process, in essence, only a single modality is used. The inventor thought of considering multimodal reasoning methods during the training of the model. In the reasoning process, not only textual reasoning is used, but also images are generated, so that visual information is added to the reasoning process, so that the reasoning model has the reasoning ability of multimodal chain thinking. In addition, in order to improve the training effect, the collected data sets are marked. Avoid deviations in the results of reasoning problem processing, thereby improving the accuracy of the reasoning answers.

[0030] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0031] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the training method of the inference model depends, the specific application environment architecture or specific hardware architecture is described here.

[0032] refer to Figure 1 , Figure 1 A schematic diagram of a scenario of a training method for an inference model provided in an embodiment of the present application, such as Figure 1 As shown, the server provided by this embodiment includes: a receiving device 101, a processing device 102 and a display device 103.

[0033] It is understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the training method of the inference model. In other feasible implementations of the present application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, split certain components, or arrange the components differently. The specific configuration can be determined according to the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0034] In a specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can receive a data set collected for training.

[0035] The processing device 102 may perform a series of processing on the data set to obtain an inference model.

[0036] The display device 103 can be used to display the output of the inference model.

[0037] It should be understood that the above-mentioned processor can be implemented by the processor reading instructions in the memory and executing the instructions, or it can be implemented by a circuit.

[0038] In addition, the network architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0039] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0040] Figure 2 A flow chart of the training method of the inference model provided in the embodiment of the present application is shown as follows: Figure 2 As shown, the embodiment of the present application provides a method for training an inference model, and the method is described in detail as follows:

[0041] S201: Collect a data set for training; wherein each data in the data set includes reasoning questions, reasoning processes and reasoning answers, and the reasoning process includes images and texts.

[0042] In this example, the training dataset was collected using a mixture of images and text, and the reasoning process employed a chain-like approach. The dataset, sourced from sources such as video data and science and engineering problem sets, encompasses a wide range of content, with the goal of enabling the inference model to learn the interleaving of images and text.

[0043] S202: Label the data set according to the reasoning question, the reasoning process, and the reasoning answer to obtain a labeled data set.

[0044] In this embodiment, the reasoning question carries a system prompt word.

[0045] In this example, in order to achieve better training results, supervised fine-tuning is performed on the dataset, and symbols are used to mark the images and texts in the dataset. The symbol prompt word template of the training data is as follows:

[0046] <|begin_of_sentence|>SYSTEM_RPOMPT<|User|> IMAGEQUESTION<|Assistant|> <think> IMAGEANALYSIS< / think> ANSWER<|end_of_sentence|>

[0047] Among them, <|begin_of_sentence|> and <|end_of_sentence|> are the data start and end markers for each data point in the dataset. SYSTEM_RPOMPT is a system prompt word used to indicate the question type, such as geometry and physics questions. <|User|> and <|Assistant|> respectively indicate the questioning role and answering role. For a single-turn dialogue, only one pair appears. For multi-turn dialogues or multiple rounds of questions, pairs appear multiple times. and appear in pairs, the middle part IMAGE is the image, Indicates the start mark of an image and the end mark of an image. <think> and< / think> Appear in pairs, the middle part of ANALYSIS is the reasoning process, <think> is the inference start marker,< / think> It marks the end of reasoning, where QUESTION is the reasoning question and ANSWER is the reasoning answer.

[0048] For example, take a physics question as an example:

[0049] The reasoning question is: Schematic image, as shown in the schematic image, there is an object with a mass m on a right-angled splitter M placed on the horizontal ground. If m slides down M at a constant speed and M remains stationary, then how much pressure and friction does M exert on the ground?

[0050] The reasoning process is as follows: First, use the isolation method to analyze. Isolate m. M is subject to gravity mg and the support force of the inclined plane, as shown in Force Analysis Graph 1. Since m slides down the inclined plane at a constant speed, the equilibrium condition shows that... Then isolate M. M is subject to a vertical downward gravity force Mg, and the ground exerts... on it, as shown in Force Analysis Graph 2. Based on Newton's third law, we obtain... Therefore, the ground exerts no friction on M.

[0051] The reasoned answer is: the pressure of M on the ground is equal to (M+m)g, and the friction force on the ground is 0.

[0052] Specifically, step S202 includes S2021 to S2024:

[0053] S2021: For each data in the data set, add a question role tag at the beginning of the reasoning question, and add an answer role tag at the beginning of the reasoning process to obtain the data set after the first tag.

[0054] For example, add a question role marker at the beginning of the reasoning question, and add an answer role marker at the beginning of the reasoning process:

[0055] <|User|> Schematic diagram. As shown in the schematic diagram, an object of mass m is placed on a right-angled splitter M on the horizontal ground. If m slides down M at a constant speed while M remains stationary, what are the pressure and friction forces exerted by M on the ground? <|Assistant|> Let's analyze this using the isolation method. First, isolate m. M is subject to gravity mg and the support force of the inclined plane..., as shown in Force Analysis Diagram 1. Since m slides down the inclined plane at a constant speed, the equilibrium condition shows that... Then isolate M. M is subject to a vertical downward gravity force Mg, and the ground exerts... on it, as shown in Force Analysis Diagram 2. According to Newton's third law,... Therefore, the ground exerts no friction on M. The pressure exerted by M on the ground is equal to (M + m) g, and the friction force on the ground is 0.

[0056] S2022: For each data in the first labeled data set, add an inference start marker at the beginning of the inference process and add an inference end marker at the end of the inference process to obtain a second labeled data set.

[0057] For example, an inference start marker is added at the beginning of the inference process, and an inference end marker is added at the end of the inference process:

[0058] <|User|>Draw an image. As shown in the image, there is an object of mass m on a right-angled wedge M placed on the horizontal ground. If m slides down M at a constant speed while M remains stationary, how much pressure and friction does M exert on the ground? <|Assistant|> <think> Let's analyze using the isolation method. First, isolate m. M is subject to gravity mg and the support force of the inclined plane..., as shown in Force Analysis Image 1. Since m slides down the inclined plane at a constant speed, the equilibrium condition shows... Next, isolate M. M is subject to a vertical downward gravity force Mg, and the ground exerts... on it, as shown in Force Analysis Image 2. Using Newton's third law, we obtain... Therefore, the ground exerts no friction on M.< / think> The pressure of M on the ground is equal to (M+m)g, and the friction force on the ground is 0.

[0059] S2023: For each data in the second marked data set, add an image start marker at the beginning of each image and add an image end marker at the end of each image to obtain a third marked data set.

[0060] For example, an image start marker is added at the beginning of each image, and an image end marker is added at the end of each image:

[0061] <|User|> Schematic diagram. As shown in the schematic diagram, there is an object of mass m on a right-angled wedge M placed on the horizontal ground. If m slides down M at a constant speed while M remains stationary, how much pressure and friction does M exert on the ground? <|Assistant|> <think>First use the isolation method to analyze, first isolate m, m is subject to gravity mg, the support force of the inclined surface..., as shown in the force analysis image 1, Force Analysis Image 1. Since m slides down the slope at a constant speed, the equilibrium condition shows that... . Then, M is isolated. M is subjected to a vertical downward gravity force Mg, and the ground exerts... on it, as shown in Force Analysis Image 2. Force analysis image 2; from Newton's third law, we get...so there is no friction force from the ground on M.< / think> The pressure of M on the ground is equal to (M+m)g, and the friction force on the ground is 0.

[0062] S2024: For each data in the third labeled data set, add a data start marker and a system prompt word at the beginning of the questioning role marker, and add a data end marker at the end of the inference answer to obtain the final labeled data set.

[0063] For example, a data start marker and a system prompt word are added at the beginning of the question role marker, and a data end marker is added at the end of the inference answer. After completing the marking of each initial general data, the final general data set is:

[0064] <|begin_of_sentence|>Physics Question<|User|> Schematic diagram. As shown in the schematic diagram, there is an object of mass m on a right-angled wedge M placed on the horizontal ground. If m slides down M at a constant speed while M remains stationary, how much pressure and friction does M exert on the ground? <|Assistant|> <think>First use the isolation method to analyze, first isolate m, m is subject to gravity mg, the support force of the inclined surface..., as shown in the force analysis image 1, Force Analysis Image 1. Since m slides down the slope at a constant speed, the equilibrium condition shows that... . Then, M is isolated. M is subjected to a vertical downward gravity force Mg, and the ground exerts... on it, as shown in Force Analysis Image 2. Force analysis image 2; from Newton's third law, we get...so there is no friction force from the ground on M.< / think> The pressure of M on the ground is equal to (M + m) g, and the friction force on the ground is 0. <|end_of_sentence|>

[0065] S203: Encode the images and texts in the labeled dataset to obtain an encoded dataset.

[0066] Specifically, for each data in the labeled data set, an image encoder is used to encode the image, and a text encoder is used to encode the text to obtain an encoded data set.

[0067] S204: The multimodal pre-training model is trained according to the encoded data set and the pre-set hyperparameters of the multimodal pre-training model to obtain an inference model; wherein the inference model is used to output an inference process and an inference answer according to the problem to be inferred.

[0068] In this embodiment, the multimodal pre-training model rarely has data with interleaving of images and text during pre-training. Therefore, at this stage, the model needs to be trained on the dataset.

[0069] Among them, the pre-set hyperparameters of the multimodal pre-training model include the minimum number of iterations, the maximum number of iterations and the loss function change threshold.

[0070] The multimodal pre-training model is a mixed-modality autoregressive model, pre-trained in an end-to-end manner on mixed modalities. The multimodal pre-training model itself has the ability to parse multi-resolution images and text. In the multimodal pre-training model, images and text are represented based on tokens. Images are similar to tokens in text. Images and text are encoded using an image encoder and text encoder, respectively, and then concatenated and fed into the multimodal pre-training model. Figure 3The architecture diagram of the multimodal pre-training model provided in the embodiment of the present application is as follows: Figure 3 As shown in the figure, the multimodal pre-trained model receives interleaved inputs of images and text, facilitating the generation of cross-modal reasoning. When outputting images, the model is subsequently connected to an image de-tokenizer to convert the image tokens back into images.

[0071] In this embodiment, the hyperparameters of the multimodal pre-training model also include learning rate and batch parameters.

[0072] Specifically, step S204 includes S2041 to S2044:

[0073] S2041: Load the multimodal pre-trained model.

[0074] In this embodiment, loading the multimodal pre-training model includes loading the multimodal pre-training model structure, optimizer and related parameters, etc.

[0075] S2042: Iteratively train the multimodal pre-trained model based on the encoded dataset.

[0076] S2043: During the training process, if the current iteration number is greater than the minimum iteration number and the loss function change value is greater than the loss function change threshold, then continue iterative training.

[0077] In this embodiment, if the current iteration number is greater than the minimum iteration number and the loss function change value is greater than the loss function change threshold, it means that the loss function value still has room to decrease, and the iterative training continues.

[0078] S2044: If the current number of iterations exceeds the maximum number of iterations, or / and the loss function change value is less than the loss function change threshold, stop the iteration and obtain the inference model.

[0079] In this embodiment, if the current iteration number exceeds the maximum iteration number, or / and the loss function change value is less than the loss function change threshold, it means that the loss function value decreases slowly and the inference model has reached the global optimum.

[0080] In this embodiment, after training, the inference model can learn to alternately generate images and texts, completing the multimodal reasoning process.

[0081] In summary, training with a dataset that includes both images and text in the reasoning process is done with labeling to improve training effectiveness. The reasoning process not only uses textual reasoning but also generates images, incorporating visual information into the process. This gives the reasoning model the ability to reason in a multimodal chain, preventing bias in the results of reasoning problems and improving the accuracy of the answers.

[0082] Based on the above embodiment, in this embodiment, the training process of the inference model in the preset inference task is introduced, which is detailed as follows:

[0083] S301: Collecting a data set of a preset reasoning task; wherein each data in the data set of the preset reasoning task includes a reasoning question, a reasoning process and a reasoning answer, and the reasoning process includes images and texts.

[0084] In this embodiment, each data in the dataset of the preset reasoning task also includes reasoning questions, reasoning processes, and reasoning answers. In order to achieve better training results, the dataset of the preset reasoning task also needs to be fine-tuned in a supervised manner, and the images and text in the dataset of the preset reasoning task are labeled with symbols. Among them, the symbol prompt word templates for each data in the dataset of the preset reasoning task are the same as the symbol prompt word templates in the above embodiment, but the dataset of the preset reasoning task will focus on areas such as code, geometry problems, word problems, physics problems, and path planning, which are obviously more convenient to solve using drawing. The sources are mainly teaching videos and mazes.

[0085] S302: Labeling a data set of a preset reasoning task according to the reasoning question, the reasoning process, and the reasoning answer to obtain a labeled data set of the preset reasoning task.

[0086] In this embodiment, the implementation process of step S302 is the same as that of step S202 in the above embodiment, and will not be repeated here.

[0087] S303: Encode the labeled images and texts in the data set of the preset reasoning task to obtain an encoded data set of the preset reasoning task.

[0088] In this embodiment, the implementation process of step S303 is the same as that of step S203 in the above embodiment, and will not be repeated here.

[0089] S304: Training the inference model according to the encoded data set of the preset inference task and the preset hyperparameters of the inference model to obtain the inference model of the preset inference task.

[0090] Among them, the pre-set hyperparameters of the inference model include batch parameters.

[0091] In this embodiment, the pre-set hyperparameters of the inference model also include learning rate and optimizer related parameters.

[0092] In this embodiment, the training purpose of step S304 is to enhance the performance of the reasoning model on the preset reasoning task.

[0093] In this embodiment, the inference model is loaded and a constant learning rate is used to avoid fluctuations caused by dynamic learning rate adjustments, ensuring stable convergence of the inference model. This is suitable for scenarios with limited data for the predefined inference task. An exponential moving average (EMA) is used to improve robustness and reduce the risk of overfitting by weighted averaging of historical parameters. Since the change in the loss function value during this training process was small and oscillated within a low range, the change in the loss function cannot be used to determine whether training has reached the global optimum, allowing for some overfitting to be avoided during training.

[0094] Specifically, step S304 includes S3041 to S3046:

[0095] S3041: Divide the encoded data set of the preset reasoning task into a training set and a test set.

[0096] In this embodiment, the data set of the preset reasoning task is divided into a training set and a test set according to a preset ratio. Optionally, the preset ratio may be 4:1.

[0097] S3042: Determine the number of iterative training times based on the training set and batch parameters.

[0098] Specifically, the formula for determining the number of iterative training times is as follows:

[0099]

[0100] Where, represents the number of iterative training, represents the number of data in the training set, Represents batch parameters.

[0101] In this embodiment, the batch parameter represents the amount of data required for each iterative training.

[0102] S3043: Repeatedly train the inference model using the training set according to the number of iterative training times until a preset training cycle is reached; the number of times the inference model is trained within one training cycle is the number of iterative training times.

[0103] In this embodiment, the preset training period is pre-set.

[0104] For example, the preset training cycle is 50 times, and the number of iterative training times is 10 times. The inference model is iteratively trained using the training set for 10 times as one training cycle, and the training is stopped after 50 repetitions.

[0105] S3044: Before reaching a preset training cycle, save a checkpoint every preset number of training cycles and obtain an inference model for each checkpoint.

[0106] Exemplarily, the preset number of training cycles is 5 times. During repeated training, the inference model of a checkpoint is saved every 5 training cycles, that is, when the training cycle reaches 50 times, the inference models of 10 checkpoints are saved.

[0107] In this embodiment, the purpose of step S3044 is to use the reasoning model of each checkpoint as an alternative reasoning model for the preset reasoning task, and to filter out the final reasoning model for the preset reasoning task from the alternative reasoning models through steps S3045 and S3046.

[0108] S3045: Construct a validation set based on the test set.

[0109] Specifically, from the test set, sample data with the same domain as the training set is screened out to obtain the initial validation set; from the initial validation set, sample data with unique answers and covering different difficulty levels for different fields are determined as the final validation set.

[0110] In this embodiment, the criteria for selecting the validation set from the test set are:

[0111] a. The domain of the reasoning problem for each data point has a high degree of overlap with the domain of the reasoning problem for each data point in the training set;

[0112] b. The reasoning answer for each data has a unique reasoning answer;

[0113] c. The difficulty of the reasoning questions for each data covers different difficulty levels.

[0114] In this embodiment, data that meets the above criteria is screened out from the test set as a validation set.

[0115] S3046: Use the validation set to validate the inference model of each checkpoint, select the inference model of one checkpoint, and obtain the inference model of the preset inference task.

[0116] Specifically, step S3046 includes Sa~Sc:

[0117] Sa: Input the verification set into the inference model of each checkpoint respectively. The inference model of each checkpoint outputs the inference result set of the verification set based on the verification set.

[0118] For example, the validation set is input into the inference models of 10 checkpoints respectively, and the inference models of 10 checkpoints respectively output the inference result sets of the validation set, wherein each inference result in the inference result set includes the inference process and the inference answer.

[0119] Sb: Determine whether the inference result set meets the evaluation criteria.

[0120] In this embodiment, the evaluation criteria are: whether the reasoning answer is correct; whether the correct reasoning process can be made for reasoning problems in different fields, that is, when the reasoning process requires multimodal reasoning, the multimodal reasoning process can be output; when multimodal reasoning is required, whether the image in the reasoning process is accurate and reasonable.

[0121] Sc: The inference model of the checkpoint corresponding to the inference result set that meets the evaluation criteria is determined as the inference model of the preset inference task.

[0122] In summary, on the basis of the inference model, a data set of the preset reasoning task is collected, and the data set of the preset reasoning task is marked to improve the training effect; the encoded data set of the preset reasoning task is used to conduct targeted training on the inference model to obtain the inference model of the preset reasoning task, thereby improving the reasoning accuracy of the inference model of the preset reasoning task in the target field corresponding to the preset reasoning task.

[0123] Based on the above embodiment, in this embodiment, the use process of the inference model is introduced, which is detailed as follows:

[0124] S401: receiving a question to be inferred sent by a user terminal; wherein the question to be inferred carries a system prompt word.

[0125] Optionally, the reasoning question includes images and text.

[0126] In this embodiment, the system prompt word is used to prompt the reasoning model with the type of the problem to be reasoned, such as geometry problems and physics problems.

[0127] For example, the reasoning problem is: the reasoning problem is accompanied by a schematic image, and there is a text after the schematic image: there is an object with a mass m on a right-angled splitter M placed on the horizontal ground. If m slides down M at a constant speed and M remains stationary, then how much pressure and friction does M exert on the ground?

[0128] Among them, the system prompt word for the question to be reasoned is a physics question.

[0129] S402: Input the problem to be inferred into the inference model, and the inference model performs the following steps.

[0130] Specifically, step S402 includes S4021 to S4023:

[0131] S4021: Encode the problem to be reasoned to obtain the encoded problem to be reasoned.

[0132] Specifically, an image and text in a problem to be inferred are obtained; an image encoder is used to encode the image, and a text encoder is used to encode the text to obtain the encoded problem to be inferred.

[0133] Exemplarily, an image encoder is used to encode a schematic image, and a text encoder is used to encode a text. The encoded image and the encoded text are concatenated to obtain an encoded problem to be inferred.

[0134] S4022: Obtain the question type of the question to be inferred according to the system prompt word.

[0135] For example, the system prompt word is a physics question, and the question type of the question to be reasoned is a recursive reasoning type.

[0136] In this embodiment, according to the question to be inferred, if the answer cannot be given directly and reasoning is required to obtain the reasoned answer, the question type is a recursive reasoning type, otherwise it is a text type.

[0137] S4023: If the question type is a recursive reasoning type, then output the reasoning process and reasoning answer of the question to be reasoned based on the encoded question to be reasoned.

[0138] Specifically, step S4023 includes Sd~Sf:

[0139] Sd: Perform reasoning based on the encoded question to be reasoned, and recursively generate reasoning images and reasoning text.

[0140] Se: In the current reasoning process, based on the last generated reasoning image and reasoning text, a new reasoning image and new reasoning text are generated until the reasoning answer is generated.

[0141] Sf: Outputs the reasoning process and reasoning answer of the problem to be reasoned.

[0142] In this embodiment, reasoning is performed based on the encoded problem to be reasoned, and a reasoning image and reasoning text of a preliminary analysis are generated. Based on the reasoning image and reasoning text generated by the preliminary analysis, analysis is performed again to generate a reasoning image and reasoning text of a second analysis. In this way, reasoning images and reasoning texts are recursively generated until the reasoning process and reasoning answer are generated.

[0143] Optionally, if the question type is a text type, the reasoning answer to the question to be reasoned is output based on the encoded question to be reasoned.

[0144] S403: Output the reasoning process and the reasoning answer; wherein the reasoning process includes images and texts.

[0145] Optionally, the method provided in the embodiments of this application can also be used to generate course explanation videos to help students more intuitively understand the problem-solving steps. It can also be combined with autonomous driving data to perform traffic scene understanding, navigation, and path planning.

[0146] In summary, the reasoning model determines the type of question to be reasoned based on the system prompt. If the question is of the recursive reasoning type, it performs recursive reasoning and outputs the reasoning process and answer, which includes images and text. The reasoning process considers multimodal chain thinking, not only using textual reasoning but also generating images, incorporating visual information into the reasoning process. This prevents bias in the reasoning problem processing and improves the accuracy of the reasoning answer. Furthermore, the recursive reasoning process continues based on the previously generated images and text, recursively generating reasoning images and text until the reasoning process and answer are generated, avoiding bias in the reasoning process and further improving the accuracy of the reasoning answer.

[0147] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0148] Figure 4 This is a structural diagram of the training device for the inference model provided in the embodiment of the present application. Figure 4 As shown, an embodiment of the present application further provides a training device for an inference model, including: an acquisition module 401 , a marking module 402 , an encoding module 403 and a training module 404 .

[0149] Acquisition module 401, for acquiring a data set for training; wherein each data in the data set includes reasoning questions, reasoning processes, and reasoning answers, and the reasoning process includes images and text;

[0150] A labeling module 402 is used to label the data set according to the reasoning question, the reasoning process, and the reasoning answer to obtain a labeled data set;

[0151] An encoding module 403 is used to encode the images and texts in the labeled dataset to obtain an encoded dataset;

[0152] The training module 404 is used to train the multimodal pre-training model based on the encoded data set and the pre-set hyperparameters of the multimodal pre-training model to obtain an inference model; wherein the inference model is used to output the inference process and inference answer according to the problem to be inferred.

[0153] In a possible implementation, the reasoning question carries a system prompt word; accordingly, the marking module 402 includes:

[0154] A first marking submodule is used to add a question role mark at the beginning of the reasoning question and an answer role mark at the beginning of the reasoning process for each data in the data set, thereby obtaining a first marked data set;

[0155] A second marking submodule is used to add an inference start marker at the beginning of the inference process and an inference end marker at the end of the inference process for each data in the first marked data set, so as to obtain a second marked data set;

[0156] A third marking submodule is configured to add an image start marker at the beginning of each image and an image end marker at the end of each image for each data in the second marked data set, thereby obtaining a third marked data set;

[0157] The fourth marking submodule is used to add a data start mark and a system prompt word at the beginning of the question role mark, and add a data end mark at the end of the inference answer for each data in the third marked data set to obtain the final marked data set.

[0158] In a possible implementation, the encoding module 403 is specifically configured to encode the image using an image encoder and encode the text using a text encoder for each data in the labeled data set to obtain an encoded data set.

[0159] In one possible implementation, the pre-set hyperparameters of the multimodal pre-training model include a minimum number of iterations, a maximum number of iterations, and a loss function change threshold; accordingly, the training module 404 includes:

[0160] Loading submodule, used to load multimodal pre-training models;

[0161] A first training submodule is used to iteratively train the multimodal pre-training model based on the encoded data set;

[0162] The first judgment submodule is used to continue iterative training if the current iteration number is greater than the minimum iteration number and the loss function change value is greater than the loss function change threshold during training;

[0163] The second judgment submodule is used to stop iteration and obtain an inference model if the current number of iterations exceeds the maximum number of iterations, or / and the loss function change value is less than the loss function change threshold.

[0164] In a possible implementation, the training device for the inference model further includes an acquisition module, which includes:

[0165] The acquisition submodule is used to collect the data set of the preset reasoning task; wherein the data in the data set of the preset reasoning task includes the reasoning question, the reasoning process and the reasoning answer, and the reasoning process includes images and text;

[0166] A fifth labeling submodule is used to label the data set of the preset reasoning task according to the reasoning question, the reasoning process and the reasoning answer, so as to obtain the labeled data set of the preset reasoning task;

[0167] An encoding submodule, configured to encode the labeled images and text in the dataset of the preset reasoning task to obtain the encoded dataset of the preset reasoning task;

[0168] The second training submodule is used to train the inference model according to the encoded data set of the preset inference task and the preset hyperparameters of the inference model to obtain the inference model of the preset inference task.

[0169] In one possible implementation, the preset hyperparameters of the inference model include batch parameters; accordingly, the second training submodule includes:

[0170] A partitioning unit, used to partition the encoded data set of the preset reasoning task into a training set and a test set;

[0171] A determination unit, used to determine the number of iterative training times based on the training set and batch parameters;

[0172] A repeated training unit is used to repeatedly train the inference model using the training set according to the number of iterative training times until a preset training cycle is reached; the number of times the inference model is trained in one training cycle is the number of iterative training times;

[0173] The first acquisition unit is configured to save a checkpoint every preset number of training cycles before reaching a preset training cycle, and obtain an inference model for each checkpoint;

[0174] Construction unit, used to construct a validation set based on the test set;

[0175] The screening unit is used to use the verification set to verify the inference model of each checkpoint, screen out the inference model of a checkpoint, and obtain the inference model of the preset inference task.

[0176] In a possible implementation, the formula for determining the unit is:

[0177]

[0178] Where, represents the number of iterative training, represents the number of data in the training set, Represents batch parameters.

[0179] In one possible embodiment, the construction unit includes:

[0180] The screening subunit is used to screen out sample data with the same domain as the training set from the test set to obtain the initial validation set;

[0181] The first determination subunit is used to determine, from the initial verification set, sample data with unique answers and covering different difficulty levels for different fields as the final verification set.

[0182] In one possible embodiment, the screening unit includes:

[0183] The first output subunit is used to input the verification set into the inference model of each checkpoint respectively, and the inference model of each checkpoint outputs the inference result set of the verification set based on the verification set;

[0184] The judgment subunit is used to judge whether the inference result set meets the evaluation criteria;

[0185] The second determining subunit is configured to determine the inference model of the checkpoint corresponding to the inference result set that meets the evaluation criteria as the inference model of the preset inference task.

[0186] In a possible implementation, the training device for the inference model further includes an application module, and the application module includes:

[0187] The receiving submodule is used to receive the question to be inferred sent by the user end; wherein the question to be inferred carries the system prompt word;

[0188] The execution submodule is used to input the problem to be inferred into the inference model, and the inference model performs the following steps.

[0189] The execution submodule includes:

[0190] An encoding unit, used for encoding the problem to be reasoned to obtain the encoded problem to be reasoned;

[0191] The second acquisition unit is used to acquire the question type of the question to be inferred according to the system prompt word;

[0192] A first output unit is configured to output the reasoning process and the reasoning answer of the problem to be reasoned according to the encoded problem to be reasoned if the problem type is a recursive reasoning type;

[0193] The output submodule is used to output the reasoning process and reasoning answers; the reasoning process includes images and text.

[0194] In a possible implementation, the first output unit includes:

[0195] A first generating subunit is used to perform reasoning based on the encoded problem to be reasoned, and recursively generate a reasoning image and a reasoning text;

[0196] The second generating subunit is configured to generate a new reasoning image and a new reasoning text based on the reasoning image and reasoning text generated last time during the current reasoning process, until a reasoning answer is generated;

[0197] The second output subunit is used to output the reasoning process and reasoning answer of the problem to be reasoned.

[0198] In a possible implementation, the execution submodule further includes a second output unit, and the second output unit is specifically configured to output an inference answer to the question to be inferred based on the encoded question to be inferred if the question type is a text type.

[0199] For the description of the features in the embodiment corresponding to the training device of the reasoning model, please refer to the relevant description of the embodiment corresponding to the training method of the reasoning model, and no further details will be given here.

[0200] Figure 5 This is a schematic diagram of the structure of the server provided in the embodiment of the present application. Figure 5 As shown, the server provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the server also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.

[0201] During the specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned training method embodiment of the inference model.

[0202] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0203] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0204] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0205] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0206] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned training method embodiments of the inference model when running.

[0207] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0208] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the training method embodiment of any of the above-mentioned inference models are implemented.

[0209] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in the training method embodiment of any of the above-mentioned inference models.

[0210] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0211] The above is a detailed introduction to the training method, device, server, storage medium and product of an inference model provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of ​​this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A method for training an inference model, characterized in that: include: Collecting datasets for training; Each data in the data set includes a reasoning question, a reasoning process, and a reasoning answer, and the reasoning process includes images and texts; Labeling the data set according to the reasoning question, the reasoning process, and the reasoning answer to obtain a labeled data set; Encoding the image and the text in the labeled data set to obtain an encoded data set; The multimodal pre-training model is trained according to the encoded data set and the pre-set hyperparameters of the multimodal pre-training model to obtain an inference model; wherein the inference model is used to output an inference process and an inference answer according to the problem to be inferred; in the inference process, not only textual reasoning is used, but also images are generated, so that visual information is added to the inference process.

2. The method according to claim 1, characterized in that The reasoning question carries a system prompt word; Accordingly, the data set is labeled according to the reasoning question, the reasoning process, and the reasoning answer to obtain the labeled data set, including: For each data in the data set, a questioning role tag is added at the beginning of the reasoning question, and an answering role tag is added at the beginning of the reasoning process to obtain a first-tagged data set; For each data in the first labeled data set, adding an inference start marker at the beginning of the inference process and adding an inference end marker at the end of the inference process to obtain a second labeled data set; For each data in the second marked data set, adding an image start marker at the beginning of each image and adding an image end marker at the end of each image to obtain a third marked data set; For each data in the third marked data set, a data start mark and the system prompt word are added at the beginning of the question role mark, and a data end mark is added at the end of the inference answer to obtain the final marked data set.

3. The method according to claim 1, characterized in that The step of encoding the image and the text in the labeled data set to obtain an encoded data set includes: For each data in the labeled data set, an image encoder is used to encode the image, and a text encoder is used to encode the text to obtain an encoded data set.

4. The method according to claim 1, wherein The pre-set hyperparameters of the multimodal pre-training model include a minimum number of iterations, a maximum number of iterations, and a loss function change threshold; Accordingly, the multimodal pre-training model is trained according to the encoded data set and the pre-set hyperparameters of the multimodal pre-training model to obtain an inference model, including: Load multimodal pre-trained model; Iteratively training the multimodal pre-training model according to the encoded data set; During the training process, if the current iteration number is greater than the minimum iteration number and the loss function change value is greater than the loss function change threshold, then continue iterative training; If the current number of iterations exceeds the maximum number of iterations, or / and the loss function change value is less than the loss function change threshold, the iteration is stopped to obtain the inference model.

5. The method according to claim 1, wherein After obtaining the inference model, the method further includes: Collecting a data set for a preset reasoning task; wherein each data in the data set for the preset reasoning task includes a reasoning question, a reasoning process, and a reasoning answer, and the reasoning process includes images and text; Marking the data set of the preset reasoning task according to the reasoning question, the reasoning process and the reasoning answer to obtain a marked data set of the preset reasoning task; Encoding the image and the text in the marked data set of the preset reasoning task to obtain an encoded data set of the preset reasoning task; The reasoning model is trained according to the encoded data set of the preset reasoning task and the preset hyperparameters of the reasoning model to obtain the reasoning model of the preset reasoning task.

6. The method according to claim 5, characterized in that The pre-set hyperparameters of the inference model include batch parameters; Accordingly, the inference model is trained based on the encoded data set of the preset inference task and the preset hyperparameters of the inference model to obtain the inference model of the preset inference task, including: Dividing the encoded data set of the preset reasoning task into a training set and a test set; Determining the number of iterative training times according to the training set and the batch parameters; Repeatedly training the inference model using the training set according to the number of iterative training times until a preset training cycle is reached; the number of times the inference model is trained within one training cycle is the number of iterative training times; Before reaching the preset training cycle, saving a checkpoint every preset number of training cycles and obtaining an inference model for each checkpoint; Construct a validation set based on the test set; The verification set is used to verify the reasoning model of each checkpoint, and the reasoning model of a checkpoint is screened out to obtain the reasoning model of the preset reasoning task.

7. The method according to claim 6, characterized in that The formula for determining the number of iterative training times based on the training set and the batch parameters is: Where, represents the number of iterative training, represents the number of data in the training set, Represents batch parameters.

8. The method according to claim 6, characterized in that The step of constructing a validation set based on the test set includes: From the test set, sample data with the same domain as the training set is selected to obtain an initial validation set; From the initial validation set, sample data with unique answers and covering different difficulty levels for different fields are determined as the final validation set.

9. The method according to claim 6, characterized in that The use of the verification set to verify the reasoning model of each checkpoint, screening out the reasoning model of a checkpoint, and obtaining the reasoning model of the preset reasoning task includes: Inputting the verification set into the inference model of each checkpoint respectively, and the inference model of each checkpoint outputs the inference result set of the verification set based on the verification set; Determining whether the inference result set meets the evaluation criteria; The reasoning model of the checkpoint corresponding to the reasoning result set that meets the evaluation criteria is determined as the reasoning model of the preset reasoning task.

10. The method according to claim 1, characterized in that After obtaining the inference model, the method further includes: Receive a question to be inferred sent by a user terminal; wherein the question to be inferred carries a system prompt word; The problem to be inferred is input into the inference model, and the inference model performs the following steps: Encoding the problem to be inferred to obtain an encoded problem to be inferred; Obtaining the question type of the question to be inferred according to the system prompt word; If the question type is a recursive reasoning type, then outputting the reasoning process and reasoning answer of the question to be reasoned according to the encoded question to be reasoned; The reasoning process and the reasoning answer are output; wherein the reasoning process includes an image and a text.

11. The method according to claim 10, characterized in that Outputting the reasoning process and the reasoning answer of the problem to be inferred according to the encoded problem to be inferred includes: Performing reasoning based on the encoded problem to be reasoned, and recursively generating a reasoning image and a reasoning text; In the current reasoning process, based on the reasoning image and the reasoning text generated last time, a new reasoning image and a new reasoning text are generated until a reasoning answer is generated; Output the reasoning process of the problem to be reasoned and the reasoning answer.

12. The method according to claim 10, characterized in that After obtaining the question type of the question to be inferred according to the system prompt word, the method further includes: If the question type is a text type, then the reasoning answer to the question to be reasoned is output according to the encoded question to be reasoned.

13. A server, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for training an inference model as described in any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the training method of the inference model according to any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the training method of the inference model as described in any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Model training method and device, equipment, storage medium and program product

    CN118378633A

  • Model training method and device, equipment, storage medium and product

    CN120218245A