Model training method and device, equipment, storage medium and product

By constructing the target text reasoning chain and performing supervised fine-tuning and reinforcement learning training, the problem of multimodal large language model performing poorly in complex visual reasoning tasks is solved, and the effect of the model showing the correct thinking process in the reasoning process is achieved.

CN120218245APending Publication Date: 2025-06-27SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 19 Cited by

Patent Information

Application Number
CN202510306547.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing multimodal large language models perform poorly when dealing with complex visual inference tasks and lack explicit intermediate inference processes.

Method used

By constructing a training data set, it contains the target text inference chain obtained by multimodal data transformation, supervised fine-tuning and long-thinking reinforcement learning training, and optimizes the multimodal large language model to output the target answers containing the inference process.

Benefits of technology

The reasoning ability and generalization ability of multimodal large language models in complex visual reasoning tasks has been significantly improved, so that the model can demonstrate the correct thinking process in the reasoning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218245A_ABST
    Figure CN120218245A_ABST
Patent Text Reader

Abstract

The invention relates to a model training method and device, equipment, a storage medium and a product. The method comprises the following steps: constructing a training data set according to a target text reasoning chain obtained by converting multi-modal data; according to the training data set, performing supervision fine tuning on the pre-trained multi-modal large language model to obtain a basic reasoning model; performing optimization processing on the basic reasoning model according to reinforcement learning training of long thinking to obtain a target reasoning model; the target reasoning model is used for outputting a target answer containing a reasoning process according to the input multi-modal data. Therefore, the long text constraint can be directly used for reinforcement learning, and the training efficiency is greatly improved; and by adopting long-thinking reinforcement learning training, the model can easily learn a correct thinking process in training, so that the reasoning ability of the multi-modal large language model for processing a complex visual reasoning task is improved, and the correct thinking process is displayed in the reasoning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a model training method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of artificial intelligence (AI) technology, multimodal large language models (MLLMs) have shown potential in image understanding and visual reasoning tasks. A multimodal large language model refers to a large artificial intelligence model that can process and understand multiple different input modalities (such as text, images, audio, etc.).

[0003] In traditional technologies, the inference paradigm of multimodal large language models usually relies on a simple "direct prediction" method to generate concise final answers, lacking an explicit intermediate reasoning process, resulting in poor performance in dealing with complex visual reasoning tasks (such as mathematical problem solving, multi-step logical analysis). Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a model training method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the inference ability of multimodal large language models in dealing with complex visual reasoning tasks and show the correct thinking process during the inference process.

[0005] In a first aspect, this application provides a model training method, and the method includes:

[0006] Construct a training data set according to a target text inference chain obtained by converting multimodal data. The target text inference chain includes: text description information, an input question, an inference process, and a target answer; the text description information includes: text information obtained when describing the multimodal data in text form; the inference process includes: each step of inferring the target answer according to the text description information and the input question;

[0007] Perform supervised fine-tuning on a pre-trained multimodal large language model according to the training data set to obtain a basic inference model;

[0008] Optimize the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model; the target inference model is used to output a target answer including an inference process according to the input multimodal data.

[0009] In one embodiment, constructing a training dataset based on the target text inference chain obtained by converting multimodal data includes:

[0010] Converting the image data in the multimodal data into text description information containing visual details;

[0011] Processing the text description information containing visual details through a text inference model to obtain the target text inference chain;

[0012] Constructing the training dataset based on the target text inference chain.

[0013] In one embodiment, converting the image data in the multimodal data into text description information containing visual details includes:

[0014] Obtaining multimodal data, where the multimodal data includes: image data and a question associated with the image data;

[0015] Inputting a target instruction containing the multimodal data into a target multimodal large language model for processing to obtain text description information containing visual details; the text description information containing visual details is used to describe visual features and / or visual elements captured from the image data.

[0016] In one embodiment, processing the text description information containing visual details through a text inference model to obtain the target text inference chain includes:

[0017] Inputting the text description information containing visual details and the question associated with the image data into the text inference model to output an initial text inference chain; evaluating each inference step in the initial text inference chain through a first reward model to obtain an inference score corresponding to the initial text inference chain;

[0018] Filtering the initial text inference chain according to the inference score;

[0019] Performing clustering processing on the filtered initial text inference chain to obtain the target text inference chain.

[0020] In one embodiment, performing clustering processing on the filtered initial text inference chain to obtain the target text inference chain includes:

[0021] Performing clustering processing on the filtered initial text inference chain through a preset clustering algorithm to obtain text inference chains of several categories;

[0022] Randomly extract a preset number of text inference chains from the text inference chains of the several categories as the target text inference chains.

[0023] In one embodiment, the method of performing supervised fine-tuning on the pre-trained multi-modal large language model according to the training data set to obtain a basic inference model includes:

[0024] Train the pre-trained multi-modal large language model based on the training data set, and during the training process, update at least part of the parameters in the pre-trained multi-modal large language model or update the parameters of the plug-in network corresponding to the pre-trained multi-modal large language model to obtain a basic inference model.

[0025] In one embodiment, the training of the pre-trained multi-modal large language model based on the training data set includes:

[0026] Select general data from the training data set to perform optimization training on the pre-trained multi-modal large language model; and / or

[0027] Screen out target domain data from the training data set, and perform optimization training on the pre-trained multi-modal large language model through an optimization data set composed of the target domain data and at least part of the general data; the proportion of the target domain data in the optimization data set is greater than the proportion of the target domain data in the training data set;

[0028] Wherein, the general data refers to sample data containing multiple domains randomly extracted from the training data set.

[0029] In one embodiment, the method of optimizing the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model includes:

[0030] Limit the output of the basic inference model according to a preset text length, train the basic inference model in combination with a reinforcement learning algorithm, and guide the correctness of the inference long chain through a result reward mechanism during the training process to obtain a target inference model.

[0031] In one embodiment, it further includes:

[0032] Establish a result reward mechanism according to format standardization and answer correctness; the format standardization includes: the correctness of mathematical symbols and the integrity of step numbers; wherein, rewards are given in the case of format standardization and / or correct answers; no rewards are given in the case of non-standard format and incorrect answers.

[0033] In one of the embodiments, optimizing the basic inference model according to the reinforcement learning training with long thinking to obtain a target inference model includes:

[0034] Limiting the output of the basic inference model according to a preset text length, training the basic inference model in combination with a reinforcement learning algorithm, and evaluating each step of the inference long chain through a second reward model during the training process to obtain an evaluation result;

[0035] Guiding the correctness of the inference long chain according to the evaluation result to obtain a target inference model.

[0036] Second, the present application also provides a data processing method, and the method includes:

[0037] Responding to multimodal data input on the display interface, the multimodal data including: image data and a question associated with the image data;

[0038] Inputting the multimodal data into the target inference model for processing, and outputting a target answer including the inference process; the target inference model is trained according to the model training method described in any item of the first aspect.

[0039] Third, the present application also provides a model training device, and the device includes:

[0040] A dataset construction module for constructing a training dataset according to a target text inference chain obtained by converting multimodal data, the target text inference chain including: text description information, an input question, an inference process, and a target answer; the text description information includes: text information obtained when describing the multimodal data in text form; the inference process includes: each step of inferring the target answer according to the text description information and the input question;

[0041] A fine-tuning module for performing supervised fine-tuning on a pre-trained multimodal large language model according to the training dataset to obtain a basic inference model;

[0042] An optimization training module for optimizing the basic inference model according to the reinforcement learning training with long thinking to obtain a target inference model; the target inference model is used to output a target answer including the inference process according to the input multimodal data.

[0043] Fourth, the present application also provides a data processing device, and the device includes:

[0044] A receiving module for responding to multimodal data input on the display interface, the multimodal data including: image data and a question associated with the image data;

[0045] A processing module, configured to input the multimodal data into a target inference model for processing and output a target answer including an inference process; the target inference model is trained according to the model inference ability training method described in any item of the first aspect.

[0046] In a fifth aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0047] Construct a training data set according to a target text inference chain obtained by converting multimodal data. The target text inference chain includes: text description information, an input question, an inference process, and a target answer. The text description information includes: text information obtained when describing the multimodal data in text form. The inference process includes: each step of inferring the target answer according to the text description information and the input question. According to the training data set, perform supervised fine-tuning on a pre-trained multimodal large language model to obtain a basic inference model. Perform optimization processing on the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model. The target inference model is used to output a target answer including an inference process according to the input multimodal data; or,

[0048] In response to multimodal data input on a display interface, the multimodal data includes: the image data and a question associated with the image data. Input the multimodal data into a target inference model for processing and output a target answer including an inference process. The target inference model is trained according to the model training method described in any item of the first aspect.

[0049] In a sixth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0050] Construct a training data set according to a target text inference chain obtained by converting multimodal data. The target text inference chain includes: text description information, an input question, an inference process, and a target answer. The text description information includes: text information obtained when describing the multimodal data in text form. The inference process includes: each step of inferring the target answer according to the text description information and the input question. According to the training data set, perform supervised fine-tuning on a pre-trained multimodal large language model to obtain a basic inference model. Perform optimization processing on the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model. The target inference model is used to output a target answer including an inference process according to the input multimodal data; or,

[0051] In response to multimodal data input on a display interface, the multimodal data includes: the image data and a question associated with the image data; inputting the multimodal data into a target inference model for processing, and outputting a target answer including an inference process; the target inference model is trained according to the model training method described in any one of the first aspects.

[0052] In a fifth aspect, the present application further provides a computer program product, including a computer program, which when executed by a processor, implements the following steps:

[0053] Construct a training data set according to a target text inference chain obtained by converting multimodal data, the target text inference chain includes: text description information, an input question, an inference process, and a target answer; the text description information includes: text information obtained when describing the multimodal data in text form; the inference process includes: each step of inferring the target answer according to the text description information and the input question; according to the training data set, perform supervised fine-tuning on a pre-trained multimodal large language model to obtain a basic inference model; optimize the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model; the target inference model is used to output a target answer including an inference process according to the input multimodal data; or,

[0054] In response to multimodal data input on a display interface, the multimodal data includes: the image data and a question associated with the image data; inputting the multimodal data into a target inference model for processing, and outputting a target answer including an inference process; the target inference model is trained according to the model training method described in any one of the first aspects.

[0055] The above-mentioned method, device, computer device, computer-readable storage medium, and computer program product for training the model inference ability construct a training data set through the target text inference chain obtained by converting multi-modal data; thus, it can efficiently and automatically generate high-quality target text inference chains, which is convenient for subsequent supervised fine-tuning of the multi-modal large language model and optimization training of the basic inference model, significantly improving the training effect. According to the training data set, the pre-trained multi-modal large language model is supervised and fine-tuned to obtain a basic inference model; thus, the inference ability and generalization ability of the basic inference model can be improved. The basic inference model is optimized according to the reinforcement learning training of long thinking to obtain a target inference model; the target inference model is used to output a target answer including the inference process according to the input multi-modal data. Thus, long text constraints can be directly used for reinforcement learning, greatly improving the training efficiency; and the reinforcement learning training with long thinking enables the model to easily learn the correct thinking process during training, so as to improve the inference ability of the multi-modal large language model to handle complex visual inference tasks and show the correct thinking process during the inference process. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description in the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0057] Figure 1 It is an application environment diagram of the model training method in an embodiment;

[0058] Figure 2 It is a schematic flowchart of the model training method provided in the first embodiment;

[0059] Figure 3 It is a schematic flowchart of step 201 in an embodiment;

[0060] Figure 4 It is a schematic flowchart of the model training method provided in the second embodiment;

[0061] Figure 5 It is a schematic flowchart of the model training method provided in the third embodiment;

[0062] Figure 6 It is a schematic flowchart of the data processing method provided in an embodiment;

[0063] Figure 7A It is a schematic diagram of the display interface of the target inference model applied in the terminal in an embodimentFigure 1 ;

[0064] Figure 7B Schematic diagram of the display interface of the target inference model provided in an embodiment for use on a terminal Figure 2 ;

[0065] Figure 7C Schematic diagram of the display interface of the target inference model provided in an embodiment for use on a terminal Figure 3 ;

[0066] Figure 7D Timing schematic diagram of the target inference model data processing method provided in the first embodiment;

[0067] Figure 8A Schematic diagram of the display interface of the target inference model provided in an embodiment for use on a terminal Figure 4 ;

[0068] Figure 8B Schematic diagram of the display interface of the target inference model provided in an embodiment for use on a terminal Figure 5 ;

[0069] Figure 8C Schematic diagram of the display interface of the target inference model provided in an embodiment for use on a terminal Figure 6 ;

[0070] Figure 8D Timing schematic diagram of the target inference model data processing method provided in the second embodiment;

[0071] Figure 9 Structural block diagram of the model training device in an embodiment;

[0072] Figure 10 Structural block diagram of the model training device in another embodiment;

[0073] Figure 11 Structural block diagram of the data processing device in an embodiment;

[0074] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0075] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0076] To facilitate the understanding of the technical solutions described in the embodiments of the present application, some terms that may appear in the embodiments of the present application will be explained first.

[0077] Exemplarily, the term "in response to" involved in the present application represents a state where a corresponding event occurs or a condition is satisfied. The execution timing of subsequent actions executed in response to the event or condition and the time when the event occurs or the condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent action can be immediately executed when the event occurs or the condition is established; while in other cases, the subsequent action can be executed after a period of time after the event occurs or the condition is established.

[0078] Exemplarily, the term "multimodal data" involved in the present application refers to data of the same description object obtained from different perspectives, which can include forms such as text, images, audio, and video. The core of multimodal data lies in improving the accuracy of prediction and understanding by integrating information of multiple modalities.

[0079] Exemplarily, the term "training data set" involved in the present application is used to solve the cold start problem in machine learning and recommendation systems. The core advantage of the training data set is that it can significantly improve the readability of the model and accelerate the convergence process of the model. For example, common ways to construct a training data set include: manual annotation and format filtering. The training data set can help the model learn how to write steps and summaries in a standardized manner, thereby improving the user-friendliness and correctness of the generated content.

[0080] Exemplarily, the term "text inference chain" involved in the present application is a method to reach a conclusion through a series of logical reasoning steps, which is widely used in multiple fields, including: the initial description and background information of the input problem, intermediate reasoning steps (intermediate conclusions or hypotheses generated by the model during the processing), and the final conclusion or answer. For example, in terms of multimodal interaction, the text inference chain improves the model's ability to integrate and understand information by combining different modal data such as text, images, tables, code, and language.

[0081] Exemplarily, the term "touch operation" involved in the present application refers to an operation performed on the information provided by a computer device (such as a control), through which corresponding instructions can be sent to the computer device to trigger the computer device to execute the next task. The next tasks triggered by performing different touch operations on different information can all be preset in the program. The touch operation can be manually executed by the user. For example, operations such as clicking, double-clicking, long-pressing, and swiping performed by the user on the content displayed on the screen of the computer device all belong to touch operations. In some cases, the touch operation can also be executed by the computer device based on a set program.

[0082] The model training method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. Model training is completed on the terminal 102 or the server 104 to obtain a target inference model that can output the inference process and the target answer. Optionally, when the target inference model is loaded on the terminal 102 for use, it can support offline multi-modal data question answering. When the target inference model is loaded on the server 104, it can support online multi-modal data question answering. Exemplarily, taking the model inference training on the server 104 as an example, first, a training data set is constructed according to the target text inference chain obtained by converting multi-modal data. The target text inference chain includes: text description information, the input question, the inference process, and the target answer; the text description information includes: the text information obtained when describing multi-modal data in text form; the inference process includes: each step of inferring the target answer according to the text description information and the input question; according to the training data set, the pre-trained multi-modal large language model is supervised and fine-tuned to obtain a basic inference model; the basic inference model is optimized according to the reinforcement learning training of long thinking to obtain the target inference model; the target inference model is used to output the target answer including the inference process according to the input multi-modal data. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0083] In an exemplary embodiment, as Figure 2 shown, the first model training method is provided. Taking the method applied to Figure 1 the terminal or server in

[0084] Step S201, construct a training data set according to the target text inference chain obtained by converting multi-modal data.

[0085] In this embodiment, the multimodal data can be a combination of an image and text, or a combination of an image and speech. Among them, the combination of an image and text can be understood as: an image and a corresponding question associated with the image. The combination of an image and speech can be understood as: an image and a corresponding speech associated with the image (the speech content can be one or more questions).

[0086] Optionally, when the multimodal data is a combination of an image and speech, speech recognition technology can be used to first process the speech into corresponding text, and then multiple combinations of images and text are obtained.

[0087] Furthermore, after obtaining the multimodal data of these combinations of images and text, these multimodal data are converted to obtain target text inference chains. Finally, a training data set is constructed based on these target text inference chains. Among them, the target text inference chain includes: text description information, input questions, the inference process, and the target answer; the text description information includes: the text information obtained when describing the multimodal data in text form; the inference process includes: each step of inferring the target answer based on the text description information and the input questions.

[0088] It should be noted that the target text inference chain in this embodiment is essentially different from the existing Chain-of-Thought (CoT). Among them, the Chain-of-Thought refers to: constructing a data set containing structured inference steps through manual annotation. Such manually designed inference steps often lack key cognitive processes in human thinking (such as questioning, reflection, verification), resulting in the generated inference chain showing the characteristics of "pseudo Chain-of-Thought". For example, in visual math problems, the model may mechanically list formulas while ignoring the in-depth association of key information in the image (such as the spatial relationship or numerical feature extraction of objects in the image). This pseudo-CoT cannot effectively capture the cross-modal information fusion required for complex reasoning, limiting the practical application ability of the model.

[0089] Therefore, in order to enable the model to give a text inference chain that conforms to the human thinking mode (thinking process), a series of conversion processes need to be performed on the multimodal data. Optionally, Modality Bridging technology can be used to convert visual information into text description information rich in semantics and generate multimodal CoT data in the style of human cognition.

[0090] Exemplarily, the multimodal data is divided into two parts: an image and a question, where the question is a text description surrounding the image. For example, an image containing various flowers of different colors. The question could be: How many colors of flowers are there in the picture? Further, after splicing the image and the question according to a preset prompt template, they are input into the same Multimodal Large Language Models (MLLMs), and the MLLMs outputs text description information rich in visual details.

[0091] Optionally, the preset prompt template can be a prompt entered by the user or some preset templates built into the background. For example: "Provide a detailed description that includes all necessary visual details to answer the question".

[0092] It should be understood that the specific content and form of the preset prompt template in this embodiment are not limited, and it can be set according to different field requirements. Taking the mathematics field as an example, its prompt can include: obtaining all mathematical symbols, numerical values, and step numbers in the image, etc.

[0093] Further, after obtaining the text description information rich in visual details, a pure text reasoning model can be used to perform self-reflection and verification on this text description information to obtain a large number of text reasoning chains. Secondly, in order to improve the quality of these text reasoning chains, these text reasoning chains can be further filtered and screened to obtain the target text reasoning chains. Among them, the pure text reasoning model refers to those models that focus on the processing and reasoning of text data. They can understand and analyze the information in the text and then make corresponding judgments or decisions. Such models play an important role in the field of Natural Language Processing (NLP) and are widely used in various application scenarios such as text classification, sentiment analysis, and question answering systems.

[0094] Among them, the training dataset is mainly used to solve the problems in the initial stage of machine learning and recommendation systems, that is, in the case of lack of sufficient historical data, effectively carry out model training and recommendation. Especially in deep learning models, the training dataset is crucial for improving the readability of the model, accelerating the convergence process, and the final performance. In this embodiment, the training dataset constructed based on the target text reasoning chain can seamlessly convert visual information into text reasoning chains, avoiding the reasoning deviation caused by modal differences in traditional methods. And it can significantly reduce the manual annotation cost, efficiently and automatically generate high-quality text reasoning chains, thereby improving the convergence speed and effect of subsequent model training.

[0095] Step S202, according to the training dataset, perform supervised fine-tuning on the pre-trained multimodal large language model to obtain a basic reasoning model.

[0096] In this embodiment, the pre-trained multi-modal large language model can be Qwen2.5-VL-7B. The Qwen2.5-VL-7B is supervised and fine-tuned based on the training data set obtained in the foregoing step S201, so that the fine-tuned model has preliminary complex reasoning ability.

[0097] Among them, Supervised Fine-Tuning (SFT) is an important technology in the fields of machine learning and natural language processing (NLP). It further trains on the basis of a pre-trained model using a labeled data set to improve the performance of the model on specific tasks or in specific fields. The main purpose of SFT is to enable the pre-trained model to better adapt to the requirements of specific tasks, thereby improving its effect in practical applications.

[0098] Exemplarily, the pre-trained multi-modal large language model is trained based on the training data set, and during the training process, at least some parameters in the pre-trained multi-modal large language model or the parameters of the plug-in network corresponding to the pre-trained multi-modal large language model are updated to obtain a basic reasoning model.

[0099] Among them, full-parameter fine-tuning means adjusting all the parameters of the model (the training effect is better); updating the parameters of the plug-in network corresponding to the pre-trained multi-modal large language model (such as LoRA fine-tuning) means introducing a module adapted to some model layers and adjusting the parameters of this module. In specific applications, the two methods can be combined.

[0100] Exemplarily, general data can be selected from the training data set to optimize and train the pre-trained multi-modal large language model.

[0101] Optionally, general data can be selected from the training data set to perform annealing training on the pre-trained multi-modal large language model.

[0102] It should be understood that in model training, "annealing" mainly refers to Learning Rate Annealing or Simulated Annealing. Learning Rate Annealing is a strategy of gradually reducing the learning rate during training, similar to the annealing process in physics, that is, gradually reducing the temperature of the system so that the system can reach a more stable state with lower energy. Adding high-quality data during the annealing stage, due to the lower learning rate, the model can learn the data features more carefully, thereby achieving more sufficient learning.

[0103] Exemplarily, target domain data can be screened out from the training data set, and the pre-trained multimodal large language model can be optimized and trained using an optimized data set consisting of the target domain data and at least part of the general data; the proportion of the target domain data in the optimized data set is greater than the proportion of the target domain data in the training data set.

[0104] Among them, general data refers to sample data containing multiple fields randomly extracted from the training data set.

[0105] Optionally, the target domain data can be screened out from the training data set, and the pre-trained multimodal large language model can be annealed using an optimized data set consisting of the target domain data and at least a portion of the general data.

[0106] For example, annealing training is performed to increase the proportion of data in a specific field that needs to be strengthened (for example, mathematical geometry, etc.), and ultimately to achieve the goal of improving the capabilities of the specific field to be strengthened as much as possible while maintaining general capabilities. Among them, the probability of the proportion of data in a specific field refers to changing the proportion of data in a specific field in the total training data. Assuming that there are a total of 100,000 training data, of which 5,000 are mathematical geometry data, therefore, 30,000 data can be screened out from the 95,000 data other than mathematical geometry, and the 30,000 data and 5,000 data are combined to form a new training set, and then this new training set is used for annealing training (thereby increasing the proportion of mathematical geometry data in the training data).

[0107] It should be understood that the above two annealing training methods can be used at the same time to obtain a model that is universal and has strong reasoning ability for a specific field. Optionally, a large amount of general data can be used in the early stage of pre-training, and high-quality small data sets can be added in the annealing stage, so as to prevent the small data sets from being overused in long-term training, avoiding overfitting and other negative effects.

[0108] Step S203, optimizing the basic reasoning model according to the long-term thinking reinforcement learning training to obtain the target reasoning model.

[0109] In this embodiment, long-thinking reinforcement learning training focuses on how to enable the model to think and reason deeply when facing complex problems. It usually involves using a longer context window to provide more thinking space and background information, thereby helping the model to better understand the problem and make accurate decisions.

[0110] Among them, the target reasoning model is used to output the target answer including the reasoning process according to the input multimodal data.

[0111] Exemplarily, the output of the basic inference model is restricted according to a preset text length limit, and the basic inference model is trained in combination with a reinforcement learning algorithm. During the training process, the correctness of the inference long chain is guided through a result reward mechanism to obtain a target inference model. Among them, the reinforcement learning algorithm includes any one of the group relative policy optimization algorithm, the deep Q-learning network algorithm, and the policy gradient algorithm. The preset text length can be set to 16K or 24K.

[0112] Optionally, long-thinking reinforcement learning training can also combine advanced strategies such as knowledge distillation, the Group Relative Policy Optimization (GRPO) algorithm, and partial rollback techniques, so as to effectively control the computational cost while maintaining high performance. Among them, GRPO is a new type of reinforcement learning algorithm aimed at improving the efficiency and stability of reinforcement learning. Its main feature is to optimize the policy model through intra-group relative rewards, rather than relying on the traditional critic model. This method significantly reduces the memory occupancy and computational cost during the training process while maintaining the stability and efficiency of policy updates.

[0113] In this embodiment, the basic inference model (such as Vision-R1-CI) is optimized through long-thinking reinforcement learning training, thereby significantly improving the inference ability. By directly rewarding the results to handle complex problems, the model can autonomously determine the inference depth, thereby realizing the correctness of the long chain guided by the reward.

[0114] Optionally, when the model parameters are large, the long-thinking training method can also be directly used, and this training method is particularly outstanding in terms of training speed.

[0115] In the above model inference ability training method, a training data set is constructed through the target text inference chain obtained by converting multi-modal data; thus, high-quality target text inference chains can be efficiently and automatically generated, which is convenient for subsequent supervised fine-tuning of the multi-modal large language model and optimization training of the basic inference model, significantly improving the training effect. According to the training data set, the pre-trained multi-modal large language model is supervised and fine-tuned to obtain a basic inference model; thus, the inference ability and generalization ability of the basic inference model can be improved. The basic inference model is optimized according to the long-thinking reinforcement learning training to obtain a target inference model; the target inference model is used to output a target answer including the inference process according to the input multi-modal data. Thus, long text constraints can be directly used for reinforcement learning, greatly improving the training efficiency; and by using long-thinking reinforcement learning training, the model can more easily learn the correct thinking process during training to improve the inference ability of the multi-modal large language model to handle complex visual inference tasks and show the correct thinking process during the inference process.

[0116] In an exemplary embodiment, as Figure 3 shown, the above step S201 may include steps S2011 to S2013.

[0117] Among them:

[0118] Step S2011: Convert the image data in the multimodal data into text description information containing visual details.

[0119] In this embodiment, the multimodal data can be processed by means of a multimodal large language model, and a prompt template is added during the processing, so as to facilitate the multimodal large language model to output more text description information rich in visual details. For example: There is a group of ducks in the original image. The prompt template can be a preset general template: Please provide a detailed description containing all necessary visual details to answer the question. It can also be a user-defined prompt, such as: Please provide a detailed description from details such as the number of ducks and the orientation of the duck heads.

[0120] Exemplarily, obtain multimodal data, where the multimodal data includes: image data and a question associated with the image data; input a target instruction containing the multimodal data into a target multimodal large language model for processing to obtain text description information containing visual details; the text description information containing visual details is used to describe the visual features and / or visual elements captured from the image data.

[0121] Optionally, taking the image data of geometric mathematics as an example, the text description information containing visual details may be: mathematical symbols, chart values, etc.

[0122] Among them, the target instruction is used to indicate splicing the question and the text description information containing visual details according to a preset prompt template; among them, the preset prompt template may be in the following format:

[0123] "Provide the image, question: {question}, provide a detailed description containing all necessary visual details to answer the question".

[0124] Step S2012: Process the text description information containing visual details through a text inference model to obtain a target text inference chain.

[0125] Exemplarily, the text description information containing visual details can be input into the pure text inference model DeepSeek-R1 to trigger it to generate a complex CoT including self-reflection and verification. DeepSeek-R1 will output a multimodal CoT (text inference chain associated with the image) that conforms to the human cognitive style.

[0126] Exemplarily, text description information containing visual details and questions associated with image data are input into a text inference model to output an initial text inference chain; each inference step in the initial text inference chain is evaluated by a first reward model to obtain an inference score corresponding to the initial text inference chain; the initial text inference chain is filtered according to the inference score; and the filtered initial text inference chain is clustered to obtain a target text inference chain.

[0127] Among them, clustering the filtered initial text inference chain to obtain a target text inference chain includes: clustering the filtered initial text inference chain by a preset clustering algorithm to obtain text inference chains of several categories; and randomly extracting a preset number of text inference chains from the text inference chains of several categories as the target text inference chain.

[0128] In this embodiment, the data quality of the initial text inference chain often varies. Therefore, in order to improve the training effect of the subsequent model, it is necessary to filter and screen the initial text inference chain. Among them, filtering can be performed according to hard rules or automatically through a trained model. The screening process needs to consider more about the diversity of data, so it can be screened from the perspective of data categories.

[0129] Exemplarily, the initial text inference chain is filtered according to a preset rule; the preset rule includes: the correctness of the answer and / or the correctness of the logic in the inference process.

[0130] Optionally, the correctness of the answer is relatively easy to understand. Taking a math problem as an example, it will have an accurate standard answer. If the inferred target answer is consistent with the standard answer, the inference result is considered correct; if the inferred target answer is inconsistent with the standard answer, the inference result is considered incorrect.

[0131] Optionally, regarding the correctness of the logic in the inference process, taking a math problem as an example, the logic is whether there is a contradiction in the content before and after. For example: in the first step, it is mentioned that the length of line segment AB is 3, but in the third step later, the length of line segment AB appears as 2.5, which results in a logical contradiction (that is, the descriptions of the same object before and after are inconsistent).

[0132] Exemplarily, each inference step in the initial text inference chain can also be evaluated by a first reward model to obtain an inference score corresponding to the initial text inference chain; then, the initial text inference chain is filtered according to the inference score.

[0133] In this embodiment, a reward model can be used to filter low-quality data and score the thinking process. Optionally, when using a process reward model for filtering, the reasoning process can be divided into multiple steps. For example, it can be divided into multiple segmented steps according to each comma or period, and each step is scored (i.e., the thinking process is scored). Whether each step is correct corresponds to a score. Then, by synthesizing the scores of all steps, a comprehensive score can be obtained (this comprehensive score is related to the final accuracy rate of the model).

[0134] Exemplarily, the filtered text inference chains are clustered to obtain text inference chains of several categories; from the text inference chains of several categories, a preset number of text inference chains are randomly selected as target text inference chains.

[0135] Optionally, the K-means algorithm and the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm can be used for clustering. Among them, the K-means algorithm is an unsupervised learning method, and its main purpose is to divide the samples in the dataset according to their features, so that the data points within the same cluster are highly similar, while the data points between different clusters are less similar; this algorithm continuously optimizes the cluster centers through iteration to achieve the best clustering effect. The DBSCAN algorithm is an algorithm that does not require predefining the number of clusters, can handle noise points, and can discover clusters of any shape.

[0136] Furthermore, clustering the filtered text inference chains can obtain text inference chains of several categories. For example, it can be classified into multiple categories of data such as mathematics, geometry, history, geography, etc. according to the category of the question. Then, a certain number of data are randomly selected from each category of data, so as to ensure data diversity.

[0137] Step S2013, construct a training dataset based on the target text inference chains.

[0138] In this embodiment, after obtaining the filtered and screened target text inference chains, a training dataset is constructed based on these target text inference chains.

[0139] In this embodiment, modal bridging technology is used to convert multimodal data containing images into text description information rich in visual details, and a pure text reasoning model (such as DeepSeek-R1) is used to generate complex CoT data containing self-reflection. Then, through the problems of images and image associations, detailed image description information is generated, so that key visual information (such as mathematical symbols, spatial relationships, etc.) can be completed. Finally, the text description information is input into the text reasoning model, and a high-quality CoT (i.e., the target text reasoning chain) is output. Thereby, the visual information is seamlessly converted into a text reasoning chain, avoiding the reasoning bias caused by modal differences in traditional methods. In addition, this method of constructing a training data set also significantly reduces the cost of manual annotation, and can automatically and efficiently generate high-quality CoT data, thereby improving the accuracy and effect of subsequent model training.

[0140] In an exemplary embodiment, Figure 4 As shown, a second model training method is provided, which is applied to Figure 1 The terminal or server in the example is used to illustrate, including the following steps S401 to S404. Among them:

[0141] Step S401, constructing a training data set according to the target text inference chain obtained by multimodal data conversion.

[0142] Step S402: Perform supervised fine-tuning on the pre-trained multimodal large language model according to the training data set to obtain a basic reasoning model.

[0143] In this embodiment, for the specific implementation process and technical effects of steps S401 to S402, please refer to Figure 2 The descriptions of steps S201 to S202 in the illustrated method embodiment are not repeated here.

[0144] Step S403, establishing a result reward mechanism according to format standardization and answer correctness.

[0145] Among them, format standardization includes: correctness of mathematical symbols and completeness of step numbers; among them, rewards will be given when the format is standard and / or the answer is correct; no rewards will be given when the format is not standard and the answer is incorrect.

[0146] In this embodiment, a result reward mechanism can be established based on the format specification of each step in the reasoning process and the correctness of the final output answer. Optionally, a reward (r=1) is given only when the model output meets both the format specification (such as correct mathematical symbols and complete step numbers) and the answer is correct, otherwise, no reward (r=0) is given.

[0147] Optionally, it is also possible to score separately in terms of format specification and correct answer. Only when both are satisfied can the highest score be obtained. For example, when scoring separately for the two aspects, the correct answer can be given 0.5 points and the step numbers are complete with 0.5 points, and then the final score is calculated.

[0148] Optionally, different weights can also be set according to the degree of importance. For example, the correct answer is 0.7 points and the step numbers are complete with 0.3 points, so that a biased guiding result can be obtained based on the importance of different factors.

[0149] In step S404, the output of the basic inference model is restricted according to a preset text length, and the basic inference model is trained in combination with a reinforcement learning algorithm. During the training process, the correctness of the inference long chain is guided through a result reward mechanism to obtain a target inference model.

[0150] Among them, the reinforcement learning algorithm includes any one of the group relative policy optimization algorithm, the deep Q-learning network algorithm, and the policy gradient algorithm.

[0151] Optionally, reinforcement learning is combined with any one of the group relative policy optimization algorithm (Group Relative Policy Optimization, GRPO), the deep Q-learning network algorithm (Deep Q-Network, DQN), and the policy gradient algorithm (reinforcement++), so that the computational cost can be effectively controlled while maintaining high performance.

[0152] Among them, GRPO is a new type of reinforcement learning algorithm aimed at improving the efficiency and stability of reinforcement learning. Its main feature is to optimize the policy model through intra-group relative rewards, rather than relying on the traditional critic model. This method significantly reduces the memory occupancy and computational cost during the training process, while maintaining the stability and efficiency of policy updates.

[0153] Among them, DQN is a reinforcement learning algorithm that combines deep learning and Q-Learning. Its main purpose is to approximate the action value function through a neural network, so as to solve the limitations of the traditional Q-Learning algorithm when facing a continuous state space.

[0154] Among them, reinforcement++ is a high-performance and scalable MindSpore (AI computing framework) reinforcement learning framework aimed at improving the efficiency and performance of reinforcement learning, and is applicable to various complex decision-making and control tasks.

[0155] It should be understood that this embodiment does not limit the type of reinforcement learning algorithm. The purpose of the above reinforcement learning algorithm is to accelerate and guide the correctness of long chains in the reasoning process.

[0156] In this embodiment, a training data set is constructed based on the target text inference chain obtained by converting multi-modal data; the pre-trained multi-modal large language model is supervised and fine-tuned according to the training data set to obtain a basic inference model; a result reward mechanism is established according to format standardization and answer correctness; the output of the basic inference model is restricted according to a preset text length, and the basic inference model is trained in combination with a reinforcement learning algorithm, and the correctness of the inference long chain is guided through the result reward mechanism during the training process to obtain a target inference model. Thus, long text constraints can be directly used for reinforcement learning, greatly improving the training efficiency; and by using long-thinking reinforcement learning training, the model can relatively easily learn the correct thinking process during training to improve the inference ability of the multi-modal large language model to handle complex visual inference tasks and show the correct thinking process during the inference process.

[0157] In an exemplary embodiment, as Figure 5 shown, a third model training method is provided. Taking the method applied to Figure 1 the terminal or server in

[0158] Step S501, construct a training data set according to the target text inference chain obtained by converting multi-modal data.

[0159] Step S502, according to the training data set, supervise and fine-tune the pre-trained multi-modal large language model to obtain a basic inference model.

[0160] In this embodiment, for the specific implementation process and technical effects of steps S501 to S502, please refer to Figure 2 the relevant descriptions of steps S201 to S202 in the method embodiment shown, which will not be elaborated here.

[0161] Step S503, restrict the output of the basic inference model according to a preset text length, and train the basic inference model in combination with a reinforcement learning algorithm, and evaluate each step of the inference long chain through a second reward model during the training process to obtain an evaluation result.

[0162] In this embodiment, the reinforcement learning algorithm includes any one of the population relative policy optimization algorithm, the deep Q-learning network algorithm, and the policy gradient algorithm. Among them, the difference between step S503 and Figure 4 the method shown is that the evaluation result can be obtained by evaluating each step of the inference long chain through a second reward model.

[0163] Optionally, a Reward Model can be used to score each step of the reasoning process. For example, the reasoning process can be divided into multiple steps (step 1, step 2, step 3,... are set respectively), and then each step is scored through the reward model. For instance, assume that the entire reasoning process consists of three steps in total. Among them, the score of step 1 is 0.5 points, the score of step 2 is 0 points, and the score of step 3 is 1 point; then the comprehensive score of this reasoning process is 1.5 points.

[0164] Optionally, for non - mathematical reasoning processes, it is sometimes difficult to divide specific steps. At this time, it can be divided into multiple segments according to each comma or period, and each segment is scored (i.e., the thinking process is scored). Each segment corresponds to a score, and then the scores of all segments are combined to obtain a comprehensive score, which is the final score of the entire reasoning process.

[0165] Step S504, according to the evaluation results, guide the correctness of the reasoning long chain to obtain the target reasoning model.

[0166] In this embodiment, combined with the foregoing step S503, it can be seen that the second reward model has an evaluation result for each step or each segment in the reasoning process. Therefore, it is possible to realize the supervision of the long chain and guide the model to reason in the correct direction. This guiding method is very intuitive, can accelerate the convergence speed of model training, improve the reasoning ability of the target reasoning model, and enable it to handle complex visual tasks well.

[0167] In this embodiment, through the hard - formatted result reward mechanism, the result reward can be used to guide the correctness of the long chain in the reasoning process, thereby greatly accelerating the training process, preventing the model from obtaining rewards through "guessing answers" or incorrect format paths, and enhancing the rigor of model reasoning. In addition, forcing the model to follow the structured output and adapting to the complex task evaluation criteria can improve the generalization ability of the model.

[0168] In an exemplary embodiment, as Figure 6 shown, a data - processing method is provided. Taking the example that this method is applied to the Figure 1 terminal or server as an example, it includes the following steps S601 to step S602. Among them:

[0169] Step S601, in response to the multi - modal data input on the display interface.

[0170] Among them, the multi - modal data includes: image data, and the question associated with the image data.

[0171] In this embodiment, the more common scenarios include: the user enters the application display interface (hereinafter referred to as the display interface) of the target inference model through a small program or application installed on the terminal. Optionally, at least one dialog box (as shown in Figure 7A ), or multiple display areas with different contents (as shown in Figure 8A ) can be seen on the display interface. The user can input or import (copy and paste) multimodal data under the prompt of the prompt message on the display interface.

[0172] Optionally, in a relatively special scenario, a large image (such as a photo or scanned copy of a test paper) can be imported. At this time, the target model can input multiple questions associated with the image for the image. For example, please answer the question numbered 1. Another example is, what are the answers to the five questions numbered 1 to 5?

[0173] Step S602, input the multimodal data into the target inference model for processing, and output the target answer including the inference process.

[0174] In this embodiment, when the target inference model is loaded in the terminal, the target inference model in the terminal can be used for question inference and answer in an offline or online situation. When the target inference model is loaded in the server, it can support online inference and answer.

[0175] Among them, the target inference model is trained by the model training method of the above Figures 2 - 5 shown method embodiment. Therefore, the specific training process of the target inference model will not be elaborated here.

[0176] In this embodiment, by responding to the multimodal data input on the display interface, the multimodal data includes: an image and questions associated with the image; inputting the multimodal data into the target inference model for processing, and outputting the target answer including the inference process. Thus, in the Q&A for multimodal data, the inference process of the questions associated with the image can be displayed, and the inference process and results similar to human thinking and verification can be intuitively given, with strong interactivity and good user experience.

[0177] Exemplarily, as shown in Figure 7A , a schematic diagram of the display interface of a target inference model in a terminal application is shown Figure 1, such a display interface is user-friendly for mobile terminals such as mobile phones. It uses a dialogue and Q&A form to complete the analysis and reasoning of multi-modal data, and finally obtains a target answer including the reasoning process. Optionally, at least one dialog box is displayed on the display interface, and a prompt "Please upload a picture" is included in the dialog box. When a touch operation is received for this dialog box (such as touching the area where the dialog box is located in the ways of single-click, long-press, double-click, swipe, etc.), a pop-up window or a drop-down window will be used to remind the user to upload a picture (for example, it can be prompted to upload a picture by taking a photo, uploading from the album, etc.). Optionally, at least one prompt will also be displayed below the dialog box, and this prompt is used to assist the user in using the dialogue function. For example: "I can help you search and answer questions. Come and ask me quickly~".

[0178] Optionally, as shown in Figure 7A , at least one input box is also displayed at the bottom of the display interface. This input box is used to receive the text information input by the user. This text information can be a question about the uploaded picture or some simple descriptions of the uploaded picture. Similarly, this input box also supports the upload of content in formats such as pictures and files.

[0179] Exemplarily, as Figure 7B shown, a schematic diagram of the display interface of a target inference model in a terminal application is shown Figure 2 . Suppose the user uploads a picture through the upload control in the input box and correspondingly inputs the text content associated with the picture. Optionally, suppose the picture is a mathematical geometry problem, and its corresponding question is: As shown in the figure, AB = AC, angle A = 36 degrees, the perpendicular bisector of AB intersects AC at D, then what is the degree of angle BDC? Options: A: 72, B: 36, C: 60, D: 82.

[0180] It should be understood that after the picture upload and the input of the associated text content are completed, the target inference model will enter the reasoning process, and this process will take different times due to the difficulty of the specific problem.

[0181] Exemplarily, as Figure 7C shown, a schematic diagram of the display interface of a target inference model in a terminal application is shown Figure 3 . After a certain period of reasoning, the following reasoning process is generated:

[0182] Alright, let's see. The problem shows... First, since AB = AC, the perpendicular bisector of AB. Let's call the midpoint of AB point m, so BM = MA... Wait, maybe I should consider the properties of the perpendicular bisector. But wait, angle DBA is part of angle ABC. But wait, maybe I'm a bit confused now. Let me try a different approach. Let's see. In triangle BDC, but wait, this might not be directly applicable. Wait, maybe we can... This is option A. Wait, but let me double-check. Or, maybe using So, the answer is 72 degrees, which is option a. This seems correct. Final answer: one< / answer>.

[0183] It can be seen from this that the target reasoning model gives a reasoning process similar to the human thinking process. In this reasoning process, the text length of the text reasoning chain can be restricted to avoid the reasoning process from being too long, that is, this reasoning process allows a certain step jump.

[0184] In one embodiment, taking the example of implementing the above data processing method through the interaction between the server 104 and the terminal 102, that is, Figure 1 Based on the application environment shown, the description is as follows. Refer to Figure 7D The timing schematic diagram of the data processing of the target reasoning model in the first embodiment shown includes:

[0185] Step S710D, the terminal displays the application display interface of the target reasoning model.

[0186] Among them, at least one dialog box is displayed on the application display interface.

[0187] Step S720D, the terminal responds to the picture uploaded in the dialog box and the question associated with the picture input.

[0188] Step S730D, the terminal sends the picture and the question associated with the picture to the server.

[0189] Step S740D, the server performs reasoning on the picture and the question associated with the picture through the target reasoning model, and outputs the reasoning process and the target answer.

[0190] Step S750D, the server sends the reasoning process and the target answer to the terminal.

[0191] Step S760D, the terminal displays the reasoning process and the target answer in the dialog box.

[0192] In this embodiment, the terminal displays the application display interface of the target inference model. After receiving the uploaded picture in the dialog box and the input question associated with the picture, the picture and the question associated with the picture are sent to the server. The server uses the target inference model to perform inference on the picture and the question associated with the picture, and after outputting the inference process and the target answer, the inference process and the target answer are sent to the terminal, and finally the inference process and the target answer are displayed in the dialog box of the terminal. Thus, the input multi-modal data (picture and question) can be accurately analyzed and inferred in the form of chat Q&A, and the completed inference process and target answer are fed back to the user terminal, making the entire analysis process more in line with human thinking habits and helping people solve more complex problems.

[0193] Exemplarily, as Figure 8A shown, a schematic diagram of the display interface of a target inference model applied in a terminal is shown Figure 4 , such a display interface is friendly to terminals such as tablets and computers, and it includes multiple content input boxes. Optionally, on the display interface of the terminal application, there are: a picture upload display area, a text input display area, a thinking process display area, and an answer result display area, and each area is used to display different contents. Exemplarily, there is a prompt "Please upload a picture" in the picture upload display area. When a touch operation on this picture upload display area is received (for example, touching the area where the dialog box is located in ways such as clicking, long pressing, double clicking, sliding, etc.), a pop-up window or a drop-down window will be used to remind of uploading a picture (for example, it can be prompted to upload a picture by taking a photo, uploading from the album, etc.). Optionally, it also supports uploading pictures by dragging. For example, if the picture to be uploaded is dragged to the range where the picture upload display area is located, the upload of the dragged picture can be completed.

[0194] Exemplarily, in the text input display area, the input method can be directly called to input the content associated with the uploaded picture. For example, if a photo including multiple ducklings is uploaded, the questions associated with this picture can be: Question 1: How many ducks are there in the picture? Question 2: How many ducks face north and how many face south in the picture? Question 3: What are the numbers of adult ducks and ducklings in the picture respectively?

[0195] Optionally, referring to Figure 8A , the thinking process display area is displayed at the lower position of the picture upload display area and the text input display area, and the answer result display area is displayed below the thinking process display area.

[0196] It should be understood that the embodiments of the present application do not limit the distribution relationship of the picture upload display area, the text input display area, the thinking process display area, and the answer result display area in terms of spatial position, nor the display size of each area.

[0197] Optionally, when the input text content is long, a scroll bar can be formed in the text input display area to display the input text content in a scrolling form. Similarly, in the thinking process display area, when the text reasoning chain is long, a scroll bar can also be formed in the thinking process display area to display the reasoning process in a scrolling form. Optionally, the display ranges of each display area can also be adjusted dynamically. For example, when the content of the input text is less and the content of the thinking process is more, the display size of the thinking process display area can be appropriately increased, and the display size of the text input display area can be appropriately reduced.

[0198] Exemplarily, as Figure 8B shown, a schematic diagram of the display interface of a target reasoning model in a terminal application is shown Figure 5 . Assume that the user uploads a picture in the picture upload display area by dragging, and correspondingly enters the text content associated with the picture in the text input display area. Optionally, assume that the picture is a mathematical geometry problem, and its corresponding question is: As shown in the figure, AB is parallel to CD, EG bisects angle BEF, then the degree of angle 2 is (). Options: A: 50 B: 60 C: 65 D: 90.

[0199] It should be understood that after the picture upload and the input of the associated text content are completed, the target reasoning model will enter the reasoning process, and this process will take different amounts of time due to the difficulty of the specific problem.

[0200] Exemplarily, as Figure 8C shown, a schematic diagram of the display interface of a target reasoning model in a terminal application is shown Figure 6 . After a certain period of reasoning, the following reasoning process is generated:

[0201] <Thought>, Well, let's see. The question shows that AB is parallel to CD and EG bisects angle BEF. We need to find the measure of angle 2. First, the image shows waiting, angle 1 is at point F, so it is the supplement of angle 1 because they are same-side interior angles. So angle BEF should be 180 - 50 = 130 degrees. Then, EG bisects angle BEF, thus dividing angle BEF into two equal parts. Therefore, each angle formed by the bisector EG will be 130 / 2 = 65 degrees. Since angle 2 is one of these bisected angles, angle 2 should be 65 degrees. Looking at the options, option C is 65. So the answer should be C. That is< / Thought><Answer>Final answer: C< / Answer>.

[0202] From this, it can be seen that the target reasoning model gives a reasoning process similar to the human thinking process. In this reasoning process, the text length of the text reasoning chain can be restricted to avoid the reasoning process from being too long, that is, this reasoning process allows a certain step jump.

[0203] Optionally, in combination with Figure 8C As shown, after all the reasoning processes are completed in the thinking process display area, the final answer: C is displayed in the answer result display area at the bottom.

[0204] In one embodiment, taking the interaction between the server 104 and the terminal 102 to implement the above data processing method as an example, that is, on the basis of the application environment shown in Figure 1 it is described, see Figure 8D the timing diagram of the target inference model data processing method in the second embodiment shown in, including:

[0205] Step S810D, the terminal displays the application display interface of the target inference model.

[0206] Among them, a picture upload display area, a text input display area, a thinking process display area, and an answer result display area are displayed on the application display interface.

[0207] Step S820D, the terminal responds to the picture uploaded in the picture upload display area and the question associated with the picture input in the text input display area.

[0208] Step S830D, the terminal sends the picture and the question associated with the picture to the server.

[0209] Step S840D, the server performs reasoning on the picture and the question associated with the picture through the target inference model, and outputs the reasoning process and the target answer.

[0210] Step S850D, the server sends the reasoning process and the target answer to the terminal.

[0211] Step S860D, the terminal displays the reasoning process in the thinking process display area and the target answer in the answer result display area.

[0212] In this embodiment, the application display interface of the target inference model is displayed through the terminal. Among them, a picture upload display area, a text input display area, a thinking process display area, and an answer result display area are shown on the application display interface. After the terminal responds to the picture uploaded in the picture upload display area and the question associated with the picture input in the text input display area, it sends the picture and the question associated with the picture to the server. The server uses the target inference model to reason about the picture and the question associated with the picture, and outputs the reasoning process and the target answer. The server sends the reasoning process and the target answer to the terminal, and finally the reasoning process is displayed in the thinking process display area of the terminal application display interface, and the target answer is displayed in the answer result display area. Therefore, accurate analysis and reasoning can be performed on the input multimodal data (picture and question), and the complete reasoning process and the target answer are fed back to the user terminal, making the entire analysis process more in line with human thinking habits and helping people solve more complex problems. In addition, the overall display interface is very friendly, facilitating user operation and having strong interactivity.

[0213] It should be understood that although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0214] Based on the same inventive concept, an embodiment of the present application further provides a model training device for implementing the above-mentioned model training method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following model training device embodiments can refer to the limitations on the model inference ability training method in the above text, and will not be repeated here.

[0215] In an exemplary embodiment, as Figure 9 shown, a model training device is provided, including: a dataset construction module 901, a fine-tuning module 902, and an optimization training module 903, where:

[0216] The dataset construction module 901 is used to construct a training dataset based on the target text inference chain obtained by converting multimodal data. The target text inference chain includes: text description information, the input question, the inference process, and the target answer; the text description information includes: the text information obtained when describing multimodal data in text form; the inference process includes: each step of inferring the target answer based on the text description information and the input question.

[0217] The fine-tuning module 902 is used to perform supervised fine-tuning on the pre-trained multimodal large language model according to the training dataset to obtain a basic inference model.

[0218] The optimization training module 903 is used to optimize the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model; the target inference model is used to output a target answer including the inference process according to the input multimodal data.

[0219] Exemplarily, the dataset construction module 901 is specifically used to: convert the image data in the multimodal data into text description information containing visual details; process the text description information containing visual details through a text inference model to obtain a target text inference chain; construct a training dataset based on the target text inference chain.

[0220] Exemplarily, converting the image data in the multimodal data into text description information containing visual details includes: obtaining multimodal data, where the multimodal data includes: image data and a question associated with the image data; inputting the target instruction containing the multimodal data into a target multimodal large language model for processing to obtain text description information containing visual details; the text description information containing visual details is used to describe the visual features and / or visual elements captured from the image data.

[0221] Exemplarily, processing the text description information containing visual details through a text inference model to obtain a target text inference chain includes: inputting the text description information containing visual details and the question associated with the image data into the text inference model to output an initial text inference chain; evaluating each inference step in the initial text inference chain through a first reward model to obtain an inference score corresponding to the initial text inference chain; filtering the initial text inference chain according to the inference score; performing clustering processing on the filtered initial text inference chain to obtain a target text inference chain.

[0222] Exemplarily, performing clustering processing on the filtered initial text inference chain to obtain a target text inference chain includes: performing clustering processing on the filtered initial text inference chain through a preset clustering algorithm to obtain text inference chains of several categories; randomly extracting a preset number of text inference chains from the text inference chains of several categories as the target text inference chain.

[0223] Exemplarily, the fine-tuning module 902 is specifically configured to: train the pre-trained multi-modal large language model based on the training data set, and during the training process, update at least some of the parameters in the pre-trained multi-modal large language model or update the parameters of the plug-in network corresponding to the pre-trained multi-modal large language model to obtain a basic inference model.

[0224] Exemplarily, training the pre-trained multi-modal large language model based on the training data set includes:

[0225] Selecting general data from the training data set to optimize and train the pre-trained multi-modal large language model; and / or screening out target domain data from the training data set, and optimizing and training the pre-trained multi-modal large language model through an optimization data set composed of the target domain data and at least some general data; the proportion of the target domain data in the optimization data set is greater than the proportion of the target domain data in the training data set;

[0226] Among them, the general data refers to sample data containing multiple domains randomly extracted from the training data set.

[0227] Exemplarily, the optimization training module 903 is specifically configured to: limit the output of the basic inference model according to a preset text length, train the basic inference model in combination with a reinforcement learning algorithm, and guide the correctness of the inference long chain through a result reward mechanism during the training process to obtain a target inference model.

[0228] Exemplarily, the optimization training module 903 is specifically configured to: limit the output of the basic inference model according to a preset text length, train the basic inference model in combination with a reinforcement learning algorithm, and evaluate each step of the inference long chain through a second reward model during the training process to obtain an evaluation result; guide the correctness of the inference long chain according to the evaluation result to obtain a target inference model.

[0229] In this embodiment, a training data set is constructed based on the target text inference chain obtained by converting multi-modal data. Thus, a high-quality target text inference chain can be efficiently and automatically generated, which is convenient for subsequent supervised fine-tuning of the multi-modal large language model and optimization training of the basic inference model, significantly improving the training effect. According to the training data set, the pre-trained multi-modal large language model is supervised and fine-tuned to obtain a basic inference model, thereby improving the inference ability and generalization ability of the basic inference model. The basic inference model is optimized according to the reinforcement learning training of long thinking to obtain a target inference model, which is used to output a target answer containing the inference process according to the input multi-modal data. Thus, long text constraints can be directly used for reinforcement learning, greatly improving the training efficiency. By using the reinforcement learning training of long thinking, the model can easily learn the correct thinking process during training, improve the inference ability of the multi-modal large language model to handle complex visual reasoning tasks, and show the correct thinking process during the inference process.

[0230] In an exemplary embodiment, as Figure 10 shown, another model training device is provided. Based on the device Figure 9 shown, it may further include: a result reward mechanism establishment module 904, configured to establish a result reward mechanism according to format standardization and answer correctness. The format standardization includes the correctness of mathematical symbols and the integrity of step numbers. Among them, rewards are given when the format is standard and / or the answer is correct, and no rewards are given when the format is not standard and the answer is incorrect.

[0231] In this embodiment, through the hard-formatted result reward mechanism, the result reward is used to guide the correctness of the long chain in the inference process, which can greatly accelerate the training process, prevent the model from obtaining rewards through "guessing answers" or incorrect format paths, and improve the rigor of model inference. In addition, forcing the model to follow the structured output and adapting to the complex task evaluation standard can improve the generalization ability of the model.

[0232] In an exemplary embodiment, as Figure 11 shown, a data processing device is provided. The device includes: a receiving module 1101 and a processing module 1102. Among them:

[0233] The receiving module 1101 is configured to respond to the multi-modal data input on the display interface. The multi-modal data includes image data and a question associated with the image data.

[0234] The processing module 1102 is configured to input the multi-modal data into the target inference model for processing and output a target answer containing the inference process. The target inference model is based on Figures 2 - 5It is trained by the model training method provided in the method embodiment shown above.

[0235] In this embodiment, in response to multi-modal data input on a display interface, the multi-modal data includes: image data and a question associated with the image data; the multi-modal data is input into a target inference model for processing, and a target answer including an inference process is output. Thus, in the question-and-answer of multi-modal data, the inference process of the question associated with the image can be displayed, and the inference process and result similar to human thinking and verification can be intuitively given, with strong interactivity and good user experience.

[0236] Each module in the above model inference ability training device and data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0237] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 12 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a model inference ability training method or a data processing method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0238] Those skilled in the art can understand that Figure 12The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0239] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0240] Construct a training data set according to the target text inference chain obtained by converting multi-modal data. The target text inference chain includes: text description information, input questions, inference processes, and target answers; the text description information includes: text information obtained when describing multi-modal data in text form; the inference process includes: each step of inferring the target answer according to the text description information and the input questions; according to the training data set, perform supervised fine-tuning on the pre-trained multi-modal large language model to obtain a basic inference model; optimize the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model; the target inference model is used to output a target answer including the inference process according to the input multi-modal data.

[0241] Exemplarily, constructing a training data set according to the target text inference chain obtained by converting multi-modal data includes: converting the image data in the multi-modal data into text description information including visual details; processing the text description information including visual details through a text inference model to obtain a target text inference chain; constructing a training data set based on the target text inference chain.

[0242] Exemplarily, converting the image data in the multi-modal data into text description information including visual details includes: obtaining multi-modal data, where the multi-modal data includes: image data and questions associated with the image data; inputting the target instruction including the multi-modal data into the target multi-modal large language model for processing to obtain text description information including visual details; the text description information including visual details is used to describe the visual features and / or visual elements captured from the image data.

[0243] Exemplarily, processing the text description information including visual details through a text inference model to obtain a target text inference chain includes: inputting the text description information including visual details and the questions associated with the image data into the text inference model to output an initial text inference chain; evaluating each inference step in the initial text inference chain through a first reward model to obtain an inference score corresponding to the initial text inference chain; filtering the initial text inference chain according to the inference score; clustering the filtered initial text inference chain to obtain a target text inference chain.

[0244] Exemplarily, clustering processing is performed on the filtered initial text inference chain to obtain a target text inference chain, including: performing clustering processing on the filtered initial text inference chain through a preset clustering algorithm to obtain text inference chains of several categories; randomly extracting a preset number of text inference chains from the text inference chains of several categories as the target text inference chain.

[0245] Exemplarily, according to the training data set, supervised fine-tuning is performed on the pre-trained multi-modal large language model to obtain a basic inference model, including: training the pre-trained multi-modal large language model based on the training data set, and during the training process, updating at least part of the parameters in the pre-trained multi-modal large language model or updating the parameters of the plug-in network corresponding to the pre-trained multi-modal large language model to obtain the basic inference model.

[0246] Exemplarily, training the pre-trained multi-modal large language model based on the training data set includes: selecting general data from the training data set to perform optimization training on the pre-trained multi-modal large language model; and / or

[0247] screening out target domain data from the training data set, and performing optimization training on the pre-trained multi-modal large language model through an optimization data set composed of the target domain data and at least part of the general data; the proportion of the target domain data in the optimization data set is greater than the proportion of the target domain data in the training data set;

[0248] wherein, the general data refers to sample data containing multiple domains randomly extracted from the training data set.

[0249] Exemplarily, optimizing the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model, including: restricting the output of the basic inference model according to a preset text length, training the basic inference model in combination with the reinforcement learning algorithm, and guiding the correctness of the inference long chain through a result reward mechanism during the training process to obtain the target inference model.

[0250] Exemplarily, it further includes: establishing a result reward mechanism according to format standardization and answer correctness; the format standardization includes: the correctness of mathematical symbols and the integrity of step numbers; wherein, rewards are given in the case of format standardization and / or correct answers; no rewards are given in the case of non-standard format and incorrect answers.

[0251] Exemplarily, the basic inference model is optimized according to the reinforcement learning training with long thinking to obtain the target inference model, including: restricting the output of the basic inference model according to a preset text length, training the basic inference model in combination with the reinforcement learning algorithm, and evaluating each step of the inference long chain through a second reward model during the training process to obtain an evaluation result; guiding the correctness of the inference long chain according to the evaluation result to obtain the target inference model.

[0252] In this embodiment, a training data set is constructed through the target text inference chain obtained by converting multi-modal data; thus, high-quality target text inference chains can be efficiently and automatically generated, facilitating subsequent supervised fine-tuning of the multi-modal large language model and optimizing the training process of the basic inference model, significantly improving the training effect. According to the training data set, the pre-trained multi-modal large language model is supervised and fine-tuned to obtain the basic inference model; thus, the inference ability and generalization ability of the basic inference model can be improved. The basic inference model is optimized according to the reinforcement learning training with long thinking to obtain the target inference model; the target inference model is used to output the target answer including the inference process according to the input multi-modal data. Thus, long text constraints can be directly used for reinforcement learning, greatly improving the training efficiency; and the reinforcement learning training with long thinking enables the model to more easily learn the correct thinking process during training, so as to improve the inference ability of the multi-modal large language model to handle complex visual inference tasks and show the correct thinking process during the inference process.

[0253] In another exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0254] In response to multi-modal data input on the display interface, the multi-modal data includes: image data and a question associated with the image data; inputting the multi-modal data into the target inference model for processing, and outputting the target answer including the inference process; the target inference model is trained according to the Figures 2 - 5 model training method provided in the method embodiment shown.

[0255] In this embodiment, in response to multi-modal data input on the display interface, the multi-modal data includes: an image and a question associated with the image; inputting the multi-modal data into the target inference model for processing, and outputting the target answer including the inference process. Thus, in the question and answer for multi-modal data, the inference process of the question associated with the image can be shown, intuitively giving the inference process and result similar to human thinking and verification, with strong interactivity and good user experience.

[0256] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0257] Construct a training data set according to the target text inference chain obtained by converting multimodal data. The target text inference chain includes: text description information, the input question, the inference process, and the target answer; the text description information includes: the text information obtained when describing the multimodal data in text form; the inference process includes: each step of inferring the target answer according to the text description information and the input question; according to the training data set, perform supervised fine-tuning on the pre-trained multimodal large language model to obtain a basic inference model; perform optimization processing on the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model; the target inference model is used to output a target answer including the inference process according to the input multimodal data.

[0258] Exemplarily, constructing a training data set according to the target text inference chain obtained by converting multimodal data includes: converting the image data in the multimodal data into text description information including visual details; processing the text description information including visual details through a text inference model to obtain a target text inference chain; constructing a training data set based on the target text inference chain.

[0259] Exemplarily, converting the image data in the multimodal data into text description information including visual details includes: obtaining multimodal data, where the multimodal data includes: image data and a question associated with the image data; inputting a target instruction including the multimodal data into a target multimodal large language model for processing to obtain text description information including visual details; the text description information including visual details is used to describe the visual features and / or visual elements captured from the image data.

[0260] Exemplarily, processing the text description information including visual details through a text inference model to obtain a target text inference chain includes: inputting the text description information including visual details and the question associated with the image data into the text inference model to output an initial text inference chain; evaluating each inference step in the initial text inference chain through a first reward model to obtain an inference score corresponding to the initial text inference chain; filtering the initial text inference chain according to the inference score; performing clustering processing on the filtered initial text inference chain to obtain a target text inference chain.

[0261] Exemplarily, clustering processing is performed on the filtered initial text inference chain to obtain a target text inference chain, including: performing clustering processing on the filtered initial text inference chain through a preset clustering algorithm to obtain text inference chains of several categories; randomly extracting a preset number of text inference chains from the text inference chains of several categories as the target text inference chain.

[0262] Exemplarily, according to the training data set, supervised fine-tuning is performed on the pre-trained multi-modal large language model to obtain a basic inference model, including: training the pre-trained multi-modal large language model based on the training data set, and during the training process, updating at least part of the parameters in the pre-trained multi-modal large language model or updating the parameters of the plug-in network corresponding to the pre-trained multi-modal large language model to obtain the basic inference model.

[0263] Exemplarily, training the pre-trained multi-modal large language model based on the training data set includes: selecting general data from the training data set to perform optimization training on the pre-trained multi-modal large language model; and / or

[0264] screening out target domain data from the training data set, and performing optimization training on the pre-trained multi-modal large language model through an optimization data set composed of the target domain data and at least part of the general data; the proportion of the target domain data in the optimization data set is greater than the proportion of the target domain data in the training data set;

[0265] wherein, the general data refers to sample data containing multiple domains randomly extracted from the training data set.

[0266] Exemplarily, optimizing the basic inference model according to the reinforcement learning training of long thinking to obtain a target inference model, including: restricting the output of the basic inference model according to a preset text length, and training the basic inference model in combination with the reinforcement learning algorithm, and guiding the correctness of the inference long chain through a result reward mechanism during the training process to obtain the target inference model.

[0267] Exemplarily, it further includes: establishing a result reward mechanism according to format normativity and answer correctness; the format normativity includes: the correctness of mathematical symbols and the integrity of step numbers; wherein, rewards are given when the format is normative and / or the answer is correct; no rewards are given when the format is non-normative and the answer is incorrect.

[0268] Exemplarily, the basic inference model is optimized according to the reinforcement learning training with long thinking to obtain the target inference model, including: restricting the output of the basic inference model according to a preset text length, training the basic inference model in combination with the reinforcement learning algorithm, and evaluating each step of the inference long chain through a second reward model during the training process to obtain an evaluation result; guiding the correctness of the inference long chain according to the evaluation result to obtain the target inference model.

[0269] In this embodiment, a training data set is constructed through the target text inference chain obtained by converting multi-modal data; thus, high-quality target text inference chains can be efficiently and automatically generated, which is convenient for subsequent supervised fine-tuning of the multi-modal large language model and optimization training of the basic inference model, significantly improving the training effect. According to the training data set, the pre-trained multi-modal large language model is supervised and fine-tuned to obtain the basic inference model; thus, the inference ability and generalization ability of the basic inference model can be improved. The basic inference model is optimized according to the reinforcement learning training with long thinking to obtain the target inference model; the target inference model is used to output the target answer including the inference process according to the input multi-modal data. Thus, long text constraints can be directly used for reinforcement learning, greatly improving the training efficiency; and the reinforcement learning training with long thinking enables the model to easily learn the correct thinking process during training, so as to improve the inference ability of the multi-modal large language model to handle complex visual inference tasks and show the correct thinking process during the inference process.

[0270] In another embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0271] In response to multi-modal data input on the display interface, the multi-modal data includes: image data and a question associated with the image data; the multi-modal data is input into the target inference model for processing, and a target answer including the inference process is output; the target inference model is trained according to Figures 2 - 5 the model training method provided in the method embodiment shown.

[0272] In this embodiment, in response to multi-modal data input on the display interface, the multi-modal data includes: image data and a question associated with the image data; the multi-modal data is input into the target inference model for processing, and a target answer including the inference process is output. Thus, in the question and answer for multi-modal data, the inference process of the question associated with the image can be shown, and the inference process and result similar to human thinking and verification can be intuitively given, with strong interactivity and good user experience.

[0273] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.

[0274] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0275] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0276] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0277] The above embodiments only express several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A model training method, characterized in that: The method comprises: A training data set is constructed according to a target text reasoning chain obtained by converting the multimodal data, wherein the target text reasoning chain includes: text description information, an input question, a reasoning process, and a target answer; the text description information includes: text information obtained when describing the multimodal data in text form; the reasoning process includes: various steps of reasoning to obtain the target answer according to the text description information and the input question; According to the training data set, supervised fine-tuning is performed on the pre-trained multimodal large language model to obtain a basic reasoning model; The basic reasoning model is optimized according to long-term thinking reinforcement learning training to obtain a target reasoning model; the target reasoning model is used to output a target answer including a reasoning process according to the input multimodal data.

2. The method according to claim 1, characterized in that The target text inference chain obtained by converting the multimodal data to construct a training data set includes: Converting image data in the multimodal data into text description information containing visual details; Processing the text description information containing the visual details through a text reasoning model to obtain the target text reasoning chain; Based on the target text reasoning chain, the training data set is constructed.

3. The method according to claim 2, characterized in that The converting the image data in the multimodal data into text description information containing visual details comprises: Acquiring multimodal data, the multimodal data comprising: image data, and questions associated with the image data; The target instruction containing the multimodal data is input into the target multimodal large language model for processing to obtain text description information containing visual details; the text description information containing visual details is used to describe the visual features and / or visual elements captured from the image data.

4. The method according to claim 2, characterized in that: The text description information containing visual details is processed by a text reasoning model to obtain the target text reasoning chain, including: Inputting the text description information containing visual details and the questions associated with the image data into the text reasoning model, and outputting an initial text reasoning chain; evaluating each reasoning step in the initial text reasoning chain through the first reward model to obtain a reasoning score corresponding to the initial text reasoning chain; According to the reasoning score, filtering the initial text reasoning chain; The filtered initial text reasoning chain is clustered to obtain the target text reasoning chain.

5. The method according to claim 4, characterized in that The clustering process is performed on the filtered initial text reasoning chain to obtain the target text reasoning chain, including: Clustering the filtered initial text reasoning chains using a preset clustering algorithm to obtain text reasoning chains of several categories; A preset number of text inference chains are randomly extracted from the text inference chains of the plurality of categories as the target text inference chains.

6. The method according to any one of claims 1 to 5, characterized in that: According to the training data set, the pre-trained multimodal large language model is fine-tuned to obtain a basic reasoning model, including: The pre-trained multimodal large language model is trained based on the training data set, and during the training process, at least part of the parameters in the pre-trained multimodal large language model is updated or the parameters of the plug-in network corresponding to the pre-trained multimodal large language model are updated to obtain a basic reasoning model.

7. The method according to claim 6, characterized in that The training of the pre-trained multimodal large language model based on the training data set includes: Selecting common data from the training data set to perform optimization training on the pre-trained multimodal large language model; and / or Filtering target domain data from the training data set, and optimizing the pre-trained multimodal large language model through an optimized data set consisting of the target domain data and at least part of the general data; the proportion of the target domain data in the optimized data set is greater than the proportion of the target domain data in the training data set; The general data refers to sample data containing multiple fields randomly extracted from the training data set.

8. The method according to any one of claims 1 to 5, characterized in that: The basic reasoning model is optimized according to the long-term thinking reinforcement learning training to obtain the target reasoning model, including: The output of the basic reasoning model is limited according to a preset text length, and the basic reasoning model is trained in combination with a reinforcement learning algorithm. During the training process, the correctness of the long chain of reasoning is guided by a result reward mechanism to obtain a target reasoning model.

9. The method according to claim 8, characterized in that Also includes: A result reward mechanism is established according to the format standardization and the correctness of the answer; the format standardization includes: the correctness of mathematical symbols and the completeness of the step numbering; wherein, rewards are given when the format is standard and / or the answer is correct; and no rewards are given when the format is not standard and the answer is incorrect.

10. The method according to any one of claims 1 to 5, characterized in that: The basic reasoning model is optimized according to the long-term thinking reinforcement learning training to obtain the target reasoning model, including: The output of the basic reasoning model is limited according to a preset text length, and the basic reasoning model is trained in combination with a reinforcement learning algorithm, and each step of the long reasoning chain is evaluated by a second reward model during the training process to obtain an evaluation result; The correctness of the long chain of reasoning is guided according to the evaluation results to obtain a target reasoning model.

11. A data processing method, characterized in that: The method comprises: In response to multimodal data input on the display interface, the multimodal data includes: image data, and questions associated with the image data; The multimodal data is input into a target reasoning model for processing, and a target answer including a reasoning process is output; the target reasoning model is trained according to the model training method according to any one of claims 1 to 10.

12. A model training device, characterized in that: The device comprises: A data set construction module is used to construct a training data set according to a target text reasoning chain obtained by converting multimodal data, wherein the target text reasoning chain includes: text description information, input questions, reasoning process, and target answer; the text description information includes: text information obtained when describing the multimodal data in text form; the reasoning process includes: various steps of reasoning to obtain the target answer according to the text description information and the input questions; A fine-tuning module, used to perform supervised fine-tuning on the pre-trained multimodal large language model according to the training data set to obtain a basic reasoning model; The optimization training module is used to optimize the basic reasoning model according to the long-thinking reinforcement learning training to obtain a target reasoning model; the target reasoning model is used to output a target answer including a reasoning process according to the input multimodal data.

13. A data processing device, characterized in that: The device comprises: A receiving module, configured to respond to multimodal data input on a display interface, wherein the multimodal data includes: image data, and questions associated with the image data; A processing module is used to input the multimodal data into a target reasoning model for processing and output a target answer including a reasoning process; the target reasoning model is trained according to the model reasoning capability training method according to any one of claims 1 to 10.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Cited By

  • Inference model training method and device, server, storage medium and product

    CN120525015A

  • Inference model training method, device, server, storage medium and product

    CN120525015B

  • Marketing decision-making method and device, equipment and medium

    CN120707176A

  • Product recommendation model training method, product recommendation method, equipment and readable storage medium

    CN120707255A

  • Method and device for training reasoning model and method and device for processing reasoning problem

    CN120725149A