Post-training method and device for multi-modal model
By freezing and activating the parameters of the encoding and language models in the multimodal model, and updating the parameters based on the prediction results and feedback values, the problem of instability in the multimodal model training process is solved, thereby improving training efficiency and reducing costs.
Patent Information
- Application Number
- CN202511267433.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-16
AI Technical Summary
Multimodal models, especially visual-language models, are prone to instability or even training failure during post-training, leading to low training efficiency and increased costs.
By freezing and activating the parameters of at least one encoding model and language model in the multimodal model, and updating the parameters based on the prediction results and feedback values of the training data, the model can be stably trained until the preset training objective is achieved.
It improves the training efficiency of multimodal models, reduces training costs, and ensures the stability and effectiveness of model training.
Smart Images

Figure CN121145976A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model training technology, and in particular to a method and apparatus for post-training a multimodal model. Background Technology
[0002] With the development of artificial intelligence technology, multimodal models can simultaneously understand, process, or generate two or more different types of information, breaking down the information barriers of single-modal models and becoming closer to the human cognitive way of perceiving the world through multiple channels such as vision, language, and hearing.
[0003] Visual language models are a typical type of multimodal model, combining visual encoding models and language models to simultaneously understand and process multimodal data such as text, images, or videos, and perform complex reasoning and generation tasks. However, in the post-training process of visual language models, mixing different training tasks for joint training can lead to instability in the training process. This can manifest as decreased performance on online test data, a sharp increase in the gradient norm (usually a slow increase), large fluctuations in entropy, or a sudden surge in response length, and may even cause training failure. Summary of the Invention
[0004] This invention provides a method and apparatus for post-training a multimodal model, in order to improve the training efficiency and reduce the cost of post-training.
[0005] In a first aspect, embodiments of the present invention provide a method for post-training a multimodal model, the method comprising:
[0006] The training steps and training task types based on the multimodal model post-training include activating the model parameters of the first training model and freezing the model parameters of at least one second training model in the multimodal model; wherein the multimodal model includes at least two encoding models and at least one language model, and the first training model and the second training model are respectively one of the encoding model and the language model;
[0007] Based on the first training model and the second training model, the training data is predicted and output to obtain the prediction result. Based on the prediction result and the correct result in the training data, at least one feedback value is determined, and the activation model parameters in the first training model are updated based on the feedback value; wherein, the training data includes image or video data.
[0008] Determine whether the preset training objective has been achieved. If so, obtain the target multimodal model. If not, determine whether the training model switching condition is met. If not, repeat the prediction output based on the first training model and the second training model to obtain the prediction result. If so, exchange the first training model and the second training model, and repeat the prediction output based on the first training model and the second training model to obtain the prediction result.
[0009] Secondly, embodiments of the present invention also provide a multimodal model post-training device, the device comprising:
[0010] The parameter freezing module is used to activate the model parameters of the first training model and freeze the model parameters of at least one second training model in the multimodal model, based on the training steps and training task types of the multimodal model post-training. The multimodal model includes at least two encoding models and at least one language model, and the first training model and the second training model are one of the encoding model and the language model, respectively.
[0011] The parameter update module is used to predict the training data based on the first training model and the second training model to obtain the prediction result, determine at least one feedback value based on the prediction result and the correct result in the training data, and update the activation model parameters in the first training model based on the feedback value; wherein, the training data includes image or video data;
[0012] The condition judgment module is used to determine whether the preset training objective has been achieved. If so, the target multimodal model is obtained; if not, it determines whether the training model switching condition is met. If not, it repeats the prediction output based on the first training model and the second training model to obtain the prediction result; if so, it swaps the first training model and the second training model and repeats the prediction output based on the first training model and the second training model to obtain the prediction result.
[0013] Thirdly, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal model post-training method as described in any of the embodiments of the present invention.
[0014] Fourthly, embodiments of the present invention also provide a storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to execute the multimodal model post-training method as described in any of the embodiments of the present invention.
[0015] The technical solution of this invention, based on the training steps and training task types of post-training of multimodal models, involves determining the activation parameters of a first training model and freezing the parameters of a second training model among at least two encoding models and at least one language model in the multimodal model. Based on the first and second training models, prediction outputs are made on the training data to obtain prediction results. Based on the prediction results and correct results in the training data, a feedback value is determined, and the activated model parameters in the first training model are updated based on the feedback value. This updating of the first training model's parameters is repeated until the training model switching conditions are met. The first and second training models are then swapped, and model training continues until a preset training objective is achieved, resulting in the target multimodal model. This solves the problem of instability or even training failure during post-training of multimodal models, especially visual-language models, in the prior art, improving the training efficiency and reducing the cost of post-training of multimodal models.
[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a multimodal model post-training method provided in Embodiment 1 of the present invention;
[0019] Figure 2 This is a flowchart of a multimodal model post-training method provided in Embodiment 2 of the present invention;
[0020] Figure 3 This is a schematic diagram of the structure of a multimodal model post-training device provided in Embodiment 3 of the present invention;
[0021] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. In the embodiments of this application, certain software, components, models, and other existing industry solutions may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0024] The acquisition, transmission, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.
[0025] Example 1
[0026] Figure 1 The flowchart of a multimodal model post-training method is provided in Embodiment 1 of the present invention. This embodiment is applicable to the post-training of multimodal models, especially visual language models. The method can be executed by a multimodal model post-training device, which can be implemented in hardware and / or software and can be configured in a server.
[0027] like Figure 1 As shown, the method includes:
[0028] S110. Training steps and training task types based on multimodal model post-training: activating the model parameters of the first training model and freezing the model parameters of at least one second training model in the multimodal model.
[0029] The multimodal model includes at least two encoding models and at least one language model. This embodiment is applicable to scenarios involving post-training of multimodal models, especially visual language models. Post-training is the process of further optimizing and adjusting the model based on a pre-trained model, tailored to specific tasks, domains, or application requirements. This can be supervised fine-tuning training or reinforcement learning training, etc. In this embodiment, the multimodal models, especially the encoding and language models in the visual language model, are both pre-trained models.
[0030] Visual language models (VLMs) are a typical type of multimodal model capable of simultaneously understanding and processing multimodal data such as text, images, or videos, and performing complex reasoning and generation tasks. A VLM typically includes at least two encoding models and at least one language model. The encoding models can include text encoding models and visual encoding models, etc. Text encoding models capture semantic and contextual relationships between words and phrases, while visual encoding models extract visual features from images or videos. The language model, based on the output of the encoding models, performs reasoning and / or perceptual processing to obtain the model's prediction results.
[0031] In a specific example, a visual language model (VLM) can include a visual encoding model (ViT) and a large language model (LLM). The visual encoding model (ViT) is used to classify and encode each frame of an image or video, and then input it into the large language model (LLM) for inference and / or perceptual processing, and output the prediction result.
[0032] A training step, also known as a step, refers to the entire process of inputting training data, predicting, providing feedback, and adjusting parameters in a single batch. Post-training typically includes dozens or even hundreds of training steps. For visual language models, training task types can include visual perception tasks (such as object detection, localization, OCR (Optical Character Recognition), and counting) and visual reasoning tasks (such as mathematics, science, graphs, and puzzles). These can be categorized by broad classes, such as visual perception tasks and visual reasoning tasks, or by specific task training types, such as object detection and mathematics.
[0033] The first training model and the second training model are respectively one of the encoding model and the language model in the multimodal model. The terms "first" and "second" are used only to distinguish between models with activated parameters and models with frozen parameters, and are not used to indicate order, structure, etc. In this embodiment, the number of second training models with frozen parameters is at least one. Since the visual language model includes at least two encoding models and at least one language model, the number of first training models with activated parameters is at least one, and may also be multiple.
[0034] Model parameter activation refers to adjusting the model parameters during the backpropagation process of a training step. Parameter freezing refers to not adjusting the model parameters during the backpropagation process of a training step.
[0035] In this embodiment, different first training models requiring model parameter activation and second training models requiring model parameter freezing can be set for different training steps. Furthermore, for different training steps, some or all model parameters of the first training model can be activated, and different activation model parameters can be set for different training steps. For example, the language model has a total of 10B model parameters and 100 training steps. 1B model parameters can be activated in the first 50 training steps, and 0.9B model parameters can be activated in the last 50 training steps. The 1B and 0.9B model parameters can overlap, partially overlap, or not overlap.
[0036] Similarly, for different training task types, different first training models that require model parameter activation and different second training models that require model parameter freezing can be set. Furthermore, different activation model parameters can be set for different training task types.
[0037] In this embodiment, based on the training steps and training task type of the multimodal model's post-training, at least one second training model with frozen parameters is selected from at least two encoding models and at least one language model of the multimodal model, while the other models are used as first training models with activated parameters. This avoids mutual interference caused by the synchronous adjustment of model parameters during the joint post-training of the models in the multimodal model, which could lead to unstable model training or even training failure.
[0038] Understandably, taking a Visual Language Model (VLM) consisting of a Visual Encoding Model (ViT) and a Large Language Model (LLM) as an example, ViT is analogous to the eyes in human perception, used to extract features from images, while LLM is analogous to the brain, used to process and output the received information. If both ViT and LLM are trained simultaneously using reinforcement learning, meaning their model parameters are adjusted concurrently, the perception of visual content will constantly change. The "vision" and "brain" will be constantly shifting, ultimately making it difficult for the predictions output by the VLM to align well with the perceived visual content. The synchronous adjustment of model parameters between ViT and LLM leads to mutual interference, causing instability in the training process.
[0039] The technical solution of this embodiment freezes the parameters of one of the visual encoding model ViT and the large language model LLM. For example, ViT is used as the first training model for parameter activation, and LLM is used as the second training model for parameter freezing. This is equivalent to evolving "vision" without changing the "brain". Alternatively, LLM is used as the first training model for parameter activation, and ViT is used as the second training model for parameter freezing. This is equivalent to evolving the "brain" without changing "vision". Ultimately, the goal of automatically and stably training the visual language model after reinforcement learning is achieved.
[0040] S120. Based on the first training model and the second training model, predict the training data to obtain the prediction result. Based on the prediction result and the correct result in the training data, determine at least one feedback value, and update the activation model parameters in the first training model based on the feedback value.
[0041] The training data includes image or video data. In this embodiment, the training data input to the multimodal model may include image data or video data, and may also include text data to indicate information such as the type of training task. Furthermore, the training data may refer to a single piece of training data or a batch of multiple pieces of training data; this embodiment does not impose any limitations on this.
[0042] Whether the encoding model in the multimodal model is used as the first training model for parameter activation and the language model is used as the second training model for parameter freezing, or the language model is used as the first training model for parameter activation and the encoding model is used as the second training model for parameter freezing, the model training follows the normal model training process. The encoding model extracts features from each frame of the image or video data in the training data and encodes the text data in the training data. The language model, based on the training task type of the multimodal model, performs inference or perception processing on the text data and image feature data included in the training data to obtain the prediction result.
[0043] The correct result in the training data can be either the correct answer corresponding to the training data or the labeled result in the training data. The specific form of the correct result can match the type of training task. For example, the perception task can be the labeled result in the training data, and the reasoning task can be the correct answer corresponding to the training data.
[0044] Feedback values can be reward values and / or penalty values, which are important means of guiding model optimization and making its output more consistent with expectations. A reward value is a positive feedback signal used to encourage the model to take the desired action or produce high-quality output. A penalty value is a negative feedback signal used to correct the model's behavior that deviates from expectations and prevent adverse consequences. Feedback values can also be divided into effective feedback values (with reward and / or penalty values) and ineffective feedback values (without reward values).
[0045] In this embodiment, the reward value can be determined based on the reward function and / or reward model, based on the prediction result and the correct result in the training data. Correspondingly, the penalty value can be determined based on the penalty function and / or penalty model, based on the prediction result and the correct result in the training data. The specific type of the reward function and / or reward model can be set based on different training task types, and this embodiment does not impose any restrictions on this. For example, in a visual reasoning task, the reward function can be used to determine whether the output prediction result is correct. If it is, a reward value (i.e., a valid feedback value) is fed back; otherwise, 0 (i.e., an invalid feedback value) is fed back.
[0046] Furthermore, the feedback value can also be a comprehensive feedback value calculated based on at least two reward values and / or penalty values. For example, in a visual perception task, two reward functions can be set, such as the intersection-union ratio (IUGR) and the format accuracy rate. The IUGR represents the ratio between the area of the intersection region and the area of the union region between the output predicted box and the correct box, while the format accuracy rate represents the ratio of the accuracy of the output format of the predicted box to the format of the correct box in the training data of the visual perception task. Specifically, when calculating the feedback value based on at least two reward values and / or penalty values, different weights can be assigned to different reward values and / or penalty values.
[0047] In this embodiment, updating the activation model parameters in the first training model based on the feedback value is achieved through reinforcement learning, which converts the feedback value into the update direction of the activation model parameters, and then updates the activation model parameters through backpropagation.
[0048] Taking the Visual Language Model (VLM) comprising a Visual Encoding Model (ViT) and a Large Language Model (LLM) as an example, if the training data is an image containing the question mark "1×1=?", ViT generates image features, and LLM performs inference based on these features, ultimately outputting a prediction of "2". Based on the prediction and the correct result "1", the output is deemed incorrect. The reward function determines whether to deny the prediction or the penalty function determines whether to penalize it. If ViT and LLM are jointly trained, the negative feedback causes LLM to attempt to correct its inference method, and ViT to attempt to adjust its image feature extraction method. This simultaneous adjustment leads to mutual interference between the two models, resulting in unstable training. In this embodiment, however, since one of the ViT or LLM parameters is frozen (e.g., ViT parameters are frozen while LLM parameters are activated), the visual perception process remains unchanged, and only the inference process is adjusted based on the negative feedback from the training data. This avoids mutual interference between the two models and improves the stability of model training.
[0049] S130. Determine whether the preset training objective has been achieved. If yes, obtain the target multimodal model. If no, determine whether the training model switching condition is met. If not, repeat the prediction output based on the first training model and the second training model to obtain the prediction result. If yes, exchange the first training model and the second training model, and repeat the prediction output based on the first training model and the second training model to obtain the prediction result.
[0050] In this embodiment, the approach involves freezing the parameters of either the encoding model or the language model in the multimodal model, while activating the parameters of the remaining models and adjusting the parameters of the activated model. Therefore, in this embodiment, the frozen model can be temporarily fixed, and the first and second training models can be switched after the activated model has achieved the desired training effect; alternatively, the encoding model and language model can undergo multiple alternating periods of parameter freezing and activation training throughout the entire post-training process.
[0051] Accordingly, in the first scenario described above, that is, the model with frozen parameters is temporarily fixed, and the first training model is switched to the second training model after the model with activated parameters has achieved the desired training effect. The training model switching condition refers to the first training model having achieved the desired training effect. Furthermore, determining whether the training model switching condition is met can also include: judging whether the training model switching condition is met based on at least one of the following: training steps, training task type, amount of training data, and model performance of the multimodal model.
[0052] Specifically, determining whether the training model switching condition is met based on the training steps can be done by dividing the training steps for the encoding model and the language model separately, and then determining whether the training model switching condition is met based on the current training step and the training steps corresponding to the first training model. Taking a total of 100 training steps for the visual language model as an example, the training steps for the encoding model can be set to 40 steps and the training steps for the language model to be 60 steps. If the parameters are first frozen using the encoding model as the second training model and the parameters are activated using the language model as the first training model, then the training model switching condition is met when the current training step is 61.
[0053] Furthermore, for different training task types, different training steps can be set for the encoding model and the language model, thereby enabling the determination of training model switching conditions based on training steps and training task types.
[0054] The determination of whether the training model switching condition is met is based on the amount of training data. This can be achieved by ensuring that the amount of training data used to update the parameters of the first training model is greater than or equal to a first preset number. Furthermore, the first preset number can be calculated based on the total amount of training data, the model type of the first training model, and the type of training task. For example, if there are 1000 training data points, for a visual reasoning task, 40% of the training data can be allocated to the encoding model and 60% to the language model. If the parameters are first frozen using the encoding model as the second training model and then activated using the language model as the first training model, then the training model switching condition is met when the amount of training data for the first training model is greater than or equal to 600.
[0055] The performance of a multimodal model can be evaluated by the performance of feedback values during training. For example, it can be at least one of the following: the number of reward values is greater than or equal to a second preset number; the proportion of reward values to the total number of feedback values is greater than or equal to a first preset proportion threshold; the rate of change of the proportion of reward values is less than or equal to a first preset rate of change threshold.
[0056] In the first scenario described above, achieving the preset training objective can mean that both the encoding model and the language model in the multimodal model have reached the desired training effect. Furthermore, whether the preset training objective has been achieved can be determined based on at least one of the following: training steps, training task type, amount of training data, and the model performance of the multimodal model.
[0057] Specifically, determining whether a preset training objective has been achieved based on training steps could involve checking if the current step has reached the pre-set total number of training steps. Furthermore, different total training steps can be set for different training task types, thus enabling the determination of preset training objectives based on training steps and training task type. Determining whether a preset training objective has been achieved based on the amount of training data could involve checking if the amount of input training data is greater than or equal to a second preset number.
[0058] Determining whether the preset training objective has been achieved based on the performance of the multimodal model can be achieved by considering the feedback values of each encoding and language model during training, which all meet the following conditions: the number of reward values is greater than or equal to a second preset number; the proportion of reward values to the total number of feedback values is greater than or equal to a first preset proportion threshold; or, the rate of change of the reward value proportion is less than or equal to a first preset rate of change threshold. Alternatively, a comprehensive evaluation of the multimodal model's performance can be conducted, considering factors such as: performance metrics (e.g., accuracy, recall, mean squared error) no longer increasing; the multimodal model reaching the preset number of iterations or training steps; overfitting; and a decrease in the learning rate.
[0059] Correspondingly, in the second scenario described above, that is, during the post-model training process, when alternating between parameter freezing and parameter activation training for the encoding model and the language model, the training model switching condition can be determined by whether the current step has reached a pre-set switching training step. Furthermore, different switching training steps can be set for different training task types. Simultaneously, for the same training task type, a series of switching training steps can be set, thereby enabling the determination of the training model switching condition based on the training step and training task type. For example, it can be set to switch between the first and second training models every 10 training steps.
[0060] The training model switching condition can also be based on whether the number of training data used to train the current first training model is greater than or equal to a pre-set third preset number. For example, it can be set to input 100 training data points individually or in batches into the current first training model, and then switch between the first and second training models.
[0061] The training model switching condition can also involve setting multiple stages for the multimodal model's performance. When the multimodal model meets the performance evaluation value corresponding to the first stage, a switch is made between the first and second training models. After the switch, model training is re-run until the multimodal model meets the performance evaluation value corresponding to the first stage. This process is repeated until both the encoding and language models meet the performance evaluation value corresponding to the first stage after switching. This process is then repeated again to determine whether the multimodal model meets the performance evaluation value corresponding to the second stage, until the multimodal model meets the performance evaluation value corresponding to the final stage. The model performance evaluation value can be represented by the number of reward values, the proportion of reward values, and the rate of change of the reward value proportion, or by performance metrics such as accuracy, recall, and mean squared error.
[0062] In the second scenario described above, whether the preset training objective has been achieved can also be determined by at least one of the following: training steps, training task type, amount of training data, and model performance of the multimodal model. The process for determining the preset training objective is similar to that in the first scenario, and this embodiment does not impose any limitations on it.
[0063] Furthermore, after obtaining the target multimodal model, the method further includes: training steps and training task types based on the multimodal model post-training; activating the model parameters of the first training model and the second training model; predicting the training data based on the first training model and the second training model to obtain prediction results; determining at least one feedback value based on the prediction results and the correct results in the training data; and updating the activated model parameters in the first training model and the second training model based on the feedback value.
[0064] In this embodiment, a technical solution is also provided to further improve the model performance by continuing joint training of the encoding model and the language model after the training steps of the multimodal model have reached a certain training stage or the model performance has reached a certain standard.
[0065] Specifically, after achieving the aforementioned preset training objectives, for the obtained target multimodal model, both the encoding model and the language model are activated. Based on the training data, joint training of the encoding and language models continues, with synchronized adjustments to the activation parameters of both models. Since the above embodiments achieve independent training of the encoding and language models through separate parameter freezing and activation, resulting in automatic and stable reinforcement learning post-training effects, further joint training can further improve the performance of the multimodal model.
[0066] Furthermore, before performing the steps of this embodiment, the following may also be included:
[0067] S1. Generate metadata corresponding to each training data based on the training task type, data source, and validator service of each training data.
[0068] S2. Input each training data and its corresponding metadata into the multimodal model. The multimodal model then predicts and outputs the corresponding training data based on the metadata.
[0069] The data source indicates the training dataset to which the training data belongs. In this embodiment, the training datasets corresponding to different training task types can be the same or different. The validator service can be a set of predefined functions, models, rules, or standards used to determine whether the prediction results output by the model are correct, what the accuracy of the results is, whether the format is correct, whether the corresponding threshold is reached, what the corresponding feedback value is, what the weight of different reward functions is, etc. It is a service that can be deployed independently and can be asynchronously called with model training.
[0070] In this embodiment, metadata is generated for each training data, and the metadata and training data are input together into a multimodal model. The multimodal model makes predictions on the training data according to the training task type in the metadata, and generates feedback values based on the prediction results and the correct results in the training data through the validator service specified in the metadata.
[0071] The technical solution of this embodiment enables multimodal models to perform unified training on multiple training tasks without the need to configure model parameters separately for each training task or to train the model separately. This improves model training efficiency, enhances the refinement, stability, and effectiveness of model training, and reduces model training costs.
[0072] Furthermore, S2 can include:
[0073] S20. Input at least two training data points of at least two training task types and their corresponding metadata into the multimodal model as a batch. The multimodal model performs prediction output on each training data point corresponding to the metadata based on the metadata to obtain at least two prediction results.
[0074] Accordingly, determining at least one feedback value based on the prediction result and the correct result in the training data in S120 may further include: S121, verifying each prediction result according to each prediction result, the correct result in the corresponding training data, and the verifier service in the corresponding metadata, to obtain at least one feedback value corresponding to the training data.
[0075] For multiple training data sets and their corresponding metadata corresponding to various training task types, they are input into the multimodal model in the same batch. The multimodal model performs the corresponding prediction task on the training data based on the training task type in the metadata, and outputs the prediction result corresponding to each training data set. At the same time, through the validator service in the metadata of the training data, a feedback value is calculated for the prediction result and the correct result corresponding to each training data set.
[0076] In this embodiment, multiple training data and their corresponding metadata corresponding to various training task types are input into the multimodal model in the same batch, realizing unified training for multiple training tasks. This eliminates the need to configure model parameters separately for each target training task, based on the training data in each training dataset, and train the model separately, thereby improving model training efficiency. At the same time, the model parameters of various task types can influence and promote each other, avoiding the technical problem of individual training dragging each other down, and improving model performance.
[0077] Specifically, S121 can be implemented through the following steps: generating a data verification request for each training data based on the metadata, correct result, and prediction result of each training data; sending the data verification request for each training data to the server, so that the server determines and calls the target validator corresponding to each training data to respond based on the data verification request for each training data, and obtains at least one feedback value corresponding to the data verification result of each training data; wherein, multiple validators are pre-deployed on the server.
[0078] The validator is used to calculate or judge data verification requests and generate feedback values as data verification results. Therefore, in this embodiment, the calculation of feedback values is delegated to specialized validators. Each validator is responsible for handling a specific verification request or a set of verification requests, rather than defining a single reward function for each training task in a general way. The granularity of verification calculation is finer, and different validators can be flexibly set for verification based on the different training task types, training data quality, and prediction difficulty of each training data, which improves the accuracy of verification and is flexibly applicable to different model training needs.
[0079] In this embodiment, the validator service is implemented through validators deployed on the server. The metadata of each training data also includes a validator identifier corresponding to that training data. Based on the metadata, correct result, and prediction result of each training data, a data verification request is generated for each training data. The server receiving the data verification request can parse this request to determine the metadata, correct result, and prediction result corresponding to the training data, and determine the target validator matching the training data based on the validator identifier in the metadata to respond.
[0080] The technical solution of this invention, based on the training steps and training task types of post-training of multimodal models, involves determining the activation parameters of a first training model and freezing the parameters of a second training model among at least two encoding models and at least one language model in the multimodal model. Based on the first and second training models, prediction outputs are made on the training data to obtain prediction results. Based on the prediction results and correct results in the training data, a feedback value is determined, and the activated model parameters in the first training model are updated based on the feedback value. This updating of the first training model's parameters is repeated until the training model switching conditions are met. The first and second training models are then swapped, and model training continues until a preset training objective is achieved, resulting in the target multimodal model. This solves the problem of instability or even training failure during post-training of multimodal models, especially visual-language models, in the prior art, improving the training efficiency and reducing the cost of post-training of multimodal models.
[0081] Example 2
[0082] Figure 2 This is a flowchart of a multimodal model post-training method provided in Embodiment 2 of the present invention. Based on the above embodiments, the present invention adds a step of updating the parameters of the encoding model based on auxiliary training data after content hiding of the training data.
[0083] like Figure 2 As shown, the method includes:
[0084] S210. Perform content hiding processing on at least one training data to obtain auxiliary training data.
[0085] Content hiding can involve hiding a portion of the training data. For example, if the training data consists of frames from an image or video, the hidden portion refers to a specific area within the image. Content hiding can be achieved through methods such as mosaicking, overlaying based on a pre-defined image, or adjusting transparency; this embodiment does not impose any limitations on these methods. Auxiliary training data refers to the training data obtained after content hiding. The same training data can correspond to one or more auxiliary training data sets.
[0086] In this embodiment, the size of the region to be hidden and its position in the image can be randomly generated. Alternatively, a random algorithm can be used to randomly generate the region size and image position, or different region sizes and different image positions can be preset, and a random algorithm can be used to randomly select from the preset region size and image position.
[0087] The size and location of the region to be hidden in the image can be determined by first defining the region of interest and then randomly generating it within the region of interest. The region of interest can be a predefined central region of the image, or it can be the boundary of multiple image regions identified based on parameters such as pixels, color, and transparency in the image. Alternatively, the region of interest can be divided according to different training task types.
[0088] The size of the region to be hidden and its position in the image can be flexibly set according to different training tasks. For example, for object detection tasks, the content of the target object region in the training data image can be hidden.
[0089] It should be noted that this embodiment uses visual data such as image data or video data as training data as an example. Training data can also be text data, audio data, etc. Accordingly, when the training data is text data, the part to be hidden refers to certain characters in the text data. Similarly, the length of the hidden characters and their positions in the text data can be determined in a similar manner to the above embodiment. Likewise, when the training data is audio data, the part to be hidden refers to audio segments in the audio data, and the length of the hidden audio segments and their positions in the audio data can also be determined in a similar manner to the above embodiment.
[0090] In this embodiment, content hiding can be performed on all training data or on a portion of the training data. When performing content hiding on a portion of the training data, the data can be selected randomly, or, based on the different training task types, training data matching the training task type can be selected for content hiding.
[0091] Understandably, the original training method for pre-training encoding models was contrastive learning, essentially aiming to align the semantics of images and text in the training data. Therefore, the encoding model learns what descriptive text corresponds to each training data point; that is, the encoding model describes the image content information of the training data through sentences. This is essentially a static feature extraction method. However, in the reinforcement learning training process of multimodal models, the task objective is to maximize the reward value of a certain task. In this case, the training method of the encoding model is to learn which features in the training data are helpful in solving the problem. As the training task type changes, the so-called "features helpful in solving the problem" also changes dynamically. Therefore, for the same training data, the detailed features that the encoding model should focus on will change depending on the training task type and the input problem. This makes the original static feature extraction method of the encoding model unable to adapt to the dynamic changes in the post-training methods of multimodal reinforcement learning, affecting the stability of the model training process.
[0092] In this embodiment, content hiding is performed on the training data to obtain auxiliary training data. This auxiliary training data, along with the training data, is then input into the encoding model for training. This self-supervised training method, which hides the training data, allows the encoding model to fully notice and learn richer, more localized details within the training data. It can generate not only an overall description of the training data during encoding but also detailed descriptions of specific aspects of the training data. This powerful representational ability enables the encoding model to provide more effective and discriminative visual information when facing reinforcement learning tasks that require dynamic attention to different regions. The encoding model can better adapt to the training process after reinforcement learning, making model training more stable and thus improving training efficiency and performance.
[0093] Furthermore, S210 may include:
[0094] Perform content hiding processing on at least two different regions of at least one training data point to obtain at least two auxiliary training data points; and / or,
[0095] Based on the training task type of the training data, determine the target regions in the training data that require content hiding processing;
[0096] The target region in the training data is hidden to obtain auxiliary training data corresponding to the training data.
[0097] In this embodiment, different regions can be determined for the same training data. Similarly, the determined different regions can be determined according to the random generation method mentioned in the above embodiments, the method of first determining the region of interest and then randomly generating regions within the region of interest, and the method of determining regions based on the training task type. This embodiment does not limit the method or number of different regions to be determined.
[0098] Content hiding is performed at least twice on different regions of the same training data. This can be done either by hiding each region separately or by hiding different combinations of regions. For example, if three regions A, B, and C are identified in the training data, the final auxiliary training data can be obtained by hiding A, B, C, AB, AC, BC, and ABC respectively. That is, the training data will ultimately yield seven auxiliary training data sets.
[0099] In this embodiment, multiple auxiliary training data are generated based on different regions of the same training data. This allows the encoding model to fully learn the various detailed features within the training data, rather than learning only the detailed features once per training data point. This method of generating auxiliary training data further improves the detailed representation ability of the encoding model, thereby enhancing the stability of the model's subsequent training process.
[0100] In this embodiment, the target region in the training data that needs content hiding can be determined according to the training task type of the training data. For example, when the training task type is solving math problems, the mathematical formula region or character region in the training data image can be hidden as the target region. When the training task type is object recognition, the target object region in the training data image can be hidden as the target region.
[0101] In this embodiment, this method of determining hidden regions based on training task type can more efficiently guide the encoding model to pay attention to the target region in the image corresponding to each training task type, thereby improving training efficiency.
[0102] Furthermore, when the same training data is used to train models for different training task types, different target regions can be determined for the training data based on different training task types, and different auxiliary training data can be generated. This embodiment does not impose any restrictions on this.
[0103] S220. Based on the training data and the corresponding auxiliary training data, supervised training is performed on at least one encoding model in the multimodal model, and the parameters of the encoding model are updated.
[0104] In this embodiment, image or video data is used as the training data, and the encoding model is a visual encoding model. The encoding model can also include text encoding models and audio encoding models; this embodiment does not impose any limitations on this.
[0105] In this embodiment, training data and corresponding auxiliary training data are combined into sample pairs, which are then input into the encoding model to train the encoding model and update its model parameters.
[0106] Furthermore, S220 may include:
[0107] S221. Input the training data and the corresponding auxiliary training data into the encoding model;
[0108] S222. Extract the feature codes of the training data and the corresponding auxiliary training data respectively through the encoding model, and calculate the deviation value between the feature codes of the training data and the feature codes of the auxiliary training data.
[0109] S223. Update the parameters of the coding model based on the deviation value.
[0110] In this embodiment, the deviation between the feature encoding of the training data and the feature encoding of the auxiliary training data represents the content features of the hidden parts of the training data. Therefore, based on this deviation, the encoding model can fully notice the feature encoding corresponding to the details in the training data, thereby improving the detail representation capability of the encoding model.
[0111] S230. Training steps and training task types based on multimodal model post-training: activating the model parameters of the first training model and freezing the model parameters of at least one second training model in the multimodal model.
[0112] S240. Based on the first training model and the second training model, predict the training data to obtain the prediction result. Based on the prediction result and the correct result in the training data, determine at least one feedback value, and update the activation model parameters in the first training model based on the feedback value.
[0113] S250: Determine whether the preset training objective has been achieved. If yes, proceed to S280; otherwise, proceed to S260.
[0114] S260. Determine whether the training model switching conditions are met. If yes, execute S270; otherwise, return to execute S240.
[0115] S270. Swap the first training model with the second training model, and return to execute S240.
[0116] S280, Obtain the target multimodal model.
[0117] The process of training the multimodal model after reinforcement learning has been described in the above embodiments, and will not be repeated here.
[0118] The technical solution in this embodiment involves hiding the content of the training data and constructing auxiliary training data before training the multimodal model using reinforcement learning. Based on this auxiliary and training data, the encoding model undergoes self-supervised training. This allows the encoding model to fully consider the detailed features in the training data, generating not only an overall description of the training data but also detailed descriptions of local details. Subsequent post-model training is then performed based on the supervised-trained encoding and language models. This enables the encoding model to better adapt to the post-reinforcement learning training mode, providing more effective and discriminative visual information when facing reinforcement learning tasks that require dynamic attention to different regions. This improves the stability and efficiency of the model training process.
[0119] Example 3
[0120] Figure 3 This is a schematic diagram of a multimodal model post-training device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes:
[0121] The parameter freezing module 310 is used to activate the model parameters of the first training model and freeze the model parameters of at least one second training model in the multimodal model, based on the training steps and training task types of the multimodal model post-training; wherein the multimodal model includes at least two encoding models and at least one language model, and the first training model and the second training model are respectively one of the encoding model and the language model;
[0122] The parameter update module 320 is used to predict the training data based on the first training model and the second training model to obtain the prediction result, determine at least one feedback value based on the prediction result and the correct result in the training data, and update the activation model parameters in the first training model based on the feedback value; wherein, the training data includes image or video data;
[0123] The condition judgment module 330 is used to determine whether the preset training objective has been achieved. If so, the target multimodal model is obtained; if not, it is determined whether the training model switching condition is met. If not, the prediction output based on the first training model and the second training model is repeatedly executed to obtain the prediction result; if so, the first training model and the second training model are swapped, and the prediction output based on the first training model and the second training model is repeatedly executed to obtain the prediction result.
[0124] The technical solution of this invention, based on the training steps and training task types of multimodal model post-training, involves determining the activation parameters of a first training model and freezing the parameters of a second training model within at least two encoding models and at least one language model in the multimodal model. Based on the first and second training models, prediction outputs are made on the training data to obtain prediction results. Based on the prediction results and correct results in the training data, a feedback value is determined, and the activated model parameters in the first training model are updated based on the feedback value. This updating of the first training model's parameters is repeated until the training model switching conditions are met. The first and second training models are then swapped, and model training continues until a preset training objective is achieved, resulting in the target multimodal model. This solves the problem of unstable training or even training failure during joint training of multimodal models, especially visual-language models, in the prior art, improving the training efficiency and reducing the cost of multimodal model post-training.
[0125] Based on the above embodiments, optionally, the condition judgment module 330 includes:
[0126] The training model switching condition judgment unit is used to determine whether the training model switching condition is met based on at least one of the training steps, training task type, training data quantity, and multimodal model performance.
[0127] Optionally, based on the above embodiments, the apparatus further includes:
[0128] An auxiliary training data determination module is used to perform content hiding processing on at least one training data to obtain auxiliary training data.
[0129] The encoding model parameter update module is used to perform supervised training on at least one encoding model in the multimodal model based on training data and corresponding auxiliary training data, and to update the parameters of the encoding model.
[0130] Based on the above embodiments, optionally, the auxiliary training data determination module includes:
[0131] The first auxiliary training data determination unit is used to perform content hiding processing on different regions of at least one training data at least twice to obtain at least two auxiliary training data.
[0132] The target region determination unit is used to determine the target region in the training data that needs to be hidden based on the training task type of the training data.
[0133] The second auxiliary training data determination unit is used to hide the target region in the training data to obtain auxiliary training data corresponding to the training data.
[0134] Based on the above embodiments, optionally, the encoding model parameter update module includes:
[0135] A data input unit is used to input training data and corresponding auxiliary training data into the encoding model;
[0136] The feature encoding deviation value determination unit is used to extract the feature codes of the training data and the corresponding auxiliary training data through the encoding model, and calculate the deviation value between the feature codes of the training data and the feature codes of the auxiliary training data.
[0137] The coding model parameter update unit is used to update the parameters of the coding model according to the deviation value.
[0138] Optionally, based on the above embodiments, the apparatus further includes:
[0139] The metadata generation module is used to generate metadata for each training data based on the training task type, data source, and validator service of each training data.
[0140] The metadata input module is used to input each training data and its corresponding metadata into the multimodal model. The multimodal model makes predictions and outputs based on the metadata for each training data corresponding to the metadata.
[0141] Based on the above embodiments, optionally, the metadata input module includes:
[0142] The prediction result determination unit is used to input at least two training data of at least two training task types and their corresponding metadata into a multimodal model as a batch. The multimodal model performs prediction output on each training data corresponding to the metadata based on the metadata to obtain at least two prediction results.
[0143] The parameter update module 320 includes:
[0144] The feedback value determination unit is used to verify each prediction result based on each prediction result, the correct result in the corresponding training data, and the validator service in the corresponding metadata, and to obtain at least one feedback value corresponding to the training data.
[0145] Based on the above embodiments, optionally, the feedback value determination unit is specifically used for:
[0146] Based on the metadata, correct results, and prediction results of each training data set, a data validation request is generated for each training data set.
[0147] Each training data data validation request is sent to the server, so that the server determines and calls the target validator corresponding to each training data data to respond based on the data validation request of each training data data, and obtains at least one feedback value corresponding to the data validation result of each training data data; wherein, multiple validators are pre-deployed on the server.
[0148] Optionally, based on the above embodiments, the apparatus further includes:
[0149] The model parameter activation module is used to activate the model parameters of the first and second training models based on the training steps and training task types of the multimodal model post-training.
[0150] The activation model parameter update module is used to predict the training data based on the first training model and the second training model to obtain the prediction result, determine at least one feedback value based on the prediction result and the correct result in the training data, and update the activation model parameters in the first training model and the second training model based on the feedback value.
[0151] The multimodal model post-training device provided in the embodiments of the present invention can execute the multimodal model post-training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0152] Example 4
[0153] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0154] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0155] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0156] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as multimodal model post-training methods.
[0157] In some embodiments, the multimodal model post-training method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the multimodal model post-training method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the multimodal model post-training method by any other suitable means (e.g., by means of firmware).
[0158] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0159] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable multimodal model post-training device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0160] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0162] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0163] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0164] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0165] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for post-training a multimodal model, characterized in that, include: The training steps and training task types based on the multimodal model post-training include activating the model parameters of the first training model and freezing the model parameters of at least one second training model in the multimodal model; wherein the multimodal model includes at least two encoding models and at least one language model, and the first training model and the second training model are respectively one of the encoding model and the language model; Based on the first training model and the second training model, the training data is predicted and output to obtain the prediction result. Based on the prediction result and the correct result in the training data, at least one feedback value is determined, and the activation model parameters in the first training model are updated based on the feedback value; wherein, the training data includes image or video data; Determine whether the preset training objective has been achieved. If so, obtain the target multimodal model. If not, determine whether the training model switching condition is met. If not, repeat the prediction output based on the first training model and the second training model to obtain the prediction result. If so, exchange the first training model and the second training model, and repeat the prediction output based on the first training model and the second training model to obtain the prediction result.
2. The method according to claim 1, characterized in that, Determine whether the conditions for switching training models are met, including: The training model switching conditions are determined based on at least one of the following: the training steps, the type of training task, the amount of training data, and the performance of the multimodal model.
3. The method according to claim 1, characterized in that, Before activating the model parameters of the first trained model, the training steps and training task types based on the multimodal model post-training also include: Content hiding is performed on at least one training data to obtain auxiliary training data; Based on the training data and the corresponding auxiliary training data, supervised training is performed on at least one encoding model in the multimodal model, and the parameters of the encoding model are updated.
4. The method according to claim 3, characterized in that, The step of performing content hiding processing on at least one training data to obtain auxiliary training data includes: Perform content hiding processing on at least two different regions of at least one training data point to obtain at least two auxiliary training data points; and / or, Based on the training task type of the training data, determine the target regions in the training data that require content hiding processing; The target region in the training data is hidden to obtain auxiliary training data corresponding to the training data.
5. The method according to claim 3, characterized in that, Based on training data and corresponding auxiliary training data, supervised training is performed on at least one encoding model in the multimodal model, and the parameters of the encoding model are updated, including: The training data and the corresponding auxiliary training data are input into the encoding model; The feature codes of the training data and the corresponding auxiliary training data are extracted using the encoding model, and the deviation between the feature codes of the training data and the feature codes of the auxiliary training data is calculated. The parameters of the coding model are updated based on the deviation value.
6. The method according to claim 1, characterized in that, Before activating the model parameters of the first trained model, the training steps and training task types based on the multimodal model post-training also include: Based on the training task type, data source, and validator service of each training data, generate metadata corresponding to each training data; Each training data point and its corresponding metadata are input into a multimodal model, which then makes a prediction output based on the metadata for each training data point corresponding to the metadata.
7. The method according to claim 6, characterized in that, Each training data point and its corresponding metadata are input into a multimodal model. The multimodal model then performs a prediction output based on the metadata for each training data point corresponding to the metadata, including: At least two training data sets of at least two training task types and their corresponding metadata are input into a multimodal model as a batch. The multimodal model makes a prediction output for each training data set corresponding to the metadata based on the metadata, and obtains at least two prediction results. The step of determining at least one feedback value based on the prediction result and the correct result in the training data further includes: Each prediction result is verified based on the correct result in the corresponding training data and the validator service in the corresponding metadata, and at least one feedback value corresponding to the training data is obtained.
8. The method according to claim 7, characterized in that, The step of verifying each prediction result based on each prediction result, the correct result in the corresponding training data, and the validator service in the corresponding metadata, to obtain at least one feedback value corresponding to the training data, includes: Based on the metadata, correct results, and prediction results of each training data set, a data validation request is generated for each training data set. Each training data data validation request is sent to the server, so that the server determines and calls the target validator corresponding to each training data data to respond based on the data validation request of each training data data, and obtains at least one feedback value corresponding to the data validation result of each training data data; wherein, multiple validators are pre-deployed on the server.
9. The method according to any one of claims 1-8, characterized in that, After obtaining the target multimodal model, the process further includes: Based on the training steps and training task types of multimodal model post-training, the model parameters of the first and second training models are activated. Based on the first training model and the second training model, the training data is predicted and output to obtain the prediction result. Based on the prediction result and the correct result in the training data, at least one feedback value is determined, and the activation model parameters in the first training model and the second training model are updated based on the feedback value.
10. A multimodal model post-training device, characterized in that, include: The parameter freezing module is used to activate the model parameters of the first training model and freeze the model parameters of at least one second training model in the multimodal model, based on the training steps and training task types of the multimodal model post-training. The multimodal model includes at least two encoding models and at least one language model, and the first training model and the second training model are one of the encoding model and the language model, respectively. The parameter update module is used to predict the training data based on the first training model and the second training model to obtain the prediction result, determine at least one feedback value based on the prediction result and the correct result in the training data, and update the activation model parameters in the first training model based on the feedback value; wherein, the training data includes image or video data; The condition judgment module is used to determine whether the preset training objective has been achieved. If so, the target multimodal model is obtained; if not, it determines whether the training model switching condition is met. If not, it repeats the prediction output based on the first training model and the second training model to obtain the prediction result; if so, it swaps the first training model and the second training model and repeats the prediction output based on the first training model and the second training model to obtain the prediction result.
Citation Information
Patent Citations
Living body detection model training method and device, medium and electronic equipment
CN116721315A
Multi-modal pre-training model training method and device and multi-modal data processing method and device
CN116861995A
Generative recommendation large model generation method and device and nonvolatile storage medium
CN118467843A
Ultrasonic multi-mode pre-training method and device, computer equipment and storage medium
CN119886266A
Multi-modal large model training method and device
CN120068981A