A multimodal model multi-task unified training method and device

CN121092994BActive Publication Date: 2026-09-25SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511201680.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-09-25
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

[0003]但如此设置,一方面割裂了不同模态、不同任务训练数据对模型能力提升的相互影响,使得模型在训练另一方面任务能力时可能会对之前训练的任务能力造成灾难性遗忘,只能训练出某一或某几项任务能力突出的模型,而无法实现多任务能力均突出的模型;且每一任务能力单独训练,训练效率低,成本高

Benefits of technology

[0009]本申请实施例提供的一种多模态模型多任务统一训练方法及装置,所述方法包括:根据本次训练的训练类型、任务类型以及本次训练所使用的训练样本数据集的数据属性信息中的至少一项,确定对训练样本数据集中训练样本数据进行结构化处理的目标数据格式模版;其中,所述任务类型包括至少两种任务类型,所述训练样本数据包括图片或视频;根据所述目标数据格式模版对所述训练样本数据集中的每一训练样本数据进行分析处理,确定所述每一训练样本数据的元数据;根据所述元数据以及对应的所述训练样本数据集中的每一训练样本数据对初始多模态模型进行多任务统一训练,得到目标多模态模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121092994B_ABST
    Figure CN121092994B_ABST
Patent Text Reader

Abstract

The application provides a multimodal model multitask unified training method and device, the method comprising: determining a target data format template for structuring processing of training sample data in a training sample data set according to at least one of a training type, a task type of the present training, and data attribute information of the training sample data set used in the present training; analyzing and processing each training sample data in the training sample data set according to the target data format template to determine metadata of each training sample data; and performing multitask unified training on an initial multimodal model according to the metadata and each training sample data in the corresponding training sample data set to obtain a target multimodal model. In this way, the application unifies the training sample data and the corresponding metadata to input the initial multimodal model for training to generate the target multimodal model, which can not only improve the training efficiency and reduce the cost, but also improve the model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a unified training method and apparatus for multimodal models and multitasking. Background Technology

[0002] Existing large language models are evolving towards multimodal and multi-task processing capabilities. This means that training large language models, including pre-training and post-reinforcement learning training, requires training data encompassing multiple modalities and tasks. Multiple modalities include training data containing images, text, audio, and video; multiple tasks include prediction-intensive tasks such as visual prediction tasks (including math, science, charting, and puzzle tasks); and perception-intensive tasks such as visual perception tasks (including detection, grounding, counting, and optical character recognition (OCR) tasks). Current techniques typically select different training datasets and configure corresponding training parameters based on different task requirements, training the model separately on a specific task capability to acquire and enhance that capability.

[0003] However, this setup has several drawbacks. First, it severs the mutual influence of training data from different modalities and tasks on model performance improvement. This can lead to catastrophic forgetting of previously trained task capabilities when training for another task, resulting in a model that excels in only one or a few tasks, rather than a model that excels across multiple tasks. Furthermore, training each task capability separately is inefficient and costly. Second, the training datasets for each task capability vary in quality, difficulty, content, and structure. Directly using the training data in these datasets and a uniform model parameter configuration for each task not only results in low training efficiency and poor model stability, but also leads to identical prediction biases on both high- and low-difficulty training data. Consequently, the model's performance improves slowly on high-difficulty training data, resulting in mediocre model performance. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a multi-modal model multi-task unified training method and apparatus, which generates a target multimodal model by uniformly inputting training sample data and corresponding metadata into the initial multimodal model for multi-task unified training. This not only improves training efficiency and reduces costs, but also improves model performance.

[0005] This application provides a unified training method for multi-modal models across multiple tasks, the method comprising: Based on at least one of the following: training type, task type, and data attribute information of the training sample dataset used in this training, determine the target data format template for structuring the training sample data in the training sample dataset; wherein, the task type includes at least two task types, and the training sample data includes images or videos; The metadata of each training sample in the training sample dataset is analyzed and processed according to the target data format template to determine the metadata of each training sample. The initial multimodal model is trained using the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model.

[0006] This application embodiment also provides a multimodal model multi-task unified training device, the training device comprising: The determination module is used to determine the target data format template for structuring the training sample data in the training sample dataset based on at least one of the following: training type, task type, and data attribute information of the training sample dataset used in this training; wherein, the task type includes at least two task types, and the training sample data includes images or videos; The processing module is used to analyze and process each training sample data in the training sample dataset according to the target data format template, and determine the metadata of each training sample data. The training module is used to perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model.

[0007] This application embodiment also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the training method described above are performed.

[0008] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the training method described above.

[0009] This application provides a method and apparatus for unified multi-task training of a multimodal model. The method includes: determining a target data format template for structured processing of training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training; wherein the task type includes at least two task types, and the training sample data includes images or videos; analyzing and processing each training sample data in the training sample dataset according to the target data format template to determine the metadata of each training sample data; and performing unified multi-task training on an initial multimodal model based on the metadata and the corresponding training sample data in the training sample dataset to obtain a target multimodal model. In this way, this application uniformly inputs each training sample data and its corresponding metadata in the training sample dataset into the initial multimodal model for training, generating the target multimodal model, without having to configure model parameters separately for each target training task, based on the training data in each training dataset, and train the model separately. This can improve model training efficiency, improve the refinement of model training, improve model training stability and model training effect, and also reduce model training cost.

[0010] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a multi-modal model multi-task unified training method provided in this application embodiment; Figure 2 This is one of the structural schematic diagrams of a multimodal model multi-task unified training device provided in an embodiment of this application; Figure 3 This is a second schematic diagram of the structure of a multimodal model multi-task unified training device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0014] First, the applicable scenarios for this application will be introduced. This application can be applied to the field of computer technology.

[0015] Research has revealed that existing large language models are evolving towards multimodal and multi-task processing capabilities. This means that training large language models, including pre-training and post-reinforcement learning training, requires training data encompassing multiple modalities and tasks. Multiple modalities include training data containing images, text, audio, and video; multiple tasks include prediction-intensive tasks such as visual prediction tasks like math, science, charting, and puzzle tasks; and perception-intensive tasks such as visual perception tasks like object detection, grounding, counting, and optical character recognition (OCR). Current techniques typically select different training datasets and configure corresponding training parameters based on different task requirements, training the model separately for a specific task capability to acquire and enhance that capability.

[0016] However, this setup has several drawbacks. First, it severs the mutual influence of training data from different modalities and tasks on model performance improvement. This can lead to catastrophic forgetting of previously trained task capabilities when training for another task, resulting in a model that excels in only one or a few tasks, rather than a model that excels across multiple tasks. Furthermore, training each task capability separately is inefficient and costly. Second, the training datasets for each task capability vary in quality, difficulty, content, and structure. Directly using the training data in these datasets and a uniform model parameter configuration for each task not only results in low training efficiency and poor model stability, but also leads to identical prediction biases on both high- and low-difficulty training data. Consequently, the model's performance improves slowly on high-difficulty training data, resulting in mediocre model performance.

[0017] Based on this, embodiments of this application provide a multimodal model multi-task unified training method and apparatus to improve training efficiency, improve model performance and reduce training costs.

[0018] Please see Figure 1 , Figure 1 This is a flowchart illustrating a multi-modal model multi-task unified training method provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the method includes: S101. Based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training, determine the target data format template for structuring the training sample data in the training sample dataset. S102. Analyze and process each training sample data in the training sample dataset according to the target data format template to determine the metadata of each training sample data. S103. Perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model.

[0019] The steps in the embodiments of this application are described in detail below.

[0020] For step S101, the task type includes at least two task types, and the training sample data includes images or videos.

[0021] The training type can be determined based on the training stage. For example, training types may include pre-training, fine-tuning training, and post-training based on reinforcement learning. The data format requirements for different training types may differ.

[0022] The training tasks can include, for example, visual prediction tasks such as Math, Science, Charting, and Puzzle; and visual perception tasks such as Detection, Grounding, Counting, and Optical Character Recognition (OCR). The data format requirements may differ for different task types.

[0023] The data attribute information of the training sample dataset may include data source, data type, data quality, etc. The data format requirements corresponding to the data attribute information of different training sample datasets may be different.

[0024] In addition, it is necessary to determine the initial multimodal model to be used in this training. The initial multimodal model can be determined based on the data types of the input and output. A multimodal model refers to an artificial intelligence system that can simultaneously process and understand multiple different types of information (modalities), such as an artificial intelligence system that can simultaneously process and understand text, images, audio, and video.

[0025] For example, the initial multimodal model may include a large language model of text + speech, a large language model of text + vision, a large language model of text + speech + vision, etc., wherein the data format requirements for the training process of different initial multimodal models may be different.

[0026] The target data format template refers to a predefined structure used to standardize the storage and organization of training sample data. It specifies which metadata should be included in each training sample data entry and how this metadata should be represented.

[0027] As can be seen from the above, the data format requirements for different initial multimodal models can be different, the data format requirements for different training types can be different, the data format requirements for different task types can be different, and the data format requirements for the data attribute information of different training sample datasets can be different.

[0028] For example, the data format requirements for each are illustrated below. Firstly, the data format requirements for the initial multimodal model may include: for a large language model consisting of text and speech, setting target data format requirements such as whether to include mathematical formulas, sound type (human voice, instrumental sound), language type (Chinese, English), and coding rate; for a large language model consisting of text and vision, setting target data format requirements such as image type (people, objects, landscapes), whether to include mathematical formulas, and language type (Chinese, English). Secondly, the data format requirements for training types may include: different training types correspond to different configuration loss functions or reward functions, penalty functions, and their corresponding weight values ​​and thresholds. Thirdly, the data format requirements corresponding to the task type may include: for question-and-answer tasks, specific target data format requirements such as the number of question-and-answer rounds can be set; for generation tasks, specific target data format requirements such as the size of the generated content and the difficulty value can be set; for target detection tasks in visual perception tasks, specific target data format requirements such as the number of targets to be identified, the attributes of the targets to be identified, and the difficulty value of the identification can be set; for localization tasks, specific target data format requirements such as the number of targets to be located, the attributes of the targets to be located, and the difficulty of localization can be set, etc.

[0029] Continuing with step S101, various examples of the target data format template determination process are provided in the embodiments of this application.

[0030] In the first embodiment provided in this application, determining the target data format template for structuring the training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training includes: S10111. Analyze each item in the data attribute information of the training type, task type, and training sample dataset used in this training, and determine the corresponding initial template requirements.

[0031] S10112. Integrate all initial template requirements and determine the target template requirements.

[0032] S10113. Based on the target template requirements, determine the target data format template for structuring the training sample dataset.

[0033] For step S10111, the training type, task type, and data attribute information of the training sample dataset used in this training each have their own corresponding initial template requirements, and the corresponding initial template requirements may be different.

[0034] Here, the initial template requirements include data dimension requirements.

[0035] For step S10112, this step may include: combining all initial template requirements, determining all data dimension requirements, and deleting duplicate data dimension requirements to obtain the target template requirements.

[0036] Regarding step S10113, this step may include searching for a template in a pre-built template database or generating a template by combining templates according to the target template requirements, to obtain the target data format template for this structured processing of the training sample dataset.

[0037] In the second embodiment provided in this application, determining the target data format template for structuring the training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training includes: S10121. Based on the training type of this training, determine the first data format template corresponding to the training type.

[0038] S10122. Based on the task type of this training, determine the second data format template corresponding to the task type.

[0039] S10123. Based on the data attribute information of the training sample dataset used in this training, determine the third data format template corresponding to the data attribute information of the training sample dataset.

[0040] S10124. Determine at least one of the first data format template, the second data format template, and the third data format template as the target data format template.

[0041] The first data format template, the second data format template, and the third data format template mentioned in steps S10121-S10124 can be extracted from a pre-built template database based on their respective corresponding data (the type of training this time, the type of task performed in this training, and the data attribute information of the training sample dataset used in this training).

[0042] Specifically, based on predetermined selection rules, specific templates from the first data format template, the second data format template, and the third data format template can be determined as the target data format template.

[0043] For example, identify the data dimensions of the first, second, and third data format templates, and determine the data format template with the most data dimensions as the target data format template. Alternatively, obtain the top N (N≤3) data format templates with the most data dimensions and determine them as the target data format template. The determined target data format template must meet all the training parameter configuration requirements for training the multi-dimensional model. The specific determination method can be adaptively determined and is not limited here.

[0044] In Embodiment 3 of this application, determining the target data format template for structuring the training sample data in the training sample dataset based on at least one of the following: training type, task type, and data attribute information of the training sample dataset used in this training, includes: S10131. Determine the training type, task type, and priority of each reference data item in the data attribute information of the training sample dataset used in this training.

[0045] S10132. Obtain the candidate data format template corresponding to the reference data with the highest priority, and determine whether the candidate data format template meets the requirements of the target template.

[0046] S10133. If the target template requirements are met, the candidate data format template corresponding to the reference data with the highest priority is determined as the target data format template.

[0047] S10134. If the target template requirements are not met, obtain the candidate data format template corresponding to the reference data of the next priority, and determine whether the target template requirements are met until the requirements are met. Then, determine the multiple candidate data format templates obtained as the target data format template.

[0048] For step S10131, the priority of each reference data item can be determined in advance.

[0049] For step S10132, for example, the target template requires that data dimension information can be set. If the data dimension of the candidate data format template is the same as or more than the data dimension set by the target template, it is considered to meet the target template requirements, and step S10133 is executed; otherwise, step S10134 is executed.

[0050] Other requirements can also be set for the target template, which are not limited here.

[0051] Regarding step S10133, the target data format template here includes a template. Template determination is highly efficient.

[0052] Regarding step S10134, the target data format template includes multiple templates. Thus, the candidate data format templates for each reference data item in the data attribute information of the training sample dataset used in this training can be flexibly combined according to different training types, task types, and the data attribute information of the training sample dataset used in this training, thereby determining the target data format template more flexibly and comprehensively.

[0053] Regarding step S102, in one embodiment provided in this application, the step of analyzing and processing each training sample data in the training sample dataset according to the target data format template to determine the metadata of each training sample data includes: S1021. Identify all metadata dimensions in the target data format template.

[0054] S1022. For each training sample data in the training sample dataset, determine the metadata of the training sample data under each metadata dimension according to all metadata dimensions, and obtain at least one metadata of the training sample data.

[0055] Regarding step S1021, each target data format template has at least one metadata dimension. Furthermore, the metadata dimension may also include sub-dimensions.

[0056] In one embodiment provided in this application, the metadata dimension includes at least one of the following: training type, training stage, task type, data type, data source, verifier, image type, generated content size, difficulty value, number of identified targets, identified target attributes, identification difficulty, number of located targets, located target attributes, location difficulty, data quality, reward function, etc.

[0057] For example, when the metadata dimension is a reward function, the corresponding sub-dimension may also include the reward function threshold, the weight of each reward function, etc.

[0058] To further understand the metadata dimensions, the following will be used as an explanation: Training types can include pre-training, fine-tuning training, and reinforcement learning-based post-training, which are used to determine the basic parameter configuration for model training corresponding to each training sample data. For example, different training types correspond to different activation parameters.

[0059] The training phase is used to determine the model training phase corresponding to each training sample data, such as the training batch to which each training sample data belongs.

[0060] Task types can include visual prediction tasks, such as Math, Science, Charting, and Puzzle, and visual perception tasks, such as Detection, Grounding, Counting, and Optical Character Recognition (OCR), used to determine the specific parameter configuration for model training corresponding to each training sample data, such as the corresponding prompt words.

[0061] Data types can include text, image, audio, and video types, and are used to determine the specific parameter configuration for model training corresponding to each training sample data, such as the corresponding codec.

[0062] The data source can be determined based on which training dataset each training data belongs to, which is used to determine the dataset from which each training sample data is obtained.

[0063] The verifier can select at least one verifier from a pre-defined pool of verifiers, or can select at least one verifier from a pre-defined pool of verifiers based on data values ​​from other dimensions. This is used to calculate the difference between the predicted value and the actual value, thereby determining the feedback value based on the difference and a threshold.

[0064] Image type is used to determine whether each training sample data contains only image content or also text annotation content, and can be obtained directly based on the structure of each training data.

[0065] The generated content size is used to determine whether the generated content is a large image of 100M or a small image of 100K, which can be obtained directly from the storage size of each training sample data.

[0066] The difficulty value can be obtained by directly weighting the presence of annotations and the size of the generated content; alternatively, a pre-trained model can be used to score each training data point based on dimensions such as the completeness of the annotations, their matching with the image content, and the relevance of the generated content, thus determining the difficulty value.

[0067] The number of targets to be identified can be determined by the results of pre-labeling or pre-trained models. If the training data is not for the type of identification task, the metadata in this dimension can be null.

[0068] The target attribute to be identified can be any one or more such as text, numbers, formulas, tables, faces, animals, or objects. It is determined by the recognition results of pre-labeled or pre-trained models. If the training data is not for the recognition task type, the metadata in this dimension can be null.

[0069] The difficulty of identification can be obtained by directly weighting the number of targets to be identified and the difficulty values ​​of the target attributes; alternatively, it can be determined by scoring each training data point using a pre-trained model. If the training data is not for a specific identification task, the metadata for this dimension can be null.

[0070] Number of targets, target attributes, and location difficulty: Similar to recognition tasks, these are used to accurately outline target regions from images. If not training a location task, the values ​​can be empty.

[0071] Data quality can be scored using a pre-trained quality assessment model.

[0072] The reward function can be determined from at least one preset reward function, or it can be determined based on data values ​​from other dimensions. For example, at least one reward function type can be determined based on the type of task performed by the model, and at least one specific reward function can be determined based on factors such as generation difficulty, recognition difficulty, localization difficulty, and data quality (including determining the feedback value, threshold, and weight of the reward function; for example, training data with higher difficulty has a lower threshold, and the reward feedback value is higher when the threshold is met. Weights can also be set between different reward functions for different difficulty values ​​(for example, reward functions include the accuracy of the prediction result and the accuracy of the prediction format; different weights can be set for these two reward functions based on the difficulty value)).

[0073] Regarding step S1022, at least one metadata can be obtained from a training sample data. For example, assuming there are two determined target data format templates, each template includes two metadata dimensions, then four metadata can be determined for each training sample data.

[0074] It's important to note that metadata, based on a predefined target data format template, scans and analyzes every training sample in the entire training dataset to extract or calculate summary information describing the overall and partial features of each training sample. This information is not the content of the training data samples themselves, but rather describes the attributes of these samples.

[0075] Regarding step S1022, in one embodiment provided in this application, determining the metadata of each training sample data in the training sample dataset according to the metadata dimension further includes: determining at least one second metadata in a second metadata dimension based on the first metadata of each training sample data in the first metadata dimension; and determining the metadata of the training sample data in each metadata dimension based on the first metadata and the second metadata.

[0076] Specifically, this embodiment may include: for each training sample data in the training sample dataset, determining at least one first metadata for the training sample data under each first metadata dimension; determining second metadata under each second metadata dimension corresponding to the first metadata based on the first metadata; and then determining the metadata of the training sample data under each metadata dimension based on all the determined first metadata and all the second metadata.

[0077] For example, the first metadata dimension includes the number of identified targets and the attributes of the identified targets, while the second metadata dimension includes the identification difficulty. For each training sample in the training sample dataset, two first metadata elements are determined for that training sample data under the number of identified targets and the identification attributes. Based on the number of identified targets and the difficulty value of the identification attributes, the identification difficulty value under the identification difficulty dimension corresponding to the number of identified targets and the difficulty value of the identification attributes is determined. Then, based on the determined number of identified targets, identification attributes, and identification difficulty value, the metadata for that training sample data under each metadata dimension is determined.

[0078] Regarding step S103, for the training process of obtaining the target multimodal model, this application provides multiple training methods, as detailed below: In the first embodiment provided in this application, the step of performing multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset includes: S10311. Obtain metadata corresponding to at least two training sample data of at least two task types.

[0079] S10312. Input the metadata of at least two training sample data of the at least two task types and the corresponding training sample data as the same training batch into the initial multimodal model for unified training.

[0080] Specifically, step S10311 may include: determining all task types of the initial multimodal model; determining at least two task types to be trained in each batch of training; and then, for each task type, obtaining metadata corresponding to at least one training sample data of that task type, thus obtaining metadata corresponding to at least two training sample data of at least two task types.

[0081] For step S10312, at least two metadata entries and the training sample data corresponding to each metadata entry are used as training data for the same training batch. Then, they are input into the initial multimodal model for unified training. The model is verified based on the correct results of the input data and the prediction results of the initial multimodal model to obtain a verification feedback value. The model parameters of the initial multimodal model are updated based on the feedback value corresponding to each training sample data.

[0082] In the second embodiment provided in this application, the metadata includes data source, task type, and validator. The step of performing multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model includes: S10321. Obtain metadata corresponding to at least two training sample data; S10322. Obtain the corresponding training sample data based on the data source in the metadata; S10323. The initial multimodal model predicts and outputs based on the task type in the metadata using the acquired training sample data. S10324. Validate the predicted output based on the correct answer in the training sample data and the validator in the metadata, and obtain the feedback value corresponding to the training sample data. S10325. Adjust the initial multimodal model parameters based on the feedback values ​​to obtain the target multimodal model.

[0083] In step S10323, the acquired training sample data and metadata are input into the initial multimodal model. The initial multimodal model performs the corresponding prediction task according to the task type specified in the metadata to obtain the prediction output of the training sample data.

[0084] For step S10324, this step may include: comparing the model’s predicted output with a pre-prepared correct answer, and using a validator specified in the metadata to generate a quantitative evaluation result, i.e., a feedback value.

[0085] The validator can be a predefined set of functions, models, rules, or standards used to judge whether the model output is correct, what the accuracy rate is, whether the format is correct, whether the corresponding threshold is reached, what the corresponding feedback value is, and what the weight of different reward functions is, etc.

[0086] The feedback value is a quantitative indicator used to reflect the degree of matching between the model's predicted output and the expected output (correct answer) on a specific training sample.

[0087] For step S10325, this step may include: taking the feedback value obtained in the previous step as a signal and adjusting the parameters inside the initial multimodal model through machine learning algorithms such as backpropagation.

[0088] In the third embodiment provided in this application, the initial multimodal model is subjected to multi-task unified training based on each training sample data in the training sample dataset and the corresponding metadata to obtain the target multimodal model, and the method further includes: S10331. Determine whether the multi-task performance of the initial multimodal model corresponding to the current training sample data meets the corresponding performance requirements; S10332. If all requirements are met, the target multimodal model is obtained; S10333. If none of the requirements are met, continue to perform multi-task unified training on the initial multimodal model based on each training sample data in the training sample dataset and the corresponding metadata to obtain the target multimodal model. S10334. If the performance of some tasks meets the requirements, the initial multimodal model will be trained in a unified manner for multiple tasks in the next training batch based on the metadata of other task types and the corresponding training sample data to obtain the target multimodal model.

[0089] Regarding step S10331, this step may include: inputting each training sample data and each corresponding metadata of the training sample data into the initial multimodal model, for the currently input training sample data, identifying the task processing type that the model can handle; and determining the task performance of the initial multimodal model in handling that task based on the determined task processing type. For example, this can be done by determining whether the model gradient has converged, or whether the statistical result of the feedback value is greater than a preset value, or by testing the model performance with test data corresponding to the task processing type to obtain a test score; and determining whether the task performance meets the requirements. If the condition is met, proceed to step S10332; otherwise, proceed to step S10333 or S10334.

[0090] For step S10322, this step includes identifying whether the performance of each of the remaining tasks in the initial multimodal model meets the requirements.

[0091] If the performance of all other tasks fails to meet the requirements, proceed to step S10333; otherwise, proceed to step S10334.

[0092] For the method of judging whether the performance of the corresponding task meets the requirements in step S10333, please refer to step S10331, which will not be elaborated here.

[0093] For step S10334, the description of this step can be found in steps S10311-S10312.

[0094] It's important to note that current model training typically involves inputting training data into an initial multimodal model in batches. This initial model then performs prediction operations on the task portion of that batch of training data and outputs the prediction results. These predictions are then compared to the corresponding correct answers in the training data, and at least one of the loss, reward, or penalty values ​​is calculated as feedback. The model parameters are then updated based on this feedback. A batch of training data can include 100, 200, or 1000 data points, etc. Typically, a batch of training data includes only one training dataset and performs only one model training task. This training method, on the one hand, isolates the mutual influence of training data from different modalities and tasks on model performance improvement. On the other hand, the quality, difficulty, content, and structure of the training data in the datasets used to train each task vary considerably. Directly using the training data and model parameter configurations in the training datasets not only results in low training efficiency but also poor model stability.

[0095] This approach allows for the use of training samples from multiple training datasets (which can be mixed and sorted) to perform various model training tasks. By acquiring the metadata of each training sample from a batch of data, the model can be trained using each individual sample. This achieves unified multi-task training, improving model training efficiency, stability, and overall performance.

[0096] In the fourth embodiment provided in this application, the step of performing multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model includes: S10341. For each training sample data in any batch, use the initial multimodal model to make a prediction based on the metadata of the training sample data to obtain the prediction result of the training sample data.

[0097] S10342. Generate a data verification request for each training sample based on the metadata, correct answer, and prediction result of each training sample.

[0098] S10343. Send the data verification request for each training sample data to the server, so that the server determines and calls the target verifier corresponding to each training sample data according to the data verification request for each training sample data, and obtains the data verification result for each training sample data. S10344 Receive the data verification results of each training sample data returned by the server, and update the model parameters of the initial multimodal model based on the data verification results of each training sample data to obtain the target multimodal model.

[0099] Regarding step S10341, in this step, for each training sample data in the multiple training sample data included in any batch, the initial multimodal model deployed in the client can make a prediction based on the task type in the metadata of the training sample data to obtain the prediction result of the training sample data.

[0100] Regarding step S10342, in this step, the metadata of each training sample data also includes the validator identifier corresponding to the training sample data, and each training sample data includes the corresponding correct answer.

[0101] In this way, the server that receives the data verification request can parse the data verification request to determine the metadata, correct answer and prediction result corresponding to the training data, and determine the target validator for verification and response based on the validator identifier in the metadata.

[0102] Regarding step S10343, in this step, multiple validators are pre-deployed on the server. The validators are used to calculate the data verification results of the corresponding training sample data for a specific task based on the given data verification request.

[0103] Regarding step S10344, in this step, the client receives the data verification results of each training sample data in this batch returned by the server, and updates the model parameters of the initial multimodal model based on the data verification results of each training sample data. When the model training termination condition is met, the target multimodal model is obtained.

[0104] In the fifth embodiment provided in this application, the initial multimodal model is subjected to multi-task unified training based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model, including: S10351. Obtain the prediction output of the initial multimodal model corresponding to each training sample data. The training sample data includes images or videos.

[0105] S10352. Based on the predicted output and the index value of at least one preset index, determine the prediction result of the initial multimodal model corresponding to each training data.

[0106] S10353. Based on the correct results and prediction results corresponding to the training sample data, determine the feedback value corresponding to the training data and generate a training log corresponding to each training data; wherein, the training log is used to update the indicator value of at least one preset indicator used in the next round of training.

[0107] S10354. Based on the feedback value, adjust the model parameters of the initial multimodal model to obtain the target multimodal model.

[0108] For step S10352, this step may include: determining the value of the real-time monitoring indicator based on the value of the preset aggregation indicator, for example, determining the output length threshold based on the average output length correctly predicted in the previous training sample data, judging whether the output length of the current training sample data has reached the threshold based on the output length threshold, and if it has, truncating it, and taking the truncated prediction output as the prediction result.

[0109] For step S10353, this step may include: the training log records information such as whether the prediction result of each training task and each training sample data is correct, the length of the prediction result, whether it is truncated, whether a reflection word appears, and the number of times the reflection word appears. The index value of the preset index is obtained by statistically analyzing the data in the log, and then at least one parameter threshold and / or reflection word is dynamically updated by using the index value updated in real time.

[0110] In the sixth embodiment provided in this application, the step of performing multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model includes: S10361. Input the training sample dataset and each training sample data in the metadata set and each metadata corresponding to the training sample data into the initial model in sequence, and calculate the optimization index for each input.

[0111] S10326. After calculating the optimization index for each input, the model parameter update strategy is executed on the initial multimodal model according to the optimization index until the expected performance is achieved, and the training ends to obtain the target multimodal model.

[0112] For step S10361, the optimization metric is at least one of loss value, reward value, and penalty value.

[0113] In this step, the optimization index can be determined based on the deviation between the input and output data.

[0114] Regarding step S10362, the corresponding model parameter update strategy can be different depending on the value of the optimization index.

[0115] Furthermore, after each input of training sample data, the model parameters can be updated until the expected performance is achieved, at which point the training ends and the target multimodal model is obtained.

[0116] Furthermore, when the target multimodal model handles tasks including question-answering and recognition tasks, the training method after obtaining the target multimodal model further includes: S201. Obtain the data to be identified.

[0117] S202. Input the data to be identified into the target multimodal model and determine the task type corresponding to the data to be identified.

[0118] S203. Based on the determined task type, the target multimodal model analyzes the data to be identified and outputs the processing result of the data to be identified.

[0119] For step S201, the data to be identified is any one of the following: image data, text data, and voice data.

[0120] Regarding step S202, in this step, the task type corresponding to the data to be identified is either a question-and-answer task or a recognition task.

[0121] Regarding step S203, the processing result in this step is either an answer result or an identification result.

[0122] In this way, this application uniformly inputs the training sample data and its corresponding metadata from the training sample dataset into the initial multimodal model for training, generating the target multimodal model, without having to configure model parameters separately for each target training task, based on the training data in each training dataset, and train the model separately. This can improve model training efficiency, improve the refinement of model training, improve model training stability and model training effect, and also reduce model training cost.

[0123] Based on the same inventive concept, this application also provides a training device corresponding to the training method. Since the principle of the device in this application is similar to the training method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0124] Please see Figure 2 , Figure 3 , Figure 2 This is one of the structural schematic diagrams of a multimodal model multi-task unified training device provided in an embodiment of this application. Figure 3 This is a second schematic diagram of a multimodal model multi-task unified training device provided in an embodiment of this application. Figure 2 As shown, the training device 200 includes: The determining module 210 is used to determine a target data format template for structuring the training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training; wherein, the task type includes at least two task types, and the training sample data includes images or videos. Processing module 220 is used to analyze and process each training sample data in the training sample dataset according to the target data format template, and determine the metadata of each training sample data. The training module 230 is used to perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model.

[0125] Optionally, when determining the target data format template for structuring the training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training, the determining module 210 is used to: Analyze each item in the data attribute information of the training type, task type, and training sample dataset used in this training to determine the corresponding initial template requirements; Integrate all initial template requirements to determine the target template requirements; Based on the target template requirements, determine the target data format template for structuring the training sample dataset.

[0126] Optionally, when the processing module 220 analyzes and processes each training sample data in the training sample dataset according to the target data format template to determine the metadata of each training sample data, the processing module 220 is used to: Identify all metadata dimensions in the target data format template; For each training sample in the training sample dataset, determine the metadata of the training sample under each metadata dimension based on all metadata dimensions, and obtain at least one metadata of the training sample.

[0127] Optionally, when the training module 230 performs multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset, the training module 230 is used to: Obtain metadata corresponding to at least two training sample data points of at least two task types; The metadata of at least two training sample data of the at least two task types and the corresponding training sample data are input into the initial multimodal model as the same training batch for unified training.

[0128] Optionally, the metadata includes data source, task type, and validator. When the training module 230 performs multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model, the training module 230 is used to: Obtain metadata corresponding to at least two training sample data points; Obtain the corresponding training sample data based on the data source in the metadata; The initial multimodal model predicts outputs based on the task type in the metadata from the acquired training sample data; The predicted output is validated based on the correct answer in the training sample data and the validator in the metadata, and the feedback value corresponding to the training sample data is obtained. The initial multimodal model parameters are adjusted based on the feedback values ​​to obtain the target multimodal model.

[0129] Optionally, when the training module 230 is used to perform multi-task unified training on the initial multimodal model based on each training sample data in the training sample dataset and the corresponding metadata to obtain the target multimodal model, the training module 230 is used to: Determine whether the multi-task performance of the initial multimodal model corresponding to the current training sample data meets the corresponding performance requirements; If all requirements are met, the target multimodal model is obtained; If none of the requirements are met, then the initial multimodal model will continue to be trained using each training sample data in the training sample dataset and the corresponding metadata to obtain the target multimodal model. If the performance of some tasks meets the requirements, the initial multimodal model will be trained in a unified manner on multiple tasks in the next training batch based on the metadata of other task types and the corresponding training sample data to obtain the target multimodal model.

[0130] Optionally, when the processing module 220 determines the metadata of each training sample data in each metadata dimension based on the metadata dimension in the training sample dataset, the processing module 220 is configured to: Based on the first metadata in the first metadata dimension of each training sample data, determine at least one second metadata in the second metadata dimension; Based on the first metadata and the second metadata, determine the metadata of the training sample data under each metadata dimension.

[0131] Optionally, when the training module 230 is used to perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model, the training module 230 is used to: For each training sample data in any batch, the initial multimodal model is used to make a prediction based on the metadata of the training sample data to obtain the prediction result of the training sample data; Based on the metadata, correct answer, and prediction result of each training sample, generate a data verification request for each training sample; Each training sample data data data data verification request is sent to the server, so that the server determines and calls the target verifier corresponding to each training sample data data data according to the data verification request of each training sample data data data to respond, and obtains the data verification result of each training sample data data data; wherein, multiple verifiers are pre-deployed on the server; The system receives the data validation results of each training sample data returned by the server, and updates the model parameters of the initial multimodal model based on the data validation results of each training sample data to obtain the target multimodal model.

[0132] Optionally, when the training module 230 is used to perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model, the training module 230 is used to: Obtain the prediction output of the initial multimodal model corresponding to each training sample data; wherein, the training sample data includes images or videos; Based on the predicted output and the index value of at least one preset index, the prediction result of the initial multimodal model corresponding to each training data is determined. Based on the correct results and prediction results corresponding to the training sample data, the feedback value corresponding to the training data is determined, and a training log corresponding to each training data is generated; wherein, the training log is used to update the indicator values ​​of at least one preset indicator used in the next round of training; The model parameters of the initial multimodal model are adjusted based on the feedback value to obtain the target multimodal model.

[0133] Optionally, when determining the target data format template for structuring the training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training, the determining module 210 is used to: Based on the training type of this training, determine the first data format template corresponding to the training type; Based on the task type of this training, determine the second data format template corresponding to the task type; Based on the task type performed by the target model, determine the third data format template corresponding to that task type; Based on the data attribute information of the training sample dataset used in this training, determine the third data format template corresponding to the data attribute information of the training sample dataset; At least one of the first data format template, the second data format template, and the third data format template is determined as the target data format template.

[0134] Optionally, when determining the target data format template for structuring the training sample data in the training sample dataset based on at least one of the training type, task type, and data attribute information of the training sample dataset used in this training, the determining module 210 is used to: Determine the training type, task type, and priority of each reference data item in the data attribute information of the training sample dataset used in this training; Obtain the candidate data format template corresponding to the highest priority reference data, and determine whether the candidate data format template meets the target template requirements; If the target template requirements are met, the candidate data format template corresponding to the highest priority reference data will be determined as the target data format template. If the target template requirements are not met, obtain the candidate data format templates corresponding to the reference data of the next priority, and determine whether they meet the target template requirements until the requirements are met. Then, determine the multiple candidate data format templates obtained as the target data format template.

[0135] Optionally, the metadata dimensions include at least one of the following: image type, generated content size, difficulty value, number of identified targets, identified target attributes, identification difficulty value, number of located targets, located target attributes, location difficulty, data source, data quality, reward function, and validator.

[0136] Optionally, when the training module 230 is used to perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model, the training module is used to: Each training sample data and each corresponding metadata in the training sample dataset and metadata set are sequentially input into the initial model, and the optimization metric for each input is calculated; wherein the optimization metric is at least one of loss value, reward value, and penalty value; After calculating the optimization index for each input, the initial multimodal model is subjected to a model parameter update strategy based on the optimization index until the expected performance is achieved, at which point the training ends and the target multimodal model is obtained.

[0137] Optional, such as Figure 3 As shown, the training device 200 uses module 240, which is used for: When the target multimodal model processes tasks including question-answering tasks and recognition tasks, it acquires data to be recognized; the data to be recognized is any one of the following: image data, text data, and voice data; The data to be identified is input into the target multimodal model, and the task type corresponding to the data to be identified is determined. Based on the determined task type, the target multimodal model analyzes the data to be identified and outputs the processing result of the data to be identified; wherein the processing result is the answer result or the identification result.

[0138] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.

[0139] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, they can perform the operations described above. Figure 1 The steps in the method embodiment shown are specifically implemented in the method embodiment and will not be repeated here.

[0140] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 The steps in the method embodiment shown are specifically implemented in the method embodiment and will not be repeated here.

[0141] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0144] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0145] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0146] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A unified training method for multi-modal models across multiple tasks, characterized in that, The method includes: Based on at least one of the following: training type, task type, and data attribute information of the training sample dataset used in this training, determine the target data format template for structuring the training sample data in the training sample dataset; wherein, the task type includes at least two task types, and the training sample data includes images or videos; The metadata of each training sample in the training sample dataset is analyzed and processed according to the target data format template to determine the metadata of each training sample. The initial multimodal model is trained using the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model. The metadata includes data source, task type, and validator. The training process for obtaining the target multimodal model includes: acquiring metadata corresponding to at least two training sample data sets; acquiring the corresponding training sample data according to the data source in the metadata; the initial multimodal model predicting output based on the acquired training sample data according to the task type in the metadata; validating the prediction output based on the prediction output, the correct answer in the training sample data, and the validator in the metadata, and obtaining the feedback value corresponding to the training sample data; adjusting the parameters of the initial multimodal model based on the feedback value to obtain the target multimodal model; or, The training process for obtaining the target multimodal model includes: for each training sample data in any batch, using the initial multimodal model to make predictions based on the metadata of the training sample data, and obtaining the prediction result of the training sample data; generating a data verification request for each training sample data based on the metadata, correct answer, and prediction result of each training sample data; sending the data verification request for each training sample data to the server, so that the server determines and calls the target validator corresponding to each training sample data to respond based on the data verification request of each training sample data, and obtaining the data verification result of each training sample data; wherein, multiple validators are pre-deployed on the server; receiving the data verification result of each training sample data returned by the server, and updating the model parameters of the initial multimodal model based on the data verification result of each training sample data, to obtain the target multimodal model.

2. The training method according to claim 1, characterized in that, The step of determining a target data format template for structuring the training sample data in the training sample dataset based on at least one of the following: training type, task type, and data attribute information of the training sample dataset used in this training; including: Analyze each item in the data attribute information of the training type, task type, and training sample dataset used in this training to determine the corresponding initial template requirements; Integrate all initial template requirements to determine the target template requirements; Based on the target template requirements, determine the target data format template for structuring the training sample dataset.

3. The training method according to claim 1, characterized in that, The step of analyzing and processing each training sample data in the training sample dataset according to the target data format template to determine the metadata of each training sample data includes: Identify all metadata dimensions in the target data format template; For each training sample in the training sample dataset, determine the metadata of the training sample under each metadata dimension based on all metadata dimensions, and obtain at least one metadata of the training sample.

4. The training method according to claim 1, characterized in that, The step of performing multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset includes: Obtain metadata corresponding to at least two training sample data points of at least two task types; The metadata of at least two training sample data of the at least two task types and the corresponding training sample data are input into the initial multimodal model as the same training batch for unified training.

5. The training method according to claim 1, characterized in that, The initial multimodal model is trained using a multi-task unified method based on each training sample data in the training sample dataset and the corresponding metadata to obtain the target multimodal model, and the method further includes: Determine whether the multi-task performance of the initial multimodal model corresponding to the current training sample data meets the corresponding performance requirements; If all requirements are met, the target multimodal model is obtained; If none of the requirements are met, then the initial multimodal model will continue to be trained using each training sample data in the training sample dataset and the corresponding metadata to obtain the target multimodal model. If the performance of some tasks meets the requirements, the initial multimodal model will be trained in a unified manner on multiple tasks in the next training batch based on the metadata of other task types and the corresponding training sample data to obtain the target multimodal model.

6. The training method according to claim 3, characterized in that, For each training sample in the training sample dataset, the metadata for that training sample in each metadata dimension is determined according to the aforementioned metadata dimension, further including: Based on the first metadata in the first metadata dimension of each training sample data, determine at least one second metadata in the second metadata dimension; Based on the first metadata and the second metadata, determine the metadata of the training sample data under each metadata dimension.

7. The training method according to any one of claims 1 to 6, characterized in that, The step of performing multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model includes: Obtain the prediction output of the initial multimodal model corresponding to each training sample data; wherein, the training sample data includes images or videos; Based on the predicted output and the index value of at least one preset index, the prediction result of the initial multimodal model corresponding to each training data is determined. Based on the correct results and prediction results corresponding to the training sample data, the feedback value corresponding to the training data is determined, and a training log corresponding to each training data is generated; wherein, the training log is used to update the indicator values ​​of at least one preset indicator used in the next round of training; The model parameters of the initial multimodal model are adjusted based on the feedback value to obtain the target multimodal model.

8. A unified training device for multimodal models and multitasking, characterized in that, The training device includes: The determination module is used to determine the target data format template for structuring the training sample data in the training sample dataset based on at least one of the following: training type, task type, and data attribute information of the training sample dataset used in this training; wherein, the task type includes at least two task types, and the training sample data includes images or videos; The processing module is used to analyze and process each training sample data in the training sample dataset according to the target data format template, and determine the metadata of each training sample data. The training module is used to perform multi-task unified training on the initial multimodal model based on the metadata and each training sample data in the corresponding training sample dataset to obtain the target multimodal model. The metadata includes data source, task type, and validator. The training module is specifically used for: acquiring metadata corresponding to at least two training sample data sets; acquiring the corresponding training sample data based on the data source in the metadata; the initial multimodal model predicting output based on the acquired training sample data according to the task type in the metadata; validating the prediction output based on the prediction output, the correct answer in the training sample data, and the validator in the metadata, and acquiring the feedback value corresponding to the training sample data; adjusting the parameters of the initial multimodal model based on the feedback value to obtain the target multimodal model; or... The training module is specifically used for: for each training sample data in any batch, using an initial multimodal model to make predictions based on the metadata of the training sample data, to obtain the prediction result of the training sample data; generating a data verification request for each training sample data based on the metadata, correct answer, and prediction result of each training sample data; sending the data verification request for each training sample data to the server, so that the server determines and calls the target validator corresponding to each training sample data to respond based on the data verification request of each training sample data, to obtain the data verification result of each training sample data; wherein, multiple validators are pre-deployed on the server; receiving the data verification result of each training sample data returned by the server, and updating the model parameters of the initial multimodal model based on the data verification result of each training sample data, to obtain the target multimodal model.

Citation Information

Patent Citations

  • Insurance data analysis method and system based on multi-modal pre-training large model

    CN119513558A

  • Multi-modal information tagging method, apparatus and device, and storage medium and product

    WO2025148651A1