Training method and device of multi-modal information processing model, and sample screening method and device

By annotating large multimodal models with generation capability labels and conducting supervised training, the problem of inconsistent quality of multimodal data samples is solved, efficient screening and training are achieved, and the performance and training efficiency of the model are improved.

CN119441873BActive Publication Date: 2025-10-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411480622.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-10-10
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively screen high-quality multimodal data samples, resulting in inefficient model training and a lack of strategies for mining the relationship between multimodal data and the target model.

Method used

Multimodal data is labeled through a large multimodal model to generate capability labels. Supervised training is performed based on the capability labels to screen out useful samples and build a multimodal information processing model.

Benefits of technology

It improves the efficiency and accuracy of model training, screens out high-quality training samples, enhances the generalization ability of the model, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441873B_ABST
    Figure CN119441873B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for training a multi-modal information processing model, a sample screening method and device, computer technology field, especially in the technical fields of artificial intelligence, large model, data screening, information processing and the like. The specific implementation scheme is: based on the multi-modal large model, the multi-modal data is labeled to obtain the ability label of the multi-modal data; the ability label is used to describe the model ability that can be trained by the multi-modal data; based on the multi-modal data and the ability label, the supervised training is carried out on the to-be-trained model, so as to train the to-be-trained model into a multi-modal information processing model capable of screening useful samples for the target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to the fields of artificial intelligence, large models, data screening, information processing, and the like. Background Art

[0002] With the development of the Internet, the amount of data generated by social media, IoT devices, etc. has increased dramatically. This has made it possible to train more complex and large models.

[0003] Massive data samples help the model fully learn the rules and knowledge contained therein, so as to facilitate accurate regression analysis and classification. Summary of the Invention

[0004] The present disclosure provides a training method, a sample screening method and a device for a multimodal information processing model.

[0005] According to one aspect of the present disclosure, a method for training a multimodal information processing model is provided, comprising:

[0006] Multimodal data is labeled based on a large multimodal model to obtain capability labels for the multimodal data. Capability labels are used to describe the model capabilities that can be trained using multimodal data.

[0007] Based on multimodal data and capability labels, the model to be trained is supervised and trained to be a multimodal information processing model that can filter out useful samples for the target model.

[0008] According to one aspect of the present disclosure, a sample screening method is provided, comprising:

[0009] Get the target dataset;

[0010] Inputting the target data set into the multimodal information processing model to obtain capability labels of data samples in the target data set output by the multimodal information processing model;

[0011] Based on the capability labels of the data samples in the target dataset, training samples for training the target model are screened from the target dataset.

[0012] According to another aspect of the present disclosure, a training device for a multimodal information processing model is provided, comprising:

[0013] The annotation module is used to annotate multimodal data based on the multimodal large model to obtain the capability labels of the multimodal data; the capability labels are used to describe the model capabilities that can be trained with multimodal data;

[0014] The training module is used to perform supervised training on the model to be trained based on multimodal data and capability labels, so as to train the model to be trained into a multimodal information processing model that can filter out useful samples for the target model.

[0015] According to another aspect of the present disclosure, there is provided a sample screening device, comprising:

[0016] Acquisition module, used to obtain the target data set;

[0017] An output module, configured to input a target data set into a multimodal information processing model and obtain capability labels of data samples in the target data set output by the multimodal information processing model;

[0018] The screening module is used to screen out training samples for training the target model from the target dataset based on the capability labels of the data samples in the target dataset.

[0019] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0020] at least one processor; and

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0023] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0024] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0025] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0027] Figure 1 is a flowchart of a method for training a multimodal information processing model according to an embodiment of the present disclosure;

[0028] Figure 2 This is a schematic diagram of a process for obtaining capability labels for multimodal data according to an embodiment of the present disclosure;

[0029] Figure 3 is a flowchart of supervised training provided according to an embodiment of the present disclosure;

[0030] Figure 4 is a schematic diagram of a process for training a model to be trained according to an embodiment of the present disclosure;

[0031] Figure 5 is a flow chart of a sample screening method provided according to an embodiment of the present disclosure;

[0032] Figure 6 This is a schematic diagram of an overall process provided according to an embodiment of the present disclosure;

[0033] Figure 7 is a structural diagram of a training device for a multimodal information processing model provided according to an embodiment of the present disclosure;

[0034] Figure 8 is a structural schematic diagram of a sample screening device provided according to an embodiment of the present disclosure;

[0035] Figure 9 It is a block diagram of an electronic device used to implement the training method and / or sample screening method of the multimodal information processing model of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0037] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0038] In the era of big models, the importance of data has been further emphasized and enhanced. Training big models requires vast amounts of data as input to generate accurate and reliable predictions. Data is a key factor in model training, determining its performance and accuracy. Without sufficient data, models cannot effectively learn and predict.

[0039] In addition to data quantity, data quality also plays a crucial role in model training. High-quality data can provide more accurate training signals, thereby improving model performance. Therefore, the process of selecting high-quality data requires extensive data analysis and processing.

[0040] However, with the development of technology and the accumulation of time, the time span for the formation of data samples is long, and the quality standards of data samples are constantly improving and changing. At present, the automatic screening of data samples is limited to screening plain text data. As for multimodal data, on the one hand, a large number of data samples have been generated through years of accumulation, and on the other hand, the quality of multimodal data is constantly being optimized, resulting in the current quality of multimodal data samples being inconsistent. For example, although the current open source data is huge in volume, not all data needs to be involved in training due to inconsistent data quality. Or, when some of the data is needed to train a model, there is a lack of strategies to automatically screen high-quality samples by mining the connection between data samples and the target model to be trained.

[0041] In view of this, the embodiments of the present disclosure provide a solution. In this solution, we first rely on the reasoning and decision-making capabilities of the multimodal large model to help train a multimodal information processing model that can be labeled. The multimodal information processing model can mine the relationship between multimodal data samples and the target model that needs to be trained. In the embodiment of the present disclosure, this relationship is represented by the capability label of the multimodal data sample, so as to measure which model capabilities can be trained by different data samples. Thus, high-quality training samples can be mined based on the capability label, and low-quality redundant data samples can be eliminated. In the face of massive data, it can provide high-quality data for the target model while filtering out redundant low-quality data to improve the training efficiency of the target model.

[0042] In the first aspect of the embodiments of the present disclosure, a method for training a multimodal information processing model is provided. In the training method, a multimodal information processing model capable of labeling (also referred to as a label expert model) can be trained with the help of a multimodal large model. Specifically, Figure 1 As shown, the training method of the multimodal information processing model provided by the embodiment of the present disclosure includes the following contents:

[0043] S101: Label the multimodal data based on the multimodal large model to obtain capability labels of the multimodal data; the capability labels are used to describe the model capabilities that can be trained using the multimodal data.

[0044] Multimodal data generally includes information in at least two of the following modalities: text modality, image modality, audio modality, etc. Among them, video can be understood as data in the image modality.

[0045] A multimodal large model is one that combines and processes multiple modal information, including text, images, audio, and video. This model can handle multiple types of data and, by fusing data from different modalities, enables more efficient information response and intelligent recognition.

[0046] It is understandable that the multimodal data in S101 is described by taking one data sample as an example. When there are multiple multimodal data, each multimodal data can be executed with reference to this, and the embodiments of the present disclosure are not limited to this.

[0047] Labeling these multimodal data based on the multimodal big model is to conduct in-depth analysis of the input multimodal data through the powerful understanding and reasoning capabilities of the multimodal big model, and generate corresponding capability labels for each multimodal data based on the analysis results.

[0048] As the name suggests, capability labels indicate which model capabilities can be trained using multimodal data. Examples include speech recognition, text analysis, emotion recognition, image description, image classification, object location, character recognition within images, semantic understanding, and even complex capabilities such as finding relevant answers to questions from images. Specific capabilities can be defined based on business needs, and model capabilities can even be categorized into different levels based on needs. During implementation, model capabilities can also be categorized from different perspectives and granularity, potentially encompassing thousands of categories.

[0049] On the basis of having capability labels, when a multimodal data sample A can only train a very small amount of model capabilities, and these model capabilities can be covered by multiple other data samples, the data sample A is an unimportant, dispensable, redundant data sample. In the case of screening out a small number of partial samples from a massive amount of data samples, the data sample A will be eliminated. Therefore, through the capability label, the relationship between the data sample and the target model can be explicitly expressed, so as to facilitate the mining of reliable and useful samples to train the target model. In summary, the accuracy and comprehensiveness of the capability label are crucial to the subsequent training of the target model. Therefore, the embodiment of the present disclosure trains an accurate and reliable multimodal information processing model that can predict capability labels through subsequent steps. The model will be trained with the help of effective samples mined from the multimodal large model, and please refer to S102 for details.

[0050] S102: Based on the multimodal data and the capability labels, supervised training is performed on the model to be trained, so as to train the model to be trained into a multimodal information processing model that can filter out useful samples for the target model.

[0051] The model to be trained is a model in its initial state. This model can be a lightweight model capable of processing multimodal information, or a large model with a complex structure and large model parameters. This initial model can be trained in a supervised manner based on labeled multimodal data and corresponding capability labels. The trained model should be able to accurately extract capability labels from multimodal data, allowing it to filter out useful samples from multiple data samples based on capability labels that are conducive to training the target model.

[0052] In the disclosed embodiments, a large multimodal model is used to construct training samples for the model to be trained, allowing the subsequent training of a multimodal information processing model through supervised training. This multimodal information processing model can accurately predict the capability labels of the multimodal data, thereby intuitively and conveniently reflecting the relationship between the multimodal data and the target model. As a result, the capability labels predicted by the multimodal information processing model can be further utilized to screen high-quality samples for the target model.

[0053] In the embodiment of the present disclosure, the method for labeling multimodal data based on the multimodal large model to obtain the capability label of the multimodal data is as follows: Figure 2 Shown, including:

[0054] S201: Construct input information of a multimodal large model based on multimodal data and a preset prompt word template.

[0055] That is, through the preset prompt word template, we help the multimodal large model to deeply understand and sort out the characteristics and task requirements of multimodal data from the perspective of mining model capabilities, so as to accurately mine the capability labels of multimodal data.

[0056] During implementation, appropriate prompt word templates can be designed based on the characteristics of multimodal data and task requirements. These templates should be able to guide the multimodal large model to focus on the key information in the data and guide the multimodal large model to generate the desired output. During implementation, the preset prompt word templates may include at least one of the following requirements:

[0057] (1) Multimodal large models are required to pay attention to the detailed content of image modality information in multimodal data;

[0058] Specifically, while ensuring that the prompt word template can guide the multimodal large model to process multimodal data including images, the multimodal large model is required to pay special attention to the detailed content of the image. This helps the multimodal large model more accurately understand the model capabilities that can be trained by image modality information, thereby generating accurate and rich capability labels.

[0059] For example, in medical image analysis, the prompt word template can explicitly require the model to pay attention to detailed information such as the size, shape, and location of the lesion, with the goal of improving diagnostic accuracy, so that the multimodal large model can accurately output capability labels.

[0060] (2) When multimodal data includes question-answer pairs, the multimodal large model is required to combine the question-answer pairs to understand the query intent of the question in the question-answer pair for the image modality information in the multimodal data;

[0061] That is, when processing multimodal data containing question-answer pairs and images, the multimodal large model can combine the questions in the question-answer pairs to accurately understand the query intent of the question for the image modal information, so as to mine relevant content from the image modal information and predict the capability label of the multimodal data based on the mining results.

[0062] (3) The multimodal large model is required to analyze and reason about the input information to provide reasonable capability labels for the multimodal data.

[0063] This requires that large multimodal models be able to perform comprehensive analysis and reasoning based on the input multimodal data and assign appropriate capability labels. This helps accurately label the target model's capabilities based on the analysis and reasoning process and results, allowing the target model to accurately select the required training samples based on the target model's task requirements when handling complex tasks.

[0064] For example, the prompt word template can require the large model to first understand the question in a chain of thought, and then search for the answer in the image modal information step by step based on the key points in the question, thereby providing an accurate answer. In this way, the multimodal large model can accurately and meticulously analyze the model capabilities that can be trained with multimodal data based on the analysis and reasoning process. It can also be understood that the ability labels of multimodal data can be sorted out based on the analysis and reasoning methods and analysis and reasoning steps of the multimodal large model. Through reasoning and analysis, the model capabilities can be more accurately grasped and classified. For example, if the input image and the question are "Please count the number of red apples in the image", the multimodal large model should be guided to give reasonable ability labels such as "object counting ability", "object recognition ability", and "color recognition ability".

[0065] In the disclosed embodiment, attention is paid to the detailed content of the image modal information in the multimodal data, and the information content conveyed by the image modal information can be understood in detail and accurately, so as to better construct a complete situational picture and enhance the grasp of the overall situation of the image modal information and the cognition of the ability to detail. In the case where the multimodal data includes question-answer pairs, the question in the question-answer pairs is combined with the query intent of the image modal information in the multimodal data to understand the question, and different labels of intent recognition capabilities can be mined. The multimodal large model is required to provide reasonable and accurate capability labels based on the analysis and reasoning of the input information. The capability labels obtained through analysis and reasoning can accurately and meticulously express the model capabilities that can be trained by the multimodal data, so as to accurately screen data samples for the target model. In summary, the preset prompt word template plays a key role in multimodal data processing, which can guide the model to pay attention to specific information, understand the query intent, and provide accurate capability labels, thereby improving the ability of the model to be trained in mining the relationship between data samples and the target model.

[0066] S202: Input the input information into the multimodal big model to obtain the capability label of the multimodal big model for the multimodal data.

[0067] The multimodal model analyzes and understands the input information, obtaining capability labels that it labels for the multimodal data. These capability labels represent the model capabilities that can be trained using the multimodal data.

[0068] During implementation, an accurate and reliable multimodal large model can be selected to process the input information, such as the GPT4-o (a latest artificial intelligence model) large model, to ensure the accuracy of the ability labels obtained for multimodal data annotation.

[0069] In the disclosed embodiment, by constructing input information based on multimodal data and a preset prompt word template, the preset prompt word template can guide the multimodal large model to focus more on how to mine capability labels. Even when the types of model capabilities are subdivided into thousands or tens of thousands of types, the preset prompt word template can guide the multimodal large model to fully and deeply mine the capability labels of the multimodal data, which helps to accurately capture the important elements in the multimodal data, thereby establishing the relationship between the multimodal data and the target model.

[0070] In the disclosed embodiments, in order to ensure that the multimodal information processing model obtained by training the model to be trained has better label prediction capabilities and can better filter data samples for the target model, the model structure of the model to be trained and the model structure of the target model may be identical. That is, the model structure of the model to be trained and the model structure of the target model are consistent in terms of hierarchical composition, connection methods between layers, and parameter settings of each layer. This allows the model to be trained to filter effective data samples for the target model based on nearly identical knowledge of the target model on a specific dataset.

[0071] In other embodiments, in order to quickly train a multimodal information processing model that can be labeled, the model to be trained can also be a student model of the target model.

[0072] As the name implies, the teacher model is typically a complex and high-performance model, while the student model is a simpler model with fewer parameters. The student model's goal is to approximate the teacher model's performance by learning its outputs. During training, the student model not only learns how to generate output directly from input data but also learns the teacher model's intermediate representation of the input data, thereby capturing more of the teacher model's knowledge.

[0073] This means that before training the target model, the target model can be pre-trained to generate a student model of the target model. The student model has a simple structure, lower memory and hardware requirements, and higher inference speed. Therefore, using the student model can improve the training efficiency of multimodal information processing models while maintaining the accuracy of capability labeling.

[0074] In the disclosed embodiment, the model structure of the model to be trained is the same as the model structure of the target model. There is no need to redesign the model architecture, and the structure of the target model can be quickly applied to new tasks or data sets, saving a lot of time and energy. Due to the same structure, its performance has a certain degree of predictability. If it is known that the target model can perform well under the target task, then the model to be trained with the same structure is also likely to achieve good results under similar conditions, reducing uncertainty, thereby providing a strong foundation for the target model to mine useful data samples. The student model usually has fewer parameters and a simpler structure, which can greatly reduce the amount of computation and memory usage, thereby achieving more efficient reasoning and deployment. In addition, since the student model is small in size, the amount of data and time required for training are relatively small, thereby reducing the cost of data collection and annotation, as well as the computational cost during training, and can also accurately predict capability labels.

[0075] In the disclosed embodiments, the multimodal large model in the disclosed embodiments is more complex than the model to be trained, taking into account hardware and time costs. Therefore, the multimodal large model's ability to label capability labels can be transferred to the model to be trained. The model to be trained can then utilize fewer hardware resources and have a shorter inference process. This allows for faster labeling of capability labels for massive amounts of data.

[0076] Of course, in other embodiments, the ability of the multimodal large model to mine capability labels can be transferred to other large models required by the user, so that the user can reasonably use the trained multimodal information processing model according to needs.

[0077] In the embodiment of the present disclosure, the training model can be trained by supervised training. The training process is as follows: Figure 3 As shown, including the following:

[0078] S301: Input the multimodal data into the model to be trained to obtain the predicted label output by the model to be trained for the multimodal data.

[0079] The multimodal data is input into the model to be trained. The model to be trained will process the input data according to the current parameter settings and output the corresponding prediction label.

[0080] It can be understood that multimodal data contains sub-information from multiple modalities. The model to be trained has processing modules for each modality's sub-information. For example, there may be an image processing module, an audio processing module, and a text processing module. The multiple processing modules can exchange intermediate results to obtain feature expressions that can be used to predict capability labels. Alternatively, the intermediate results obtained by the multiple processing modules can be input into a feature fusion module for feature fusion processing to obtain feature expressions that can predict capability labels.

[0081] S302: Determine the prediction loss based on the prediction label and the capability label.

[0082] The predicted labels output by the model are compared with the labels of the multimodal large model's ability to annotate multimodal data. The prediction loss is calculated using a loss function such as cross entropy loss or mean squared error loss. The choice of loss function can be determined based on the specific requirements of the task.

[0083] S303: Adjust the model parameters of the model to be trained based on the prediction loss.

[0084] Based on the calculated prediction loss, the backpropagation algorithm propagates the loss value from the output layer to the input layer layer by layer, simultaneously calculating the gradient of the parameters at each layer. Optimization algorithms such as gradient descent are then used to update the model parameters to reduce the prediction loss. This process is repeated multiple times until the performance of the trained model on the validation set no longer significantly improves, or until a specified number of training iterations have been completed.

[0085] In the disclosed embodiments, by inputting multimodal data into the model to be trained and obtaining the predictive power of its output, and then determining the prediction loss based on the predictive power and power labels, this helps identify the gap between the model to be trained and the real-world scenarios. This enables the model to better adapt to a variety of different situations, rather than being limited to specific patterns in the training data. This enhances the model's generalization ability and enables it to perform well on new, unseen data.

[0086] In the embodiment of the present disclosure, supervised training of the training model is performed based on multimodal data and capability labels, which can also be implemented as follows: Figure 4 The method shown includes the following:

[0087] S401 , based on the capability labels respectively labeled for the multiple multimodal data by the multimodal large model, an initial training sample set is selected from the multiple multimodal data.

[0088] Specifically, representative, high-quality samples are selected from the labeled multimodal data to construct an initial training sample set. This initial training sample set is used to iteratively optimize the target model based on its task requirements. These samples in the initial training sample set should cover the key capability labels that the target model should possess. For example, all model capabilities that can be included in the capability label can be classified into different levels based on task importance, and key capability labels can be constructed based on the core level of model capabilities. Furthermore, the initial training sample set can be required to maintain data diversity and balance as much as possible to avoid model overfitting or bias.

[0089] S402: Perform at least one round of iterative optimization on the target model using the initial training sample set, so as to use the iteratively optimized target model as the model to be trained.

[0090] The target model is iteratively optimized using the initial training sample set for at least one round. During this iterative optimization process, the model's parameters, structure, or training strategy are adjusted to improve the target model's performance on its target task. It is understood that the target task is the original task for which the target model is expected to possess capabilities, and does not necessarily need to be a task with labeled capabilities.

[0091] S403: Perform supervised training on the to-be-trained model based on the multiple multimodal data and their respective capability labels.

[0092] Based on multiple multimodal data and their respective capability labels, the iteratively optimized model undergoes supervised training. During the training process, the model learns how to map the input multimodal data to the corresponding capability labels and continuously optimizes its internal parameters to improve the accuracy of capability label prediction.

[0093] In the embodiment of the present disclosure, by screening the initial training sample set based on the capability labels that are labeled as multiple multimodal data based on the multimodal large model, high-quality samples required for the target task can be mined. The target model is optimized for at least one round of iterative optimization using the high-quality samples to obtain the model to be trained, thereby enabling the obtained model to be trained to have a relatively high-quality model parameter from the initial state, and based on the target task of the target model. On this basis, the model to be trained is continued to be trained so that it has the ability to label capability labels, which can make the training process of the model to be trained converge faster, improve the training efficiency, and also enable the model to be trained to have a relatively high-quality initial parameter, thereby improving the predictive ability of the capability label.

[0094] The above article describes how to train a multimodal information processing model in the first aspect. In the second aspect, in the embodiment of the present disclosure, based on the same technical concept, a sample screening method is also provided. The method is applicable to the multimodal information processing model trained above. Figure 5 As shown, the sample screening method may include the following:

[0095] S501, obtaining a target data set.

[0096] A target dataset is a collection of data samples collected for the purpose of training or testing a specific model. These data samples are typically multimodal and massive. They can be obtained through various channels, such as public datasets, self-built datasets, or datasets provided by third parties.

[0097] S502: Input the target data set into the multimodal information processing model to obtain capability labels of data samples in the target data set output by the multimodal information processing model.

[0098] The target dataset is then fed into a trained multimodal information processing model. The multimodal information processing model then outputs a capability label for each data sample based on the input data, thereby describing the training capabilities of different data samples for the target model.

[0099] S503 , based on the capability labels of the data samples in the target data set, screening out training samples for training the target model from the target data set.

[0100] That is, according to the specific content and requirements of the capability label, the screening criteria are set, and the training samples meeting the conditions are screened out. These training samples will be used for the training of the subsequent target model. In the embodiments of the present disclosure, the specific screening criteria can meet the following requirements:

[0101] The data samples in the target data set that meet the first preset requirement of the number of capabilities included in the capability label are screened out as training samples; and / or, the data samples in the target data set that meet the second preset requirement of the proportion of the total amount of training samples screened out from the target data set in the target data set are screened out as training samples.

[0102] Based on the foregoing content, the capability label represents which model capabilities the multi-modal data can train the target model. Therefore, in the obtained target data set, the model capabilities that different target data can train also differ. Taking the question and answer pairs in the target data set as an example, a simple sample can only train the image recognition capability of the model, while a complex sample can train multiple capabilities of the model, such as image recognition capability, positioning capability, classification capability, etc. Compared with the second type of data sample that can provide a large amount of model capabilities, the first type of data sample provides low-density capabilities, and the first type of data sample is a low-information-density data sample compared with the second type of data sample, and vice versa. The second type of data sample is a high-information-density data sample compared with the first type of data sample. In the case of screening a limited amount of data samples from a large amount of target data set, the second type of data sample will be preferentially selected, thereby eliminating low-information-density data samples.

[0103] Therefore, in the embodiments of the present disclosure, the first preset requirement refers to a specific threshold of the number of capabilities that should be included in the capability label. For example, the first preset requirement can be that each target data sample includes at least m different capability labels.

[0104] The second preset requirement refers to the proportion of the total amount of training samples screened out from the target data set in the target data set. That is, after screening out the data samples meeting the first preset requirement, the total proportion of these samples in the target data set is calculated.

[0105] In implementation, in the case where two preset requirements need to be met, the number of labels included in the first preset requirement can be adjusted to meet the second preset requirement. For example, after preliminary screening, each target data sample includes at least k different capability labels, and the data samples meeting the first preset requirement account for 60% of the total data set. If the second preset requirement is 80%, the screening condition needs to be adjusted, such as reducing the threshold of the number of capabilities to n (n is less than k), and then re-screening until the 80% proportion requirement is met.

[0106] Therefore, when the data set is very large, by setting the number of capability labels or the proportion of the total number of samples in the overall data set, the quality and number of training sample sets can be ensured, thereby improving the efficiency of model training and helping to avoid wasting computing resources on too much invalid or low-quality data.

[0107] In summary, in the disclosed embodiments, a target dataset is input into a multimodal information processing model. The multimodal information processing model generates corresponding capability labels for the data samples. Based on these capability labels, higher-quality training samples that better meet the training requirements of the target model are screened, thereby improving the training effectiveness and generalization capabilities of the target model. Screening the training samples used to train the target model reduces unnecessary data involved in training, reduces computing resource consumption, and improves training efficiency.

[0108] In summary, when faced with a large amount of multimodal data with training labels, such as Figure 6 As shown, the training data of the model to be trained can be generated by combining the existing multimodal large model, such as the GPT4-o large model, with the preset prompt (prompt template). A multimodal information processing model, namely the capability label expert model, is trained based on the target model through the training data. Afterwards, based on the trained multimodal information processing model, the massive multimodal data is labeled and filtered. Among them, by labeling each multimodal data, the model capability corresponding to each multimodal data can be understood, which is not only conducive to screening out high-quality training samples, but also when it is assessed that the samples of a certain model capability are insufficient, it can be found which model capabilities are lacking, and it is helpful to screen samples of different model capabilities, and even supplement the training samples corresponding to the model capability. Filtering each multimodal data can filter out the available target data for training the target model. By comparison, the solution provided by the present disclosure can filter out 50% of low-density data to ensure consistency with the accuracy of the original model.

[0109] Based on the same technical concept, the embodiment of the present disclosure also provides a training device 700 for a multimodal information processing model, such as Figure 7 Shown, including:

[0110] The labeling module 701 is used to label the multimodal data based on the multimodal large model to obtain the capability label of the multimodal data; the capability label is used to describe the model capability that can be trained with the multimodal data;

[0111] The training module 702 is used to perform supervised training on the model to be trained based on the multimodal data and capability labels, so as to train the model to be trained into a multimodal information processing model that can filter out useful samples for the target model.

[0112] In some embodiments, the annotation module includes:

[0113] A construction unit, configured to construct input information of a multimodal large model based on multimodal data and a preset prompt word template;

[0114] The labeling unit is used to input the input information into the multimodal large model to obtain the ability label of the multimodal large model for labeling the multimodal data.

[0115] In some embodiments, the preset prompt word template includes at least one of the following requirements:

[0116] The multimodal large model is required to pay attention to the detailed content of the image modality information in the multimodal data;

[0117] When multimodal data includes question-answer pairs, a large multimodal model is required to combine the question-answer pairs to understand the query intent of the question in the question-answer pair for the image modality information in the multimodal data.

[0118] Large multimodal models are required to analyze and reason about the input information to provide reasonable capability labels for multimodal data.

[0119] In some embodiments, the model structure of the model to be trained is the same as the model structure of the target model; or,

[0120] The model to be trained is the student model of the target model.

[0121] In some embodiments, the model structure of the multimodal large model is more complex than the model to be trained.

[0122] In some embodiments, the training module includes:

[0123] A prediction unit, configured to input multimodal data into the model to be trained and obtain a predicted label output by the model to be trained for the multimodal data;

[0124] a determination unit, configured to determine a prediction loss based on the prediction label and the capability label;

[0125] The first optimization unit is used to adjust the model parameters of the to-be-trained model based on the prediction loss.

[0126] In some embodiments, the training module includes:

[0127] A screening unit, configured to screen an initial training sample set from the multiple multimodal data based on capability labels respectively labeled by the multimodal large model for the multiple multimodal data;

[0128] A second optimization unit is configured to perform at least one round of iterative optimization on the target model using the initial training sample set, so as to use the iteratively optimized target model as the model to be trained;

[0129] The training unit is used to perform supervised training on the training model based on multiple multimodal data and their respective capability labels.

[0130] Based on the same technical concept, the embodiment of the present disclosure also provides a sample screening device 800, such as Figure 8 Shown, including:

[0131] Acquisition module 801, used to acquire target data set;

[0132] Output module 802, configured to input the target data set into the multimodal information processing model and obtain capability labels of data samples in the target data set output by the multimodal information processing model;

[0133] The screening module 803 is used to screen out training samples for training the target model from the target data set based on the capability labels of the data samples in the target data set.

[0134] In some embodiments, the screening module comprises:

[0135] A first screening unit is configured to screen out data samples from a target data set, the data samples having the number of capabilities included in the capability labels meeting a first preset requirement, as training samples; and / or,

[0136] The second screening unit is configured to screen out data samples whose proportion of the total amount of training samples in the target data set meets a second preset requirement from the target data set based on the capability label, as training samples.

[0137] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0138] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0139] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0140] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0141] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0142] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0143] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the training method and / or sample screening method of the multimodal information processing model. For example, in some embodiments, the training method and / or sample screening method of the multimodal information processing model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the training method and / or sample screening method of the multimodal information processing model described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the training method and / or the sample screening method of the multimodal information processing model in any other appropriate manner (for example, by means of firmware).

[0144] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0148] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0149] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0150] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0151] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for training a multimodal information processing model, comprising: Labeling the multimodal data based on the multimodal large model to obtain capability labels of the multimodal data; The capability label is used to describe the model capability that can be trained by the multimodal data, so that the target model has the ability to perform the task category corresponding to the capability label; Based on the multimodal data and the capability labels, supervised training is performed on the model to be trained to train the model to be trained into a multimodal information processing model capable of screening useful samples for the target model, including: Inputting the multimodal data into the model to be trained to obtain a predicted label output by the model to be trained for the multimodal data; determining a prediction loss based on the prediction label and the capability label; Adjust model parameters of the to-be-trained model based on the prediction loss.

2. The method according to claim 1, wherein labeling the multimodal data based on the multimodal large model to obtain capability labels of the multimodal data comprises: Constructing input information of the multimodal large model based on the multimodal data and a preset prompt word template; The input information is input into the multimodal large model to obtain a capability label that the multimodal large model labels for the multimodal data.

3. The method according to claim 2, wherein the preset prompt word template includes at least one of the following requirements: Require the multimodal large model to pay attention to the detailed content of the image modality information in the multimodal data; In the case where the multimodal data includes question-answer pairs, the multimodal large model is required to combine the question-answer pairs to understand the query intent of the question in the question-answer pairs with respect to the image modality information in the multimodal data; The multimodal large model is required to analyze and reason about the input information to provide reasonable capability labels for the multimodal data.

4. The method according to claim 1, wherein The model structure of the model to be trained is the same as the model structure of the target model; or, The model to be trained is a student model of the target model.

5. According to the method of claim 1, the complexity of the model structure of the multimodal large model is greater than that of the model to be trained.

6. The method according to claim 1, wherein The method of performing supervised training on the model to be trained based on the multimodal data and the capability label further includes: Based on the capability labels respectively labeled by the multimodal large model for the plurality of multimodal data, an initial training sample set is selected from the plurality of multimodal data; Performing at least one round of iterative optimization on the target model using the initial training sample set, so as to use the iteratively optimized target model as the model to be trained; Based on the multiple multimodal data and their respective capability labels, supervised training is performed on the model to be trained.

7. A sample screening method, applied to a multimodal information processing model trained by the method of any one of claims 1 to 6, comprising: Get the target dataset; Inputting the target data set into the multimodal information processing model to obtain capability labels of data samples in the target data set output by the multimodal information processing model; Based on the capability labels of the data samples in the target data set, training samples for training the target model are screened from the target data set.

8. The method according to claim 7, wherein: The step of screening out training samples for training a target model from the target dataset based on the capability labels of the data samples in the target dataset includes: Filtering data samples from the target data set, the data samples having the number of capabilities included in the capability labels meeting the first preset requirement, as the training samples; and / or, Based on the capability labels, data samples whose proportion of the total amount of training samples in the target data set meets a second preset requirement are screened out from the target data set as the training samples.

9. A training device for a multimodal information processing model, comprising: A labeling module, configured to label the multimodal data based on the multimodal large model to obtain capability labels of the multimodal data; The capability label is used to describe the model capability that can be trained by the multimodal data, so that the target model has the ability to perform the task category corresponding to the capability label; A training module, configured to perform supervised training on the model to be trained based on the multimodal data and the capability label, so as to train the model to be trained into a multimodal information processing model capable of screening useful samples for the target model, comprising: A prediction unit, configured to input the multimodal data into the model to be trained, and obtain a prediction label output by the model to be trained for the multimodal data; a determining unit, configured to determine a prediction loss based on the prediction label and the capability label; A first optimization unit is used to adjust the model parameters of the to-be-trained model based on the prediction loss.

10. The apparatus according to claim 9, wherein the annotation module comprises: A construction unit, configured to construct input information of the multimodal large model based on the multimodal data and a preset prompt word template; The labeling unit is used to input the input information into the multimodal large model to obtain a capability label that the multimodal large model labels for the multimodal data.

11. The device according to claim 10, wherein The preset prompt word template includes at least one of the following requirements: Require the multimodal large model to pay attention to the detailed content of the image modality information in the multimodal data; In the case where the multimodal data includes question-answer pairs, the multimodal large model is required to combine the question-answer pairs to understand the query intent of the question in the question-answer pairs with respect to the image modality information in the multimodal data; The multimodal large model is required to analyze and reason about the input information to provide reasonable capability labels for the multimodal data.

12. The device according to claim 9, wherein The model structure of the model to be trained is the same as the model structure of the target model; or, The model to be trained is a student model of the target model.

13. The device according to claim 9, wherein the complexity of the model structure of the multimodal large model is greater than that of the model to be trained.

14. The device according to claim 9, wherein The training module further includes: a screening unit, configured to screen an initial training sample set from the plurality of multimodal data based on the capability labels respectively labeled by the multimodal large model for the plurality of multimodal data; a second optimization unit, configured to perform at least one round of iterative optimization on the target model using the initial training sample set, so as to use the iteratively optimized target model as the model to be trained; A training unit is used to perform supervised training on the model to be trained based on the multiple multimodal data and their respective capability labels.

15. A training sample screening device, applied to a multimodal information processing model trained by the device of any one of claims 9 to 14, comprising: Acquisition module, used to obtain the target data set; an output module, configured to input the target data set into the multimodal information processing model and obtain capability labels of data samples in the target data set output by the multimodal information processing model; The screening module is used to screen out training samples for training a target model from the target data set based on the capability labels of the data samples in the target data set.

16. The device according to claim 15, wherein The screening module comprises: A first screening unit is configured to screen out data samples whose capability quantity included in the capability labels meets a first preset requirement from the target data set as the training samples; and / or, The second screening unit is configured to screen out data samples from the target data set based on the capability label, the data samples having a proportion of the total amount of training samples in the target data set that meets a second preset requirement, as the training samples.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method, device and equipment for training network model based on deep learning framework network

    CN116911361A

  • Auxiliary labeling method and device, object recognition method and device and electronic equipment

    CN117454149A