Method, apparatus, device, and storage medium for training a multimodal model

In the pre-training process of multimodal models, sample data of different modalities are obtained for mask processing, predicted features are generated and original mask features are replaced, and the problem that the loss value is generated only by masked features is solved, which improves the efficiency of multimodal model training.

CN114819180BActive Publication Date: 2025-06-27BEIJING SANKUAI ONLINE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210348800.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-06-27
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

In the pre-training process of multimodal models, the loss value is generated only by the masked features, resulting in limited features participating in training and low efficiency.

Method used

By obtaining sample data of different modes for mask processing, predictive features are generated, and the predictive features are replaced with the original mask feature, and the predictive features are flattened and discriminant processing is performed. Finally, the multimodal model is trained based on the prediction features and discriminant results.

Benefits of technology

The efficiency of multimodal model training is improved, feature utilization is increased, and more features are involved in the calculation of loss value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819180B_ABST
    Figure CN114819180B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for training a multimodal model, belonging to the technical field of machine learning. The method includes: obtaining first feature data and second feature data after mask processing corresponding to first sample data and second sample data; respectively performing prediction processing on the masked features to obtain a first predicted feature and a second predicted feature; replacing the masked features in the first feature data with the first predicted feature to obtain third feature data, and replacing the masked features in the second feature data with the second predicted feature to obtain fourth feature data; performing discrimination processing on each feature in the combined feature data corresponding to the third feature data and the fourth feature data to obtain a discrimination result on whether each feature in the combined feature data is replaced; training the multimodal model based on the first predicted feature, the second predicted feature and the discrimination result. Using the present application can improve the efficiency of multimodal model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of machine learning, and particularly relates to a method, apparatus, device, and storage medium for training a multimodal model. Background Art

[0002] With the development of Internet technology and machine learning technology, multimodal technology has also begun to be widely applied. Multimodal technology is a technology for learning and understanding the semantic information of data in different modalities, and application scenarios include multimodal data matching, multimodal data retrieval, and visual question answering, etc. Among them, data in different modalities can include text, images, audio, etc.

[0003] In multimodal technology, a multimodal model can be used to process data in different modalities, determine whether the data in different modalities match, or understand the semantic information jointly expressed by data in multiple modalities, and judge the category to which the data in each modality belongs, etc. Before applying the multimodal model, the multimodal model can be pre-trained to improve the effect of the multimodal model in various application scenarios.

[0004] In the current pre-training of multimodal models, a generation module such as MLM (Masked Language Model), MOC (Masked Object Classification), etc. also needs to be added to the multimodal model. When pre-training the multimodal model, the sample data or the feature data corresponding to the sample data can be masked first. After obtaining the masked feature data corresponding to the sample data in different modalities, the obtained feature data can be input into a feature concatenation module for concatenation processing to obtain concatenated feature data, and then the concatenated feature data can be input into the generation module, and the generation module can perform prediction processing on the masked features in the concatenated feature data to generate corresponding predicted features. Then, the corresponding loss value can be determined according to the predicted features generated by the generation module and the masked features, and the generation module, concatenation module, etc. in the multimodal model can be trained according to the loss value. After completing the pre-training of the multimodal model, the generation module in the multimodal model can be deleted.

[0005] During the pre-training of the multimodal model, the loss value is only generated by the masked features in the sample data. Therefore, during the pre-training of the multimodal model, only the masked features participate in the training process of the multimodal model, which results in low efficiency in training the multimodal model. Summary of the Invention

[0006] Embodiments of this application provide a method, apparatus, device, and storage medium for training a multimodal model, which can improve the training efficiency of the multimodal model. The technical solutions are as follows:

[0007] In a first aspect, a method for training a multi-modal model is provided. The method includes:

[0008] Obtain first feature data after mask processing corresponding to first sample data, and second feature data after mask processing corresponding to second sample data, where the first sample data and the second sample data belong to data of different modalities;

[0009] Perform prediction processing on the masked features in the first feature data to obtain first predicted features, and perform prediction processing on the masked features in the second feature data to obtain second predicted features;

[0010] Replace the masked features in the first feature data with the first predicted features to obtain third feature data, and replace the masked features in the second feature data with the second predicted features to obtain fourth feature data;

[0011] Perform splicing processing on the third feature data and the fourth feature data to obtain spliced combined feature data, and perform discrimination processing on each feature in the combined feature data to obtain a discrimination result on whether each feature in the combined feature data is replaced;

[0012] Train the multi-modal model based on the first predicted features, the second predicted features, and the discrimination result.

[0013] Optionally, the first sample data is image data, and the second sample data is text data.

[0014] Optionally, the performing discrimination processing on each feature in the combined feature data to obtain a discrimination result on whether each feature in the combined feature data is replaced includes:

[0015] Perform discrimination processing on each key image feature included in the combined feature data to obtain a first discrimination result on whether each image feature included in the combined feature data is replaced;

[0016] Perform discrimination processing on each word feature included in the combined feature data to obtain a second discrimination result on whether each word feature included in the combined feature data is replaced;

[0017] The training the multi-modal model based on the first predicted features, the second predicted features, and the discrimination result includes:

[0018] Train the multi-modal model based on the first predicted features, the second predicted features, the first discrimination result, and the second discrimination result.

[0019] Optionally, the multimodal model includes a generation module, a splicing module, and a discrimination module. The generation module is used to perform the prediction process, the splicing module is used to perform the splicing process, and the discrimination module is used to perform the discrimination process;

[0020] Training the multimodal model based on the first prediction feature, the second prediction feature, the first discrimination result, and the second discrimination result includes:

[0021] Training the generation module based on the first prediction feature and the second prediction feature;

[0022] Training the generation module, the splicing module, and the discrimination module based on the first discrimination result and the second discrimination result.

[0023] Optionally, training the generation module based on the first prediction feature and the second prediction feature includes:

[0024] Determining a first loss value corresponding to the masked feature in the first prediction feature and the first feature data, and determining a second loss value corresponding to the masked feature in the second prediction feature and the second feature data;

[0025] Training the generation module based on the first loss value and the second loss value.

[0026] Optionally, training the generation module, the splicing module, and the discrimination module based on the first discrimination result and the second discrimination result includes:

[0027] Determining a third loss value based on the first discrimination result and first reference information, where the first reference information is used to indicate whether each key image feature included in the spliced feature data is replaced;

[0028] Determining a fourth loss value based on the second discrimination result and second reference information, where the second reference information is used to indicate whether each word feature included in the spliced feature data is replaced;

[0029] Training the generation module, the splicing module, and the discrimination module based on the third loss value and the fourth loss value.

[0030] Optionally, the multimodal model further includes a matching module, which is used to perform a matching process based on features of different modalities in the spliced feature data to obtain a matching result of the first sample data and the second sample data;

[0031] Determine a fifth loss value based on the reference matching values corresponding to the first sample data and the second sample data and the matching result;

[0032] Train the multimodal model based on the fifth loss value.

[0033] In a second aspect, a device for training a multimodal model is provided, and the device includes:

[0034] An acquisition unit, configured to acquire first feature data after mask processing corresponding to first sample data, and second feature data after mask processing corresponding to second sample data, where the first sample data and the second sample data belong to data of different modalities;

[0035] A prediction unit, configured to perform prediction processing on the masked features in the first feature data to obtain first predicted features, and perform prediction processing on the masked features in the second feature data to obtain second predicted features;

[0036] A replacement unit, configured to replace the masked features in the first feature data with the first predicted features to obtain third feature data, and replace the masked features in the second feature data with the second predicted features to obtain fourth feature data

[0037] A discrimination unit, configured to perform a splicing process on the third feature data and the fourth feature data to obtain spliced feature data after the splicing process, and perform discrimination processing on each feature in the spliced feature data to obtain a discrimination result on whether each feature in the spliced feature data is replaced;

[0038] A training unit, configured to train the multimodal model based on the first predicted features, the second predicted features, and the discrimination result.

[0039] Optionally, the first sample data is image data, and the second sample data is text data.

[0040] Optionally, the discrimination unit is configured to:

[0041] Perform discrimination processing on each key image feature included in the spliced feature data to obtain a first discrimination result on whether each image feature included in the spliced feature data is replaced;

[0042] Perform discrimination processing on each word feature included in the spliced feature data to obtain a second discrimination result on whether each word feature included in the spliced feature data is replaced;

[0043] The training unit is configured to: train the multimodal model based on the first prediction feature, the second prediction feature, the first discrimination result, and the second discrimination result.

[0044] Optionally, the multimodal model includes a generation module, a splicing module, and a discrimination module. The generation module is configured to perform the prediction process, the splicing module is configured to perform the splicing process, and the discrimination module is configured to perform the discrimination process.

[0045] The training unit is configured to: train the generation module based on the first prediction feature and the second prediction feature; train the generation module, the splicing module, and the discrimination module based on the first discrimination result and the second discrimination result.

[0046] Optionally, the training unit is configured to:

[0047] Determine a first loss value corresponding to the masked feature in the first feature data for the first prediction feature, and determine a second loss value corresponding to the masked feature in the second feature data for the second prediction feature.

[0048] Train the generation module based on the first loss value and the second loss value.

[0049] Optionally, the training unit is configured to:

[0050] Determine a third loss value based on the first discrimination result and first reference information, where the first reference information is used to indicate whether each key image feature included in the spliced feature data is replaced.

[0051] Determine a fourth loss value based on the second discrimination result and second reference information, where the second reference information is used to indicate whether each word feature included in the spliced feature data is replaced.

[0052] Train the generation module, the splicing module, and the discrimination module based on the third loss value and the fourth loss value.

[0053] Optionally, the multimodal model further includes a matching module, configured to perform a matching process based on features of different modalities in the spliced feature data to obtain a matching result of the first sample data and the second sample data.

[0054] The training unit is further configured to: determine a fifth loss value based on a reference matching value corresponding to the first sample data and the second sample data and the matching result.

[0055] Train the multimodal model based on the fifth loss value.

[0056] In a third aspect, a computer device is provided, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the operations performed by the method for training a multi-modal model as described in the first aspect above.

[0057] In a fourth aspect, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the method for training a multi-modal model as described in the first aspect above.

[0058] In a fifth aspect, a computer program product is provided. The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the method for training a multi-modal model as described above.

[0059] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are as follows:

[0060] In the embodiments of the present application, after obtaining the predicted features corresponding to the masked features in the feature data of different modal sample data, the predicted features can be used to replace the masked features in the feature data, and then it is determined whether each feature in the replaced feature data has been replaced to obtain the discrimination result of whether each feature has been replaced. Then, the multi-modal model is trained according to the predicted features and the discrimination result. In this way, when training the multi-modal model according to the predicted features and the discrimination result, in addition to the predicted features participating in the training process of the multi-modal model, the result of whether each feature has been replaced will also participate in the training process of the multi-modal model. In this way, whether the features in the feature data are masked or not will participate in the calculation of the loss value, which can improve the utilization rate of the sample data and thus improve the training efficiency of the multi-modal model. Description of the Drawings

[0061] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 is a schematic structural diagram of a multi-modal model provided by an embodiment of the present application;

[0063] Figure 2 is a flowchart of a method for training a multi-modal model provided by an embodiment of the present application;

[0064] Figure 3It is a schematic diagram of a multimodal model structure provided by an embodiment of the present application;

[0065] Figure 4 It is a flowchart of a method for training a multimodal model provided by an embodiment of the present application;

[0066] Figure 5 It is a flowchart of a method for training a multimodal model provided by an embodiment of the present application;

[0067] Figure 6 It is a schematic diagram of the structure of a device for training a multimodal model provided by an embodiment of the present application;

[0068] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0069] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0070] A method for training a multimodal model provided by the present application can be implemented by a computer device, which can be a terminal or a server. When the computer device is a terminal, the terminal can be a mobile phone, a tablet computer, a smart wearable device, a desktop computer, a laptop computer, etc. When the computer device is a server, the server can be a single server or a server group. If it is a single server, the server can be responsible for all the processing in the following solutions. If it is a server group, different servers in the server group can be responsible for different processing in the following solutions respectively. The specific processing allocation can be set arbitrarily by technicians according to actual needs and will not be elaborated here.

[0071] The computer device may include a processor and a memory. Among them, the data required to implement the method for training a multimodal model provided by the embodiment of the present application can be stored in the memory. For example, it can be the program code corresponding to the method for training a multimodal model, the sample data required for training a multimodal model, and so on. The processor can execute the program code stored in the memory and process the sample data stored in the memory, so as to realize the training of the multimodal model.

[0072] As Figure 1 shown, in the multimodal model, there are multiple feature extraction modules, a splicing module, a matching module, etc. In the process of applying the multimodal model, different modalities of data can be processed through the extraction module, the splicing module, and the matching module, etc., so as to determine whether the data of different modalities match, or understand the semantic information jointly expressed by multiple modalities and judge its category, etc.

[0073] Generally, the training of a multi-modal model includes pre-training and fine-tuning. In fact, both pre-training and fine-tuning are to train the multi-modal model. Pre-training means training each module in the multi-modal model according to the sample data common to multiple scenarios, so as to improve the usage effect of the multi-modal model in various application scenarios. Fine-tuning means training the pre-trained multi-modal model again according to the sample data corresponding to the actual application scenario of the multi-modal model, so as to improve the usage effect of the multi-modal model in the actual application scenario.

[0074] When pre-training a multi-modal model, a generation module (Generator) can be added to the multi-modal model, such as MLM (Masked Language Model), MOC (Masked Object Classification), etc. When pre-training the multi-modal model, the sample data or the feature data corresponding to the sample data can be masked first. For example, if the sample data is text data: "An old man swimming in a pool", the masked sample data may be "An old[mask]swimming in a pool". Then, the masked sample data can be input into the feature extraction module to obtain the corresponding masked feature data. After obtaining the masked feature data corresponding to the sample data of different modalities, the obtained masked feature data can be input into the splicing module for splicing processing to obtain the spliced feature data. Then, the spliced feature data is input into the generation module, and the generation module predicts the masked features in the spliced feature data to generate the corresponding predicted features. For example, for "An old[mask]swimming in a pool", the masked feature is the word feature corresponding to "man", and the generated predicted feature may be the word feature corresponding to "woman". Then, the corresponding loss value can be determined according to the predicted features generated by the generation module and the masked features, and the feature extraction module, splicing module, etc. in the multi-modal model can be pre-trained according to the loss value. After completing the pre-training of the multi-modal model, the multi-modal model can be trained again according to the actual application scenario of the multi-modal model. When training the multi-modal model again, the generation module can be removed, and the multi-modal model can be trained again according to the samples corresponding to the actual application scenario.

[0075] It can be seen that in the related technology, each time the multi-modal model is pre-trained, the loss value is only generated by the masked features in the sample data. In this way, during the pre-training process of the multi-modal model, only the masked features participate in the training of the multi-modal model, resulting in low training efficiency of the multi-modal model.

[0076] The method for training a multimodal model provided by an embodiment of the present application can be applied to the training process of the multimodal model, which can improve the utilization rate of sample data in the training process of the multimodal model, and further improve the training efficiency of the multimodal model. Refer to Figure 2 , Figure 2 which is a flowchart of a method for training a multimodal model provided by an embodiment of the present application. The method includes:

[0077] Step 201, obtain the first feature data after mask processing corresponding to the first sample data, and the second feature data after mask processing corresponding to the second sample data.

[0078] Among them, the first sample data and the second sample data belong to data of different modalities. For example, the first sample data may be image data, text data, audio data, video data, etc., and the second sample data may be data belonging to a different modality from the first sample data.

[0079] Taking the first sample data as image data and the second sample data as text data as an example below, the method flow provided by an embodiment of the present application will be described in detail. Other situations are similar and will not be elaborated.

[0080] Among them, the first sample data (subsequently referred to as the sample image) can be input into the image feature extraction module included in the multimodal model, and the key image features in the sample image are output by the image feature extraction module. For example, if the input sample image is an image including a human body, the key image features may include the image features of the face, the image features of the limbs, and so on. After obtaining multiple key image features in the sample image, mask processing can be performed on the multiple key image features, that is, at least one of the multiple key image features can be masked, that is, at least one of the multiple key image features can be replaced with a preset image mask feature, so as to obtain the first feature data after mask processing corresponding to the sample image. It should be noted that the image feature extraction module included in the multimodal model can be a pre-trained image feature extraction module. For example, it can be Faster RCNN (a target recognition algorithm), and its training process belongs to the prior art and will not be introduced in detail here.

[0081] For the second masked feature data corresponding to the second sample data (hereinafter referred to as the sample text), the second sample data can be masked first. For example, if the sample data is text data: "An old man swimming ina pool", the corresponding masked sample data may be "An old[mask]swimming in a pool". Then, the masked sample data can be input into the text feature extraction module included in the multimodal model to obtain the corresponding second masked feature data. In the second masked feature data, it includes the word features corresponding to each unmasked word, such as "An", "old", "swimming", "in", "a", "pool", as well as the preset word feature corresponding to [mask]. Among them, the text feature extraction module can be text encoding feature (Word Embedding).

[0082] Step 202: Perform prediction processing on the masked features in the first feature data to obtain the first predicted feature, and perform prediction processing on the masked features in the second feature data to obtain the second predicted feature.

[0083] After obtaining the first masked feature data corresponding to the sample image and the second masked feature data corresponding to the sample text, the first masked feature data can be input into the generation module. The generation module predicts and processes the masked features based on the unmasked features in the first feature data to generate the first predicted feature corresponding to the masked features. And the second masked feature data can be input into the generation module. The generation module predicts and processes the masked features based on the unmasked features in the second feature data to generate the second predicted feature corresponding to the masked features.

[0084] Among them, when the first sample data is image data and the second sample data is text data, the corresponding generation module may include MLM (Masked Language Model) and MOC (Masked Object Classification). The first feature data includes image mask features and key image features that are not masked. The second feature data includes word mask features and word features corresponding to the unmasked words.

[0085] Such as Figure 3 shown Figure 3 is the structural schematic diagram of the multimodal model provided by the embodiment of the present application during training.

[0086] After obtaining the first feature data after mask processing corresponding to the sample image and the second feature data after mask processing corresponding to the sample text, the first feature data after mask processing can be input into the MOC, and the MOC performs prediction processing on the masked key image features in the first feature data, and outputs the first prediction feature for predicting the masked key image features. For example, if the sample image is an image including a human body, the key image features may include the image features of the face, the image features corresponding to the limbs respectively, the image features of the abdomen, and so on. Assuming that the masked feature is the face image feature, the MOC may generate the face image feature (i.e., generate the first prediction feature) based on the image features of the limbs and the abdomen, but may generate non-face image features due to insufficient training times of the MOC or certain errors.

[0087] Similarly, the second feature data after mask processing corresponding to the sample text can be input into the MLM, and the MLM performs prediction processing on the masked word features in the second feature data, and outputs the second prediction feature for predicting the masked word features. For example, the sample text is "An old man swimming in apool". Assuming that "man" is the masked word, the MLM can predict the word features of the masked word based on the word features corresponding to "An", "old", "swimming", "in", "a", "pool". Ideally, the word features corresponding to "man" can be predicted (i.e., generate the second prediction feature), but due to insufficient training times of the MLM or certain errors, the word features of other words other than "man" may be generated. In another possible implementation, the MLM can directly generate the masked word based on the word features corresponding to "An", "old", "swimming", "in", "a", "pool". Then the generated masked word can be input into the WordEmbedding, and the Word Embedding outputs the word features corresponding to the masked word generated by the MLM (i.e., the second prediction feature).

[0088] Step 203: Replace the masked features in the first feature data with the first prediction feature to obtain the third feature data, and replace the masked features in the second feature data with the second prediction feature to obtain the fourth feature data.

[0089] After obtaining the first prediction feature and the second prediction feature, the masked features in the first feature data can be replaced with the first prediction feature to obtain the third feature data, and the masked features in the second feature data can be replaced with the second prediction feature to obtain the fourth feature data.

[0090] Continuing with the example in step 202 above, assuming that the first predicted feature obtained according to MOC is the image feature of a human hand, the third feature data obtained includes the image feature of the human hand, the image feature of the limbs, and the image feature of the abdomen. Assuming that the second predicted feature obtained according to MLM is the word feature corresponding to "woman", the corresponding fourth feature data can be the word features corresponding to "An", "old", "woman", "swimming", "in", "a", "pool" respectively.

[0091] Step 204: Perform a splicing process on the third feature data and the fourth feature data to obtain the spliced feature data after the splicing process, and perform a discrimination process on each feature in the spliced feature data to obtain the discrimination result of whether each feature in the spliced feature data is replaced.

[0092] After obtaining the third feature data and the fourth feature data, the third feature data and the fourth feature data can be input into a splicing module, and the splicing module performs feature mapping processing and splicing processing on the third feature data and the fourth feature data, etc., and finally obtains the spliced feature data.

[0093] As Figure 3 shown, in the splicing module, a feature mapping sub-module and a feature splicing sub-module can be included. Among them, the feature mapping sub-module can be a self-attention layer (Self-attention Layers), which can include an image feature mapping sub-module and a text mapping sub-module, and the feature splicing sub-module can be a cross-attention layer (Cross-attention Layers). After obtaining the third feature data and the fourth feature data, the third feature data can be input into the image feature mapping sub-module to obtain the mapped third feature data and the mapped fourth feature data. Then, the mapped third feature data and the mapped fourth feature data are input into the feature splicing sub-module for splicing processing to obtain the spliced feature data.

[0094] It should be noted that the obtained spliced feature data consists of three parts, including the third feature data (subsequently referred to as the fifth feature data) after feature mapping and other processing, the fourth feature data (subsequently referred to as the sixth feature data) after feature mapping and other processing, and the position information used to distinguish the fifth feature data and the sixth feature data. It can be understood that each feature included in the fifth feature data after feature mapping and other processing is the key image feature included in the third feature data and the feature after feature mapping and other processing of the first predicted feature, so it can still be considered that the fifth feature data includes the key image feature of the sample image and the first predicted feature. Similarly, the sixth feature data can include the word feature of the sample text and the second predicted feature.

[0095] After obtaining the combined feature data, the combined feature data can be input into a discrimination module, which can be a discriminator. The discrimination module can perform discrimination processing on the combined feature data to determine the discrimination result of whether each feature in the combined feature data is replaced. That is, to determine the first predicted feature among the key image features included in the fifth feature data and the second predicted feature among the word features included in the sixth feature data.

[0096] Step 205: Train the multi-modal model based on the first predicted feature, the second predicted feature, and the discrimination result.

[0097] After obtaining the first predicted feature, the second predicted feature, and the discrimination result, the multi-modal model can be trained. That is, the corresponding loss values can be calculated respectively according to the first predicted feature, the second predicted feature, and the discrimination result, and the parameters of each module in the multi-modal model can be adjusted according to the loss values.

[0098] Optionally, in step 204, the discrimination module may include an image feature discrimination sub-module and a text discrimination sub-module. Among them, the image feature discrimination sub-module can be used to judge the features replaced in the image feature data in the combined feature data. The text feature discrimination sub-module can be used to judge the features replaced in the text feature data in the combined feature data. The image feature discrimination module can be an R 2 D Task (Replaced ROIDetection Task, the discrimination task of whether the picture feature area is replaced), and the text feature discrimination module can be an RTD Task (Replaced Token Detection Task, the discrimination task of whether the text is replaced). For specific processing, please refer to Figure 4 , as follows:

[0099] Step 401: Input the combined feature data into the image feature discrimination module to obtain the first discrimination result of whether each image feature included in the combined feature data is replaced.

[0100] After obtaining the combined feature data, the combined feature data can be respectively input into the image feature discrimination module. The image feature discrimination module predicts the masked key image features in the corresponding image feature part of the combined feature data to obtain the first discrimination result. That is, the combined feature data can be input into the R 2 D Task, and the R 2 D Task judges whether each feature included in the fifth feature data in the combined feature data has been replaced and outputs the discrimination result (i.e., the first discrimination result). Among them, the first discrimination result includes the result of whether each feature included in the fifth feature data has been replaced. For example, if the fifth feature data includes 5 key image features, then R2 The D Task can output a 5-bit binary string, which is the first discrimination result. If the output binary string is "00100", it means that the third key image feature among the 5 key image features is the replaced key image feature (i.e., the first predicted feature).

[0101] Step 402: Input the combined feature data into the text feature discrimination module to obtain the second discrimination result on whether each word feature included in the combined feature data is replaced.

[0102] After obtaining the combined feature data, the combined feature data can also be input into the text feature discrimination module respectively. The text feature discrimination module predicts the masked word features in the part of the corresponding text features in the combined feature data to obtain the second discrimination result. That is, the combined feature data can be input into the RTD Task, and the RTD Task determines whether each feature included in the sixth feature data in the combined feature data has been replaced and outputs the discrimination result (i.e., the second discrimination result). Among them, the second discrimination result includes the result of whether each feature included in the sixth feature data has been replaced. For example, if the sixth feature data includes 8 word features, the RTD Task can output an 8-bit binary string, which is the second discrimination result. If the output binary string is "00000100", it means that the 6th word feature among the 8 key image features is the replaced word feature (i.e., the second predicted feature).

[0103] Step 403: Train the multi-modal model based on the first predicted feature, the second predicted feature, the first discrimination result, and the second discrimination result.

[0104] After obtaining the first discrimination result and the second discrimination result, the corresponding loss values can be calculated according to the first predicted feature, the second predicted feature, the first discrimination result, and the second discrimination result. Then, the multi-modal model can be trained according to the corresponding loss values.

[0105] For the further process of training the multi-modal model according to the loss value, please refer to Figure 5 , as follows:

[0106] Step 501: Train the generation module based on the first predicted feature and the second predicted feature.

[0107] Among them, for the generated first predicted feature, the first loss value between the masked feature in the first feature data and the first predicted feature can be calculated; for the generated second predicted feature, the second loss value between the masked feature in the second feature data and the second predicted feature can be calculated. For example, the first loss value and the second loss value can be calculated according to the cross-entropy loss function.

[0108] When training the multimodal model according to the first loss value and the second loss value, the MOC in the generation module can be trained according to the first loss value, and the MLM in the generation module can be trained according to the second loss value. In addition, the text feature extraction module can also be trained according to the second loss value. Among them, the calculation of the loss value and the training of MOC and MLM belong to the prior art and will not be introduced in detail here.

[0109] Step 502: Based on the first discrimination result and the second discrimination result, train the generation module, the splicing module, and the discrimination module.

[0110] After obtaining the first discrimination result and the second discrimination result, the third loss value corresponding to the first discrimination result can be determined. The third loss value can be calculated according to the first discrimination result and the corresponding first reference information. The first reference information is used to indicate whether each feature of the fifth feature data has been replaced. This first reference information can be obtained after masking the feature data of the sample image. For example, if the sample image corresponds to 5 key image features and it is assumed that the second key image feature is masked, the corresponding first reference information is "01000".

[0111] Similarly, the fourth loss value corresponding to the second discrimination result can be determined. The fourth loss value can be calculated according to the second discrimination result and the corresponding second reference information. The second reference information is used to indicate whether each feature of the sixth feature data has been replaced. This second reference information can be obtained after masking the sample text. For example, if the sample text corresponds to 8 words and it is assumed that the 4th word is masked, the corresponding second reference information is "00010000".

[0112] After obtaining the third loss value, the generation module, the discrimination module, and the splicing module can be trained. After obtaining the fourth loss value, the generation module, the discrimination module, the text feature extraction module, and the splicing module can be trained. Among them, the calculation of the above loss values and the training process of each module belong to the prior art and will not be elaborated here.

[0113] It should be noted that the above steps 501 and 502 both belong to the training process of the multimodal model, and there is no fixed sequence in the execution time series. It can be that after obtaining the first loss value, the second loss value, the third loss value, and the fourth loss value, the corresponding modules in the multimodal model are trained according to the first loss value, the second loss value, the third loss value, and the fourth loss value in sequence.

[0114] In addition, in the feature combination model of the present application, a matching model may also be included, and the matching model may be a fully connected layer (FC layer). Before training the multimodal model based on the first predicted feature, the second predicted feature, and the discrimination result, the combined feature data may also be input into the feature matching model to obtain the matching result of the first sample data and the second sample data. Based on the reference matching value corresponding to the first sample data and the second sample data and the matching result, the fifth loss value is determined, and each module included in the multimodal model is trained according to the fifth loss value.

[0115] After obtaining the combined feature data in step 204, the combined feature data may also be input into the matching model, and the matching model determines whether the fifth feature data and the sixth feature data in the combined feature data match, that is, determines whether the image content in the sample image and the text content in the sample text match, and obtains the matching result. For example, if the sample image is an image of blue sky and white clouds, and the sample text is "white clouds floating in the blue sky", it can be considered that the sample image and the sample text match. After obtaining the matching result, the fifth loss value can be calculated based on the reference matching value corresponding to the first sample data and the second sample data and the matching result. After obtaining the fifth loss value, the corresponding models in the multimodal model can be trained in turn according to the first loss value - the fifth loss value.

[0116] After training the multimodal model multiple times, when it is determined that the decline rate of the first loss value - the fifth loss value tends to be gentle, it can be considered that the training of the multimodal model is completed. When applying the multimodal model, the generation module and the discrimination module in the training process of the multimodal model can be removed, and the remaining modules form the multimodal model in actual application.

[0117] In the embodiment of the present application, after obtaining the predicted feature corresponding to the masked feature in the feature data of different modal sample data, the predicted feature can be used to replace the masked feature in the feature data, and then it is determined whether each feature in the replaced feature data has been replaced to obtain the discrimination result of whether each feature has been replaced. Then, the multimodal model is trained according to the predicted feature and the discrimination result. In this way, when training the multimodal model according to the predicted feature and the discrimination result, in addition to the predicted feature participating in the training process of the multimodal model, the result of whether each feature has been replaced will also participate in the training process of the multimodal model. In this way, whether the feature in the feature data is masked or not will participate in the calculation of the loss value, which can improve the utilization rate of the sample data and thus improve the training efficiency of the multimodal model.

[0118] Figure 6An apparatus for training a multi-modal model provided by an embodiment of the present application. The apparatus may be the terminal or server in the above embodiments. The apparatus includes:

[0119] An acquisition unit 610, configured to acquire first feature data after mask processing corresponding to first sample data, and second feature data after mask processing corresponding to second sample data, where the first sample data and the second sample data belong to data of different modalities;

[0120] A prediction unit 620, configured to perform prediction processing on the masked features in the first feature data to obtain first predicted features, and perform prediction processing on the masked features in the second feature data to obtain second predicted features;

[0121] A replacement unit 630, configured to replace the masked features in the first feature data with the first predicted features to obtain third feature data, and replace the masked features in the second feature data with the second predicted features to obtain fourth feature data

[0122] A discrimination unit 640, configured to perform splicing processing on the third feature data and the fourth feature data to obtain spliced feature data after splicing processing, and perform discrimination processing on each feature in the spliced feature data to obtain a discrimination result on whether each feature in the spliced feature data is replaced;

[0123] A training unit 650, configured to train the multi-modal model based on the first predicted features, the second predicted features, and the discrimination results.

[0124] Optionally, the first sample data is image data, and the second sample data is text data.

[0125] Optionally, the discrimination unit 640 is configured to:

[0126] Perform discrimination processing on each key image feature included in the spliced feature data to obtain a first discrimination result on whether each image feature included in the spliced feature data is replaced;

[0127] Perform discrimination processing on each word feature included in the spliced feature data to obtain a second discrimination result on whether each word feature included in the spliced feature data is replaced;

[0128] The training unit 650 is configured to: train the multi-modal model based on the first predicted features, the second predicted features, the first discrimination result, and the second discrimination result.

[0129] Optionally, the multimodal model includes a generation module, a splicing module, and a discrimination module. The generation module is used to perform the prediction process, the splicing module is used to perform the splicing process, and the discrimination module is used to perform the discrimination process;

[0130] The training unit 650 is used to: train the generation module based on the first prediction feature and the second prediction feature; train the generation module, the splicing module, and the discrimination module based on the first discrimination result and the second discrimination result.

[0131] Optionally, the training unit 650 is used to:

[0132] Determine a first loss value corresponding to the masked feature in the first prediction feature and the first feature data, and determine a second loss value corresponding to the masked feature in the second prediction feature and the second feature data;

[0133] Train the generation module based on the first loss value and the second loss value.

[0134] Optionally, the training unit 650 is used to:

[0135] Determine a third loss value based on the first discrimination result and first reference information, where the first reference information is used to indicate whether each key image feature included in the spliced feature data is replaced;

[0136] Determine a fourth loss value based on the second discrimination result and second reference information, where the second reference information is used to indicate whether each word feature included in the spliced feature data is replaced;

[0137] Train the generation module, the splicing module, and the discrimination module based on the third loss value and the fourth loss value.

[0138] Optionally, the multimodal model further includes a matching module, which is used to perform a matching process based on features of different modalities in the spliced feature data to obtain a matching result of the first sample data and the second sample data;

[0139] The training unit 650 is further used to: determine a fifth loss value based on a reference matching value corresponding to the first sample data and the second sample data and the matching result;

[0140] Train the multimodal model based on the fifth loss value.

[0141] In an embodiment of the present application, after obtaining the predicted features corresponding to the masked features in the feature data of different modality sample data, the predicted features can be used to replace the masked features in the feature data, and then it can be determined whether each feature in the replaced feature data has been replaced to obtain the discrimination result of whether each feature has been replaced. Then, the multimodal model can be trained based on the predicted features and the discrimination result. In this way, when training the multimodal model based on the predicted features and the discrimination result, in addition to the predicted features participating in the training process of the multimodal model, the result of whether each feature has been replaced will also participate in the training process of the multimodal model. In this way, regardless of whether the features in the feature data are masked, they will all participate in the calculation of the loss value, which can improve the utilization rate of the sample data and thus improve the training efficiency of the multimodal model.

[0142] It should be noted that when training the multimodal model by the device for training the multimodal model provided in the above embodiment, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for training the multimodal model provided in the above embodiment and the method embodiment for training the multimodal model belong to the same concept, and the specific implementation process can be found in the method embodiment, which will not be elaborated here.

[0143] Figure 7 The structural block diagram of a computer device 700 provided by an exemplary embodiment of the present application is shown. The computer device 700 may be the terminal in the above embodiment, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The computer device 700 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0144] Generally, the computer device 700 includes a processor 701 and a memory 702.

[0145] The processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (digital signal processing), FPGA (field-programmable gate array), and PLA (programmable logic array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (central processing unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (graphics processing unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may further include an AI (artificial intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0146] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 701 to implement a method for training a multi-modal model provided in the method embodiments of the present application.

[0147] In some embodiments, the computer device 700 may further optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, a positioning assembly 708, and a power supply 709.

[0148] The peripheral device interface 703 can be used to connect at least one I / O (input / output) related peripheral device to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0149] The radio frequency circuit 704 is used to receive and transmit RF (radio frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (wireless fidelity) network. In some embodiments, the radio frequency circuit 704 may further include a circuit related to NFC (near field communication), and this application does not limit this.

[0150] The display screen 705 is used to display the UI (user interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input as control signals to the processor 701 for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 705, which is provided on the front panel of the computer device 700; in other embodiments, there may be at least two display screens 705, which are respectively provided on different surfaces of the computer device 700 or are in a folded design; in other embodiments, the display screen 705 may be a flexible display screen, which is provided on the curved surface or the folding surface of the computer device 700. Even, the display screen 705 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared using materials such as LCD (liquid crystal display) and OLED (organic light-emitting diode).

[0151] The camera module 706 is used to capture images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (virtual reality) shooting function or other fused shooting functions. In some embodiments, the camera module 706 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0152] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 701 for processing, or input to the radio frequency circuit 704 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the computer device 700. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 707 may further include a headphone jack.

[0153] The positioning component 708 is used to locate the current geographical location of the computer device 700 to implement navigation or LBS (location based service). The positioning component 708 may be a positioning component based on the GPS (global positioning system) of the United States, the Beidou system of China, or the Galileo system of Russia.

[0154] The power supply 709 is used to supply power to each component in the computer device 700. The power supply 709 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 709 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0155] In some embodiments, the computer device 700 further includes one or more sensors 710. The one or more sensors 710 include but are not limited to: an acceleration sensor 711, a gyroscope sensor 712, a pressure sensor 713, a fingerprint sensor 714, an optical sensor 715, and a proximity sensor 716.

[0156] The acceleration sensor 711 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the computer device 700. For example, the acceleration sensor 711 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 701 can control the display screen 705 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 711. The acceleration sensor 711 can also be used for collecting game or user's motion data.

[0157] The gyroscope sensor 712 can detect the body orientation and rotation angle of the computer device 700. The gyroscope sensor 712 can cooperate with the acceleration sensor 711 to collect the 3D actions of the user on the computer device 700. Based on the data collected by the gyroscope sensor 712, the processor 701 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0158] The pressure sensor 713 can be disposed on the side frame of the computer device 700 and / or the lower layer of the display screen 705. When the pressure sensor 713 is disposed on the side frame of the computer device 700, it can detect the holding signal of the user on the computer device 700, and the processor 701 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 713. When the pressure sensor 713 is disposed on the lower layer of the display screen 705, the processor 701 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 705. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0159] The fingerprint sensor 714 is used to collect the fingerprints of the user. The processor 701 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 714, or the fingerprint sensor 714 can identify the user's identity according to the collected fingerprints. When the identity of the user is identified as a trusted identity, the processor 701 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 714 can be disposed on the front, back, or side of the computer device 700. When there are physical buttons or manufacturer logos on the computer device 700, the fingerprint sensor 714 can be integrated with the physical buttons or manufacturer logos.

[0160] The optical sensor 715 is used to collect the ambient light intensity. In one embodiment, the processor 701 can control the display brightness of the display screen 705 according to the ambient light intensity collected by the optical sensor 715. Specifically, when the ambient light intensity is high, the display brightness of the display screen 705 is increased; when the ambient light intensity is low, the display brightness of the display screen 705 is decreased. In another embodiment, the processor 701 can also dynamically adjust the shooting parameters of the camera module 706 according to the ambient light intensity collected by the optical sensor 715.

[0161] A proximity sensor 716, also known as a distance sensor, is typically disposed on the front panel of the computer device 700. The proximity sensor 716 is used to collect the distance between the user and the front of the computer device 700. In one embodiment, when the proximity sensor 716 detects that the distance between the user and the front of the computer device 700 is gradually decreasing, the processor 701 controls the display screen 705 to switch from the lit state to the off state; when the proximity sensor 716 detects that the distance between the user and the front of the computer device 700 is gradually increasing, the processor 701 controls the display screen 705 to switch from the off state to the lit state.

[0162] Those skilled in the art can understand that Figure 7 the structure shown in does not constitute a limitation on the computer device 700, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0163] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the method of training a multi-modal model in the above embodiment. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (read-only memory), a RAM (random access memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0164] In an exemplary embodiment, a computer program product is also provided. The computer program product includes at least one instruction, and the at least one instruction can be loaded and executed by a processor to implement the operations performed by the method of training a multi-modal model as described in the above embodiment.

[0165] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0166] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between a user terminal and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the sample data involved in this application are all obtained under full authorization.

[0167] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A method for training a multi-modal model, characterized in that, The method includes: Obtaining first feature data after mask processing corresponding to first sample data, and second feature data after mask processing corresponding to second sample data, where the first sample data and the second sample data belong to data of different modalities; performing prediction processing on the masked features in the first feature data to obtain first predicted features, and performing prediction processing on the masked features in the second feature data to obtain second predicted features; Replacing the masked features in the first feature data with the first predicted features to obtain third feature data, and replacing the masked features in the second feature data with the second predicted features to obtain fourth feature data; Performing splicing processing on the third feature data and the fourth feature data to obtain spliced feature data after splicing processing, and performing discrimination processing on each feature in the spliced feature data to obtain a discrimination result on whether each feature in the spliced feature data is replaced; Training a multimodal model based on the first predicted features, the second predicted features, and the discrimination result; The first sample data is image data, and the second sample data is text data; The performing discrimination processing on each feature in the spliced feature data to obtain a discrimination result on whether each feature in the spliced feature data is replaced includes: Performing discrimination processing on each key image feature included in the spliced feature data to obtain a first discrimination result on whether each image feature included in the spliced feature data is replaced; Performing discrimination processing on each word feature included in the spliced feature data to obtain a second discrimination result on whether each word feature included in the spliced feature data is replaced; The training the multimodal model based on the first predicted features, the second predicted features, and the discrimination result includes: Training the multimodal model based on the first predicted features, the second predicted features, the first discrimination result, and the second discrimination result.

2. The method according to claim 1, wherein The multimodal model includes a generation module, a splicing module, and a discrimination module. The generation module is used to perform the prediction processing, the splicing module is used to perform the splicing processing, and the discrimination module is used to perform the discrimination processing; The training the multimodal model based on the first predicted features, the second predicted features, the first discrimination result, and the second discrimination result includes: Training the generation module based on the first predicted features and the second predicted features; Training the generation module, the splicing module, and the discrimination module based on the first discrimination result and the second discrimination result.

3. The method according to claim 2, wherein The training the generation module based on the first predicted features and the second predicted features includes: Determining a first loss value corresponding to the first predicted features and the masked features in the first feature data, and determining a second loss value corresponding to the second predicted features and the masked features in the second feature data; Training the generation module based on the first loss value and the second loss value.

4. The method according to claim 3, characterized in that Training the generation module, the splicing module, and the discrimination module based on the first discrimination result and the second discrimination result includes: Determining a third loss value based on the first discrimination result and first reference information, where the first reference information is used to indicate whether each key image feature included in the spliced feature data is replaced; Determining a fourth loss value based on the second discrimination result and second reference information, where the second reference information is used to indicate whether each word feature included in the spliced feature data is replaced; Training the generation module, the splicing module, and the discrimination module based on the third loss value and the fourth loss value.

5. The method according to claim 1, wherein The multimodal model further includes a matching module for performing a matching process based on features of different modalities in the spliced feature data to obtain a matching result between the first sample data and the second sample data; Determining a fifth loss value based on a reference matching value corresponding to the first sample data and the second sample data and the matching result; Training the multimodal model based on the fifth loss value.

6. An apparatus for training a multi-modal model, characterized in that, The apparatus includes: An acquisition unit for acquiring first feature data after mask processing corresponding to first sample data and second feature data after mask processing corresponding to second sample data, where the first sample data and the second sample data belong to data of different modalities; A prediction unit for performing a prediction process on the masked features in the first feature data to obtain first predicted features, and performing a prediction process on the masked features in the second feature data to obtain second predicted features; A replacement unit for replacing the masked features in the first feature data with the first predicted features to obtain third feature data, and replacing the masked features in the second feature data with the second predicted features to obtain fourth feature data; A discrimination unit for performing a splicing process on the third feature data and the fourth feature data to obtain spliced feature data after the splicing process, and performing a discrimination process on each feature in the spliced feature data to obtain a discrimination result indicating whether each feature in the spliced feature data is replaced; A training unit for training a multimodal model based on the first predicted features, the second predicted features, and the discrimination result; The first sample data is image data, and the second sample data is text data; The discrimination unit is used for: Performing a discrimination process on each key image feature included in the spliced feature data to obtain a first discrimination result indicating whether each image feature included in the spliced feature data is replaced; Performing a discrimination process on each word feature included in the spliced feature data to obtain a second discrimination result indicating whether each word feature included in the spliced feature data is replaced; The training unit is used for: Training the multimodal model based on the first predicted features, the second predicted features, the first discrimination result, and the second discrimination result.

7. The device according to claim 6, characterized in that The multimodal model includes a generation module, a splicing module, and a discrimination module. The generation module is used to perform the prediction process, the splicing module is used to perform the splicing process, and the discrimination module is used to perform the discrimination process; The training unit is configured to: train the generation module based on the first prediction feature and the second prediction feature; Based on the first discrimination result and the second discrimination result, train the generation module, the splicing module, and the discrimination module.

8. The device according to claim 7, characterized in that, The training unit is configured to: Determine a first loss value corresponding to the masked feature in the first feature data for the first prediction feature, and determine a second loss value corresponding to the masked feature in the second feature data for the second prediction feature; train the generation module based on the first loss value and the second loss value; The training unit is configured to: Determine a third loss value based on the first discrimination result and first reference information, where the first reference information is used to indicate whether each key image feature included in the spliced feature data is replaced; determine a fourth loss value based on the second discrimination result and second reference information, where the second reference information is used to indicate whether each word feature included in the spliced feature data is replaced; based on the third loss value and the fourth loss value, train the generation module, the splicing module, and the discrimination module.

9. The device according to claim 8, characterized in that, The multimodal model further includes a matching module, configured to perform a matching process based on features of different modalities in the spliced feature data to obtain a matching result of the first sample data and the second sample data; The training unit is further configured to: determine a fifth loss value based on the reference matching value corresponding to the first sample data and the second sample data and the matching result; Based on the fifth loss value, train the multimodal model.

10. A computer device, characterized in that, The computer device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the operations performed by the method for training a multimodal model according to any one of claims 1 to 5.

11. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the method for training a multimodal model according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Cross-modal data processing method and device, storage medium and electronic device

    CN112199462A

  • Visual language model obtaining method and device, visual language task processing method and device, equipment and storage medium

    CN113792113A