Data matching method and apparatus, and multi-modal model processing method and apparatus

By introducing multimodal model processing methods into the vision-language pretrained model, the shortcomings in data processing performance of traditional VLP models are solved, and more efficient multimodal information fusion and better image-text matching processing are achieved.

WO2025130344A1PCT designated stage expired Publication Date: 2025-06-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/127423
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-10-25
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The traditional vision-language pre-training model (VLP) has room for improvement in data processing performance, especially in terms of time-consuming training processes, ineffective modal embedding information fusion, and noise processing of training data.

Method used

A multimodal model processing method is proposed. By obtaining training sample pairs of different modalities, encoding and fusion using the initial model, determining prediction weights, and data fusion and decoding are performed based on these weights to improve the efficiency and performance of the model.

Benefits of technology

This method can improve the efficiency of the VLP model, optimize multimodal information fusion, improve the model's processing performance on image-text mismatch problems, and thus improve the processing performance of the entire model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127423_26062025_PF_FP_ABST
    Figure CN2024127423_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A data matching method, which is executed by a computer device. The method comprises: acquiring first data, and encoding the first data, so as to obtain a first data representation; acquiring second data, and encoding the second data, so as to obtain a second data representation, wherein the first data and the second data are data of different modalities (502); combining the first data representation and the second data representation to obtain a mixed data representation, and determining a fusion weight on the basis of the mixed data representation (504); on the basis of the fusion weight, fusing the first data representation and the mixed data representation to obtain a first fused representation, and fusing the second data representation and the mixed data representation to obtain a second fused representation (506); and on the basis of the first fused representation and the second fused representation, determining the degree of matching between the first data and the second data (508).
Need to check novelty before this filing date? Find Prior Art

Description

Data matching method and device, multimodal model processing method and device

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2023117771941, filed on December 21, 2023, entitled “Data matching method and device, multimodal model processing method and device,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to a data matching method, apparatus, computer equipment, storage medium and computer program product, as well as a multimodal model processing method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0004] With the development of computer technology, machine learning has emerged. By building and training machine learning models, these models can be equipped to perform tasks. In the field of vision-language learning, to improve the practicality of the model, a general vision-language pre-training model, also known as a VLP (Vision-Language Pre-training) model, is often pre-trained for image and text processing.

[0005] Traditional VLP models typically use a dual encoder structure, with parallel encoders, one for image processing and the other for text processing. The performance of this traditional VLP model for data processing needs to be further improved.

[0006] Summary of the Invention

[0007] The present application provides a data matching method, apparatus, computer device, computer-readable storage medium, and computer program product.

[0008] In one aspect, the present application provides a data matching method, performed by a computer device, comprising:

[0009] Acquire first data, encode the first data, and obtain a first data representation;

[0010] Acquire second data, encode the second data, and obtain a second data representation, wherein the first data and the second data are data of different modalities;

[0011] combining the first data representation and the second data representation to obtain a hybrid data representation, and determining a fusion weight according to the hybrid data representation;

[0012] According to the fusion weight, fusing the first data representation and the mixed data representation to obtain a first fused representation, and fusing the second data representation and the mixed data representation to obtain a second fused representation; and

[0013] A degree of matching between the first data and the second data is determined according to the first fused representation and the second fused representation.

[0014] On the other hand, the present application also provides a data matching device, comprising:

[0015] an encoding module, configured to obtain first data, encode the first data to obtain a first data representation; and obtain second data, encode the second data to obtain a second data representation, wherein the first data and the second data are data of different modalities;

[0016] a determination module, configured to combine the first data representation and the second data representation to obtain a mixed data representation, and determine a fusion weight according to the mixed data representation;

[0017] a fusion module, configured to fuse the first data representation and the mixed data representation to obtain a first fused representation, and to fuse the second data representation and the mixed data representation to obtain a second fused representation, according to the fusion weight; and

[0018] A matching module is used to determine a matching degree between the first data and the second data according to the first fusion representation and the second fusion representation.

[0019] In another aspect, the present application provides a multimodal model processing method, performed by a computer device, the method comprising:

[0020] Acquire a training sample pair, the training sample pair comprising first sample data and second sample data; the first sample data and the second sample data are sample data of different modalities;

[0021] Encoding the first sample data using an initial model to be trained to obtain a first sample data representation; Encoding the second sample data to obtain a second sample data representation;

[0022] obtaining a mixed sample data representation by combining the first sample data representation and the second sample data representation through the initial model, and determining a prediction weight according to the mixed sample data representation;

[0023] fusing the first sample data representation and the mixed sample data representation according to the prediction weights using the initial model, and performing decoding based on the fused result to obtain a first prediction feature representation;

[0024] fusing the second sample data representation and the mixed sample data representation according to the prediction weights using the initial model, and performing decoding based on the fused result to obtain a second prediction feature representation;

[0025] Determining a first loss based on the true feature representation of the first sample data and the first predicted feature representation, and the true feature representation of the second sample data and the second predicted feature representation;

[0026] determining a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weight; and

[0027] A target loss function is constructed according to the first loss and the second loss, and the initial model to be trained is trained based on the target loss function, so as to obtain a multimodal model after the training is completed.

[0028] On the other hand, the present application also provides a multimodal model processing device, comprising:

[0029] an encoding module configured to obtain a training sample pair, the training sample pair comprising first sample data and second sample data; the first sample data and the second sample data being sample data of different modalities; encoding the first sample data using an initial model to be trained to obtain a first sample data representation; and encoding the second sample data to obtain a second sample data representation;

[0030] a determination module, configured to obtain a mixed sample data representation by combining the first sample data representation and the second sample data representation through the initial model, and determine a prediction weight according to the mixed sample data representation;

[0031] a fusion module, configured to fuse the first sample data representation and the mixed sample data representation using the initial model according to the prediction weight, and perform decoding based on the fused result to obtain a first prediction feature representation;

[0032] The fusion module is further configured to fuse the second sample data representation and the mixed sample data representation using the initial model according to the prediction weight, and perform decoding based on the fused result to obtain a second prediction feature representation;

[0033] A construction module is configured to determine a first loss based on the true feature representation of the first sample data and the first predicted feature representation, and the true feature representation of the second sample data and the second predicted feature representation; determine a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weight; and construct a target loss function based on the first loss and the second loss; and

[0034] A training module is used to train the initial model to be trained based on the target loss function, and obtain a multimodal model after the training is completed.

[0035] On the other hand, the present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0036] On the other hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0037] On the other hand, the present application also provides a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.

[0038] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] FIG1 is a diagram illustrating an application environment of a data matching method and / or a multimodal model processing method in one embodiment;

[0041] FIG2 is a schematic flow chart of a multimodal model processing method according to an embodiment;

[0042] FIG3 is a schematic diagram showing the principle of mask processing in one embodiment;

[0043] FIG4 is a schematic flow chart of a multimodal model processing method according to another embodiment;

[0044] FIG5 is a schematic diagram of a flow chart of a data matching method in one embodiment;

[0045] FIG6 is a schematic flow chart of a data matching method in another embodiment;

[0046] FIG7 is a block diagram of a multimodal model according to an embodiment;

[0047] FIG8 is a structural block diagram of a data matching device in one embodiment;

[0048] FIG9 is a structural block diagram of a multimodal model processing device according to one embodiment;

[0049] FIG10 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] The data matching method and / or multimodal model processing method provided in the embodiment of the present application can be applied to the application environment shown in Figure 1. Among them, the terminal 102 communicates with the server 104 via a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal and the server can respectively execute the data matching method and / or multimodal model processing method provided in the embodiment of the present application separately, or they can collaboratively execute the data matching method and / or multimodal model processing method provided in the embodiment of the present application. Take the server executing the data matching method provided in the present application alone as an example for explanation: the server receives the first data and the second data sent by the terminal, and encodes the first data to obtain a first data representation, and encodes the second data to obtain a second data representation, the first data and the second data are data of different modalities; the first data representation and the second data representation are combined to obtain a mixed data representation, and the fusion weight is determined according to the mixed data representation; according to the fusion weight, the first data representation and the mixed data representation are fused to obtain a first fused representation, and the second data representation and the mixed data representation are fused to obtain a second fused representation; according to the first fused representation and the second fused representation, the matching degree of the first data and the second data is determined. The server feeds back the matching degree to the terminal.

[0052] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and this application does not limit this.

[0053] The multimodal model mentioned in the embodiments of the present application can specifically be a pre-training model. Among them, the pre-training model (Pre-training model), also known as the cornerstone model or the large model, refers to a deep neural network (DNN) with a large number of parameters, which is trained on a large amount of unlabeled data, and the function approximation ability of the large-parameter DNN is used to enable the PTM to extract common features from the data, and after fine-tuning (fine tune), parameter efficient fine-tuning (PEFT), prompt-tuning and other technologies, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in a few-shot or zero-shot scenario. PTM can be divided into language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multimodal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modality processed, wherein the multimodal model refers to a model that establishes two or more data modality feature representations. Pre-trained models are important tools for outputting artificial intelligence generated content (AIGC) and can also serve as a universal interface for connecting multiple specific task models.

[0054] The solutions provided in the embodiments of this application involve artificial intelligence machine learning, large models and other technologies, which are specifically introduced in detail through the following embodiments.

[0055] Before introducing the data matching method and / or multimodal model processing method provided in this application, the relevant English abbreviations are explained:

[0056] VLP: Vision-Language Pre-training, vision-language pre-training.

[0057] GIAE: Gated Interactive Masked AutoEncoders, gated interactive masked autoencoders.

[0058] GIM: Gated Interactive Mechanism, gated interactive mechanism.

[0059] MAE: Masked AutoEncoders, masked autoencoder.

[0060] Image-text mismatching issue: Image-text mismatching issue.

[0061] Transformer: A neural network architecture based on the attention mechanism that can effectively encode temporal information in the encoder part. It is widely used in natural language processing, computer vision, machine translation, speech recognition and other fields.

[0062] LightVLP: a lightweight Vision-Language Pre-training framework, a lightweight vision-language pre-training model framework, which is also the multimodal model proposed in this application.

[0063] In the field of vision-language learning, to improve model practicality, a general vision-language pre-training model, also known as a VLP (Vision-Language Pre-training) model, is typically pre-trained for image and text processing. Furthermore, in actual business processing, to adapt to diverse downstream tasks, the VLP model can be retrained to obtain a task model suitable for specific business scenarios. Therefore, training a high-performance VLP model is a highly practical project.

[0064] Traditional VLP methods often face the following challenges: (1) The training process is very time-consuming because the text and image inputs are often very long; (2) The embedding information of different modalities is not effectively fused; (3) In the real world, many training image-text alignment data contain a lot of noise and need further processing.

[0065] In view of this, this application proposes a multimodal model processing method that can solve the above-mentioned technical problems, optimize the multimodal information fusion method on the basis of improving the efficiency of the VLP model, and obtain better results. In addition, the multimodal model of this application can better handle the modal data mismatch problem in the training data set during the training process, such as the image-text mismatch problem, thereby improving the processing performance of the model.

[0066] The following is an introduction to the training method of the multimodal model involved in this application:

[0067] In an exemplary embodiment, as shown in FIG2 , a multimodal model processing method is provided, which is described by taking the method applied to a computer device (such as the terminal or server in FIG1 ) as an example, and includes the following steps 202 to 214 . Among them:

[0068] Step 202: Obtain a training sample pair, the training sample pair including first sample data and second sample data; the first sample data and the second sample data are sample data of different modalities; encode based on the first sample data through the initial model to be trained to obtain a first sample data representation; encode based on the second sample data to obtain a second sample data representation.

[0069] The initial model is a pre-built neural network model, which becomes a multimodal model after training. The initial model in this application may include two encoders and decoders with a symmetrical relationship, two decoders with a symmetrical relationship, and a cross transformer for processing interactive information.

[0070] The training sample pair is a pre-constructed data pair for model training. The training sample pair includes first sample data and second sample data. The first sample data is data of the first modality, and the second sample data is data of the second modality, and the first modality is different from the second modality. Among them, the modality is a way of carrying information in a perceptual channel, such as visual, auditory, olfactory or tactile channels. The modality in this application can be at least one of a visual modality, a text modality, or an audio modality. For example, the first sample data can be image data, and the second sample data can be text data; or, the first sample data can be image data, and the second sample data can be audio data, etc., which is not limited in the embodiments of this application.

[0071] It should be noted that, for ease of understanding, the data used and generated in the model training phase are distinguished from the data used and generated in the model application phase. The word "sample" is included in the relevant data names, such as the first sample data, the second sample data, etc.

[0072] It should be understood that the terms "first," "second," and the like used herein may be used to describe various elements, but these elements are not limited by these terms. These terms are merely used to distinguish a first element from another element. For example, first sample data may be referred to as second sample data, and similarly, first data may be referred to as second data without departing from the scope of this application.

[0073] Encoding is the process of mapping plaintext information into a latent space through calculations. This application can encode the first sample data using a first encoder in the initial model to be trained to obtain a first sample data representation; and encode the second sample data using a second encoder in the initial model to obtain a second sample data representation.

[0074] In some embodiments, the computer device may directly encode the first sample data to obtain the first sample data representation, and directly encode the second sample data to obtain the second sample data representation. In other embodiments, the computer device may mask the first sample data, then encode the masked first sample data using the first encoder in the initial model to obtain the first sample data representation. The computer device may mask the second sample data, then encode the masked second sample data using the second encoder in the initial model to obtain the second sample data representation.

[0075] For ease of description, the data input to the first encoder is referred to as the first input sequence, and the data input to the second encoder is referred to as the second input sequence. That is, the first input sequence can be the first sample data or the masked first sample data; the second input sequence can be the second sample data or the masked second sample data.

[0076] In some embodiments, the initial model to be trained is encoded multiple times for the first input sequence to obtain a first sample data representation. The initial model to be trained is encoded multiple times for the second input sequence to obtain a second sample data representation.

[0077] In some embodiments, the first encoder and the second encoder have the same network structure, so the encoding process performed on the first input sequence is the same as the encoding process performed on the second input sequence. Of course, in other embodiments, the first encoder and the second encoder can have different network structures, and the encoding process performed on the first input sequence is different from the encoding process performed on the second input sequence. This embodiment of the present application is not limited to this.

[0078] Taking the example of the first encoder and the second encoder having the same network structure, the following is explained: for any encoding, such as the l+1th encoding, the output of the lth encoding can be obtained first. A multi-head self-attention layer operation can be performed based on the output of the lth encoding. The result of the multi-head self-attention layer operation and the output of the lth encoding are combined and then layer regularization is performed to obtain the intermediate processing features in the encoding process; the intermediate processing features in the encoding process are processed by a forward neural network, and the result of the forward neural network processing is combined with the intermediate processing features in the encoding process and then layer regularization is performed to obtain the output of the l+1th encoding; l+1 is used as the new l, and the output of the lth encoding is returned to continue execution until the last encoding process is completed, and the output of the last encoding is used as the first sample data representation or the second sample data representation. Where l is a natural number greater than or equal to 0. When l is 0, the output of the lth encoding obtained is the first input sequence or the second input sequence.

[0079] In some embodiments, the computer device may encode each input data in the first input sequence (when the first input sequence includes multiple data blocks, each input data is referred to as each data block) multiple times to obtain an output corresponding to each input data, and then concatenate the outputs corresponding to each input data in the last encoding to obtain the first sample data representation. Similarly, the computer device may encode each input data in the second input sequence multiple times to obtain an output corresponding to each input data, and then concatenate the outputs corresponding to each input data in the last encoding to obtain the second sample data representation.

[0080] For example, the computer device may express the output of the (l+1)th encoding corresponding to any input data by Formula 1:

[0081] Among them, H l represents the output of the lth encoding, which can also be considered as the hidden feature of the lth layer; f LN Representation layer regularization operation, Indicates the multi-head self-attention layer operation through the encoder layer l network, Indicates processing by the forward neural network of the encoder layer l (specifically, it can be a two-layer fully connected network); It is the intermediate processing feature in the encoding process. The encoder can be the first encoder or the second encoder; H l+1 It represents the output of the l+1th encoding, and can also be considered as the hidden feature of the l+1th layer.

[0082] Furthermore, the computer device may express the first sample data by the following formula 2:

[0083] Where x1 represents the first data block in the first input sequence, x2 represents the second data block in the first input sequence, and x M represents the Mth data block (i.e. the last data block) in the first input sequence; represents the latent features corresponding to the first data block output by the last encoding layer of the first encoder, represents the latent features corresponding to the second data block output by the last encoding layer of the first encoder, represents the latent features corresponding to the Mth data block output by the last encoding layer of the first encoder, Represents the first sample data representation.

[0084] Furthermore, the computer device may express the second sample data using the following formula 3:

[0085] Where w1 represents the first data block in the second input sequence, w2 represents the second data block in the second input sequence, and w N represents the Nth data block (i.e., the last data block) in the second input sequence; represents the latent features corresponding to the first data block output by the last encoding layer of the second encoder, represents the latent features corresponding to the second data block output by the last encoding layer of the second encoder, represents the latent features corresponding to the Nth data block output by the last encoding layer of the second encoder, Represents the second sample data representation.

[0086] In some embodiments, the first encoder and / or the second encoder are composed of a recurrent neural network, a convolutional neural network, or a graph neural network. In practical applications, more or fewer network layers can be designed according to actual needs, and this is not limited in the embodiments of the present application. In some embodiments, the first encoder and / or the second encoder can be a Transformers structure.

[0087] Step 204 : obtaining a mixed sample data representation by combining the first sample data representation and the second sample data representation through the initial model, and determining a prediction weight according to the mixed sample data representation.

[0088] It should be noted that when the alignment quality of the first sample data and the second sample data in a training sample pair is high, the information of the mixed modality can be considered more during the single-modal decoding; when the first sample data and the second sample data have relatively large noise (such as image-text mismatch), the information of this single modality can be considered more during the single-modal decoding. Based on this thinking, the present application designs a gated interaction mechanism to control the ratio of reference single data representation and mixed data representation during decoding, that is, to dynamically adjust the reference ratio through fusion weights. The fusion weights can be called gating weights.

[0089] Specifically, the initial model may fuse the first sample data representation and the second sample data representation through a cross transformer structure to obtain a mixed sample data representation, and then perform a numerical transformation on the mixed modality sample representation to obtain a prediction weight.

[0090] In some embodiments, the computer device may directly perform addition or multiplication operations on the first sample data representation and the second sample modality to obtain a mixed sample data representation. In other embodiments, the computer device may also perform an attention operation on the first sample data representation and the second sample data representation to obtain a mixed sample data representation. In other embodiments, the computer device may also use more complex operations, such as performing an attention mechanism processing, and then performing matrix addition or multiplication, and then performing linear / nonlinear transformations to obtain a mixed sample data representation. This application does not limit the specific fusion method.

[0091] Furthermore, after the initial model combines the first sample data representation and the second sample data representation to obtain a mixed sample data representation, it can perform a linear transformation on the mixed sample data representation and then perform activation processing to obtain prediction weights. In other words, the mixed sample modality is converted into a numerical value, which is also the prediction weight.

[0092] Step 206 : The first sample data representation and the mixed sample data representation are fused according to the prediction weights through the initial model, and decoding is performed based on the fused result to obtain a first prediction feature representation.

[0093] Specifically, the initial model may perform weighted processing on the first sample data representation and the mixed sample data representation according to the prediction weight, and perform decoding based on the result obtained by the weighted processing to obtain the first prediction feature representation.

[0094] In some embodiments, the computer device may use the predicted weight as the weight of any one of the first sample data representation and the mixed sample data representation, and perform weighted sum processing on the first sample data representation and the mixed sample data representation.

[0095] In some embodiments, the computer device may use the predicted weight as the weight of any one of the first sample data representation and the mixed sample data representation, and use the difference between the value one and the predicted weight as the weight of the other representation, and perform weighted summation processing on the first sample data representation and the mixed sample data representation.

[0096] In some embodiments, the first sample data representation and the mixed sample data representation are fused according to the predicted weight, including: using the predicted weight as the coefficient of the mixed sample data representation, and using the difference between the value one and the predicted weight as the coefficient of the first sample data representation; and performing weighted processing on the mixed sample data representation and the first sample data representation according to the coefficient of the mixed sample data representation and the coefficient of the first sample data representation to obtain a first fused sample representation.

[0097] Specifically, the initial model may perform a weighted summation of the mixed sample data representation and the first sample data representation according to the coefficients of the mixed sample data representation and the coefficients of the first sample data representation to obtain a first fused sample representation. For example, the computer device may calculate the first fused sample representation using the following formula 4:

[0098] Among them, p is the prediction weight, is the first sample data representation, S is the mixed sample data representation, and O1 is the first fusion sample representation.

[0099] In the above embodiment, by fusing the mixed sample data representation and the first sample modality through prediction weights, the ratio of the single data representation and the mixed data representation can be flexibly referenced according to the alignment quality of the two modal data, so that the fused representation can better express the information of the single modal data.

[0100] Furthermore, the initial model may be decoded based on the first fused sample representation to obtain a first prediction feature representation. The specific decoding method may be a traditional decoding method.

[0101] In some embodiments, the initial model encodes the masked data of the first sample data, and some data blocks are randomly discarded during the masking process of the first sample data. Therefore, when decoding, the mask embedding can be inserted into the first fused sample representation according to the position of the discarded data block in the first sample data, wherein the mask embedding can be a preset feature representation, such as a feature representation composed of preset values ​​(preset values ​​such as 0 or 1). Then, the first decoder in the initial model can decode the first fused sample representation after the mask embedding is inserted to obtain a first predicted feature representation. The first predicted feature representation represents the latent features of the discarded data block.

[0102] Step 208: The second sample data representation and the mixed sample data representation are fused according to the prediction weights through the initial model, and decoding is performed based on the fused result to obtain a second prediction feature representation.

[0103] Specifically, the initial model may perform weighted processing on the second sample data representation and the mixed sample data representation according to the prediction weight, and decode based on the result obtained by the weighted processing to obtain the second prediction feature representation.

[0104] In some embodiments, the computer device may use the predicted weight as the weight of any one of the second sample data representation and the mixed sample data representation, and perform weighted sum processing on the second sample data representation and the mixed sample data representation.

[0105] In some embodiments, the computer device may use the predicted weight as the weight of any one of the second sample data representation and the mixed sample data representation, and use the difference between the value one and the predicted weight as the weight of the other representation, and perform weighted summation processing on the second sample data representation and the mixed sample data representation.

[0106] In some embodiments, the second sample data representation and the mixed sample data representation are fused according to the predicted weight, including: using the predicted weight as the coefficient of the mixed sample data representation, and using the difference between the value one and the predicted weight as the coefficient of the second sample data representation; and performing weighted processing on the mixed sample data representation and the second sample data representation according to the coefficient of the mixed sample data representation and the coefficient of the second sample data representation to obtain a second fused sample representation.

[0107] Specifically, the initial model may perform a weighted summation of the mixed sample data representation and the second sample data representation according to the coefficients of the mixed sample data representation and the coefficients of the second sample data representation to obtain a second fused sample representation. For example, the computer device may calculate the second fused sample representation using the following formula 5:

[0108] Among them, p is the prediction weight, is the second sample data representation, S is the mixed sample data representation, and O2 is the second fused sample representation.

[0109] In the above embodiment, by fusing the mixed sample data representation and the second sample modality with the prediction weight, the ratio of the single data representation and the mixed data representation can be flexibly referred to according to the alignment quality of the two modal data, so that the fused representation can better express the information of the single modal data.

[0110] Furthermore, the initial model can be decoded based on the second fused sample representation to obtain a second predicted feature representation. The specific decoding method can be a traditional decoding method.

[0111] In some embodiments, the initial model encodes the masked data of the second sample data, and some data blocks are randomly discarded during the masking process of the second sample data. Therefore, during decoding, the mask embedding can be inserted into the second fused sample representation according to the position of the discarded data block in the second sample data, wherein the mask embedding can be a preset feature representation, such as a feature representation composed of preset values ​​(preset values ​​such as 0 or 1). Then, the second decoder in the initial model can decode the second fused sample representation after the mask embedding is inserted to obtain a second predicted feature representation. The second predicted feature representation represents the latent features of the discarded data block.

[0112] Step 210 : Determine a first loss based on the true feature representation and the first predicted feature representation of the first sample data, and the true feature representation and the second predicted feature representation of the second sample data.

[0113] Specifically, the computer device may extract the true feature representation of the first sample data, and then calculate the difference between the true feature representation of the first sample data and the first predicted feature representation, using this difference as the first modal loss. The computer device may also extract the true feature representation of the second sample data, and then calculate the difference between the true feature representation of the second sample data and the second predicted feature representation, using this difference as the second modal loss. The first loss is determined based on the first modal loss and the second modal loss. The difference may be represented by a difference, a comparison value, or a quotient, etc.

[0114] In some embodiments, the true feature representation of the first sample data involved in the loss calculation may specifically be the true feature representation of the partial data blocks discarded from the first sample data before encoding; the true feature representation of the second sample data involved in the loss calculation may specifically be the true feature representation of the partial data blocks discarded from the second sample data before encoding.

[0115] In other embodiments, the true feature representation of the first sample data involved in the loss calculation may also be the true feature representation of all data blocks in the first sample data; the true feature representation of the second sample data involved in the loss calculation may also be the true feature representation of all data blocks in the second sample data.

[0116] In some embodiments, the computer device may extract the true feature representation of the first sample data and the true feature representation of the second sample data through other trained codecs.

[0117] In some embodiments, when the first sample data is image data, the difference between the true feature representation of the first sample data and the first predicted feature representation can be measured by MSE (Mean Square Error) loss; when the second sample data is text data, the difference between the true feature representation of the second sample data and the second predicted feature representation can be measured by CE (Cross-Entropy) loss.

[0118] In some embodiments, the computer device may use the mean of the differences between the true feature representations of multiple first sample data and the first predicted feature representation as the first modal loss; and use the mean of the differences between the true feature representations of multiple second sample data and the second predicted feature representation as the second modal loss.

[0119] Exemplarily, the computer device may calculate the first modal loss by the following formula:

[0120] Among them, M x represents the set consisting of the first sample data in the training sample set, |M x | is the number of first sample data in the set consisting of first sample data, f MSE represents the average error function, x m Represents the first prediction feature representation of the mth first sample data, x′ m Represents the true feature representation of the mth first sample data, L MIR is the first mode loss.

[0121] Exemplarily, the computer device may calculate the second modal loss by the following formula:

[0122] Among them, M w represents the set consisting of the second sample data in the training sample set, |M w | is the number of second sample data in the set consisting of second sample data, f CE represents the cross entropy function, y m Represents the second prediction feature representation of the mth second sample data, y′ m Represents the true feature representation of the mth second sample data, L MTR is the second mode loss.

[0123] In some embodiments, the computer device may use the sum of the first modal loss and the second modal loss as the first loss, or use the average of the first modal loss and the second modal loss as the first loss, or use the product of the first modal loss and the second modal loss as the first loss, etc. The embodiments of the present application are not limited to this.

[0124] Step 212: Determine a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weights.

[0125] The alignment tag indicates whether the first sample data and the second sample data are aligned. Alignment indicates that the content of the first and second sample data matches. Misalignment indicates that the content of the first and second sample data does not match. For example, if the first sample data is image data and the second sample data is text data, alignment indicates that the image and text match, while misalignment indicates that the image and text do not match.

[0126] It can be understood that when the first sample data and the second sample data are aligned, it means that the contents of the first sample data and the second sample data are matched, so when decoding, more reference can be made to the mixed sample data representation, that is, the proportion of the mixed sample data representation is controlled to be larger.

[0127] In some embodiments, the computer device may set the alignment label to 1 or 0, where 1 indicates alignment and 0 indicates misalignment. A second loss may be determined based on the difference between the alignment label and the predicted weight. For example, the second weight may be determined based on the difference between the alignment label and the predicted weight.

[0128] In some embodiments, the second loss is determined based on the alignment labels of the first sample data and the second sample data, and the prediction weights, including: constructing a first vector based on the alignment labels of the first sample data and the second sample data; constructing a second vector based on the prediction weights and the difference between the value one and the prediction weights; and determining the second loss based on the first vector and the second vector.

[0129] Specifically, the computer device may construct a first vector based on the alignment tags of the first sample data and the second sample data. Specifically, the value of the alignment tag and the difference between 1 and the alignment tag may be combined into a two-dimensional vector, namely, the first vector. For example, when the alignment tag is 1, the constructed first vector is [0, 1]; when the alignment tag is 0, the constructed first vector is [1, 0].

[0130] The computer device may construct a second vector using the predicted weight and the difference between the value 1 and the predicted weight, such as [1-p, p], where p is the predicted weight.

[0131] Furthermore, the computer device may determine the second loss based on the cross entropy of the first vector and the second vector. For example, the computer device may determine the second loss using the following formula:

[0132] L ITM =f CE (yitm ,[1-p,p]);(Formula 8)

[0133] Among them, y itm represents the first vector, p is the prediction weight, f CE represents the cross entropy function, L ITM Represents the second loss, L ITM Indicates the second loss.

[0134] In the above embodiment, determining the second loss based on the alignment labels of the first and second sample data and the predicted weights ensures that the predicted weights learned by the model during training are positively correlated with the degree of alignment of the input data. In other words, higher quality input data alignment results in greater predicted weights, allowing for dynamic control over the increased proportion of mixed data representation.

[0135] In step 214 , a target loss function is constructed based on the first loss and the second loss, and the initial model to be trained is trained based on the target loss function, and a multimodal model is obtained after the training is completed.

[0136] Specifically, the computer device may perform a weighted summation process on the first loss and the second loss to construct a target loss function, wherein the weighting coefficient may be 1 or another pre-set value, which is not limited in the embodiment of the present application.

[0137] Exemplarily, the computer device may construct the target loss function using the following formula:

[0138] L=L MIR +L MTR +L ITM ;(Formula 9)

[0139] Among them, L MIR Represents the first modal loss in the first loss, L MTR Represents the second modal loss in the first loss, L ITM represents the second loss, and L represents the target loss function.

[0140] Furthermore, the computer device can input multiple training sample pairs into the initial model in batches for processing, and adjust the model weights of the initial model through the target loss function to train the initial model. The training stops when the training stop condition is reached, and a multimodal model with training is obtained. Among them, the training stop condition can specifically be reaching a preset number of iterations, reaching a preset training time, or the model performance reaches a preset performance, or the change in model prediction accuracy is less than a preset change, etc., which is not limited in the embodiments of the present application.

[0141] The multimodal model processing method described above encodes first and second sample data belonging to different modalities, respectively, to obtain a first sample data representation and a second sample data representation. The first and second sample data representations are then combined to obtain a mixed sample data representation. Prediction weights are derived from this mixed sample data representation. These prediction weights are used to control the weights of the first and second sample data representations when they are fused with the mixed sample data representation. This allows for intelligent extraction of information from single and mixed modalities for subsequent decoding. During training, on the one hand, the loss between the predicted feature representation and the true feature representation is taken into account, and on the other hand, the matching loss between sample data of different modalities (that is, the second loss) is taken into account. This allows the model to learn encoding and decoding capabilities that are closer to the true feature representation. In addition, it can intelligently control the ratio of single data representation and mixed data representation used for reference during encoding and decoding, so that when the alignment quality of the training sample pairs is high, more mixed modal information can be considered during single-modal decoding; when the training sample pairs are not aligned, more single-modal information can be considered during single-modal decoding. This can better handle the problem of modal content mismatch in the training dataset and improve the processing performance of the multimodal model.

[0142] In some embodiments, before encoding the first sample data and the second sample data in the training sample pair, the method further includes: obtaining a training sample pair, the training sample pair including the first sample data and the second sample data; splitting the first sample data into multiple first data blocks, splitting the second sample data into multiple second data blocks, extracting some data blocks from the multiple first data blocks to form a first input sequence, and extracting some data blocks from the multiple second data blocks to form a second input sequence. Encoding the first sample data and the second sample data in the training sample pair to obtain a first sample data representation and a second sample data representation includes: encoding the first input sequence multiple times to obtain the first data representation, and encoding the second input sequence multiple times to obtain the second data representation.

[0143] Specifically, the computer device can train the initial model through the training sample pairs. During the training process, the first sample data and the second sample data in the training sample pairs can be directly processed, or the first sample data and the second sample data can be masked before being input into the initial model.

[0144] In some embodiments, the computer device may split the first sample data into multiple first data blocks and split the second sample data into multiple second data blocks. Then, a portion of the first data blocks is extracted according to a certain ratio to form a first input sequence, and a portion of the second data blocks is extracted to form a second input sequence. The first input sequence is input into one encoder of the initial model, and the second input sequence is input into another encoder of the initial model.

[0145] In some embodiments, for multiple first data blocks of the first sample data, the computer device may randomly mask a certain proportion of the first data blocks, and form the remaining first data blocks into a first input sequence. For multiple second data blocks of the second sample data, the computer device may randomly mask a certain proportion of the second data blocks, and form the remaining second data blocks into a second input sequence.

[0146] For example, the initial model uses two separate autoencoders to model different modal information. For one modality, such as image data, the computer device can randomly mask a certain proportion (e.g., 50%) of the patches. The remaining patches then form an input sequence and are fed into the autoencoder corresponding to the image data to obtain an image representation.

[0147] In some embodiments, taking the first modality as an image modality and the second modality as a text modality as an example, LightVLP (i.e., the multimodal model of this application) includes two autoencoders, one for processing image information and the other for processing text information. To improve the efficiency of the model, this application introduces a masked autoencoder strategy, which removes the masked tokens (masked features) from the sequences obtained by masking the sample data of the two modalities and then inputs them into their respective autoencoders.

[0148] For ease of understanding, please refer to Figure 3, which is a schematic diagram of a masking strategy in one embodiment. Taking the sample data (x1, x2, x3, x4, x5, x6) as an example, in the traditional masking strategy, referring to part A in Figure 3, some of the data blocks, such as x2, x4, and x5, are usually randomly masked to turn them into mask features, such as by replacing them with 0 or 1. Then, the sequence including the mask features is input as an input sequence to the encoder for processing. The masking strategy in this application is to refer to part B in Figure 3, randomly mask some of the data blocks, such as x2, x4, and x5, to turn them into mask features, such as by replacing them with 0 or 1. Then, the sequence after removing the mask features is input as an input sequence to the encoder for processing. Furthermore, after the input sequence is encoded by the encoder, the corresponding data representation (e1, e3, e6) is obtained, and then the mask features are inserted into the data representation for decoding. In the above embodiment, a mask encoding strategy is introduced in the process of processing data of different modalities. For each modality of data, part of the data blocks are extracted to form an input sequence, which can greatly reduce the amount of data in the encoding process of the model and improve the model processing efficiency, especially in the training stage, which can effectively improve the model training efficiency.

[0149] In some embodiments, a mixed sample data representation is obtained by combining the first sample data representation and the second sample data representation through the initial model, including: performing multiple interactive fusions of the first sample data representation and the second sample data representation through the initial model to obtain a first mixed sample representation biased towards the first modal side and a second mixed sample representation biased towards the second modal side; combining the first mixed sample representation and the second mixed sample representation to obtain a mixed sample data representation.

[0150] Specifically, in the process of fusing the first sample data representation and the second sample data representation, the initial model can perform multiple fusions through a multi-layer neural network to obtain a first mixed sample representation biased towards the first modal side and a second mixed sample representation biased towards the second modal side. Each time the interactive fusion is performed, the second intermediate mixed sample representation can be fused into the first intermediate mixed sample representation based on the attention mechanism to obtain a first intermediate mixed sample representation biased towards the first modal side. The first intermediate mixed sample representation is fused into the second intermediate mixed sample representation based on the attention mechanism to obtain a second intermediate mixed sample representation biased towards the second modal side. In this way, the first intermediate mixed sample representation output by the last layer of the neural network is the first mixed sample representation, and the second intermediate mixed sample representation output by the last layer of the neural network is the second mixed sample representation.

[0151] Furthermore, the initial model may fuse the first mixed sample representation and the second mixed sample representation to obtain a mixed sample data representation, for example, by weighted summation, concat (connection), or splicing, etc., which is not limited in the present embodiment.

[0152] In the above embodiment, the first sample data representation and the second sample data representation are interactively fused multiple times through the initial model, so that a first mixed sample representation biased towards the first modal side and a second mixed sample representation biased towards the second modal side can be obtained, and the relationship information between the first sample data and the second sample data can be fully extracted. Then, the first mixed sample representation and the second mixed sample representation are combined to obtain a mixed sample data representation.

[0153] In some embodiments, the first sample data representation and the second sample data representation are interactively fused multiple times through the initial model to obtain a first mixed sample representation biased towards the first modal side and a second mixed sample representation biased towards the second modal side, including: when performing the i+1th interactive fusion, obtaining the first intermediate mixed sample representation and the second intermediate mixed sample representation output by the i-th interactive fusion; fusing the second intermediate mixed sample representation into the first intermediate mixed sample representation to obtain the first intermediate mixed sample representation of the i+1th interactive fusion; fusing the first intermediate mixed sample representation into the second intermediate mixed sample representation to obtain the second intermediate mixed sample representation of the i+1th interactive fusion. ; Take i+1 as the new i, and return to the first intermediate mixed sample representation and the second intermediate mixed sample representation obtained by the i-th interactive fusion during the i+1-th interactive fusion, and continue to execute until the stopping condition is reached; take the first intermediate mixed sample representation output by the last interactive fusion as the first mixed sample representation, and take the second intermediate mixed sample representation output by the last interactive fusion as the second mixed sample representation; wherein, i is a natural number greater than or equal to 0. When i is 0, the first intermediate mixed representation output by the i-th interactive fusion obtained is the first sample data representation, and the second intermediate mixed representation output by the i-th interactive fusion obtained is the second sample data representation.

[0154] Specifically, for the first interactive fusion, the initial model can fuse the second sample data representation into the first sample data representation based on the attention mechanism to obtain a first intermediate mixed sample representation, and fuse the first sample data representation into the second sample data representation based on the attention mechanism to obtain a second intermediate mixed sample representation.

[0155] For the second interactive fusion, the initial model can fuse the second intermediate mixed sample representation obtained from the first interactive fusion into the first intermediate mixed sample representation based on the attention mechanism to obtain a new first intermediate mixed sample representation. The first intermediate mixed sample representation obtained from the first interactive fusion is fused into the second intermediate mixed sample representation based on the attention mechanism to obtain a new second intermediate mixed sample representation.

[0156] In this way, during each interactive fusion process, the current interactive fusion is continuously performed based on the output of the previous interactive fusion, until the final interactive fusion. The initial model can directly output the first intermediate mixed sample representation and the second intermediate mixed sample representation obtained from the last interactive fusion. The first intermediate mixed sample representation outputted this time is the first mixed sample representation, and the second intermediate mixed sample representation outputted this time is the second mixed sample representation.

[0157] In the above embodiment, after the first sample data representation and the second sample data representation are fused, the next interactive fusion is continuously performed based on the previous output through an iterative cycle, which can better utilize the information obtained from each fusion and make the information fusion more complete.

[0158] In some embodiments, the second intermediate mixed sample representation is fused into the first intermediate mixed sample representation to obtain the first intermediate mixed sample representation of the i+1th interactive fusion, including: performing an attention operation on the first intermediate mixed sample representation and the second intermediate mixed sample representation output by the i-th interactive fusion, combining the attention operation result with the first intermediate mixed sample representation output by the i-th interactive fusion, and then performing a layer regularization operation to obtain a first intermediate sample processing feature; performing a forward neural network processing on the first intermediate sample processing feature, combining the forward neural network processing result with the first intermediate sample processing feature, and then performing a layer regularization operation to obtain the first intermediate mixed sample representation of the i+1th interactive fusion.

[0159] Specifically, when performing a layer regularization operation after combining the attention operation result and the first intermediate mixed sample representation output by the i-th interactive fusion, the way of combining the attention operation result and the first intermediate mixed sample representation output by the i-th interactive fusion can be addition, multiplication, or splicing, etc., which is not limited in this embodiment of the present application. When performing a layer regularization operation after combining the forward neural network processing result and the first intermediate sample processing feature, the way of combining the forward neural network processing result and the first intermediate sample processing feature can be addition, multiplication, or splicing, etc., which is not limited in this embodiment of the present application.

[0160] Exemplarily, the computer device may determine the first intermediate mixed sample representation of the (i+1)th interactive fusion by the following formula:

[0161] in, is the first intermediate mixed sample representation obtained by the i-th interactive fusion, is the second intermediate mixed sample representation obtained by the i-th interactive fusion, Is the classic attention calculation formula, which is based on is the Q value, Attention processing is performed for K value and V value, and softmax is a normalization function. is the first intermediate mixed sample representation obtained by the i+1th interactive fusion. LN Representation layer regularization operation, It represents the processing of the feedforward neural network (specifically, a two-layer fully connected network) of the i-th layer of the interactive conversion structure.

[0162] In the above embodiment, the interactive fusion processing method based on the attention mechanism can better integrate the second intermediate mixed representation into the first intermediate mixed representation. Moreover, this stacked fusion design can prevent the gradient explosion problem during model training.

[0163] In some embodiments, the first intermediate mixed sample representation is fused into the second intermediate mixed sample representation to obtain the second intermediate mixed sample representation of the i+1th interactive fusion, including: performing an attention operation on the second intermediate mixed sample representation output by the i-th interactive fusion and the first intermediate mixed sample representation, combining the attention operation result with the second intermediate mixed sample representation output by the i-th interactive fusion and performing a layer regularization operation to obtain a second intermediate sample processing feature; performing a forward neural network processing on the second intermediate sample processing feature, combining the forward neural network processing result with the second intermediate sample processing feature and performing a layer regularization operation to obtain the second intermediate mixed sample representation of the i+1th interactive fusion.

[0164] Specifically, when performing a layer regularization operation after combining the attention operation result and the second intermediate mixed sample representation output by the i-th interactive fusion, the way of combining the attention operation result and the second intermediate mixed sample representation output by the i-th interactive fusion can be addition, multiplication, or splicing, etc., which is not limited in the embodiments of the present application. When performing a layer regularization operation after combining the forward neural network processing result and the second intermediate sample processing feature, the way of combining the forward neural network processing result and the second intermediate sample processing feature can be addition, multiplication, or splicing, etc., which is not limited in the embodiments of the present application.

[0165] Exemplarily, the computer device may determine the second intermediate mixed sample representation of the (i+1)th interactive fusion by the following formula:

[0166] in, is the first intermediate mixed sample representation obtained by the i-th interactive fusion, is the second intermediate mixed sample representation of the i-th interactive fusion, Is the classic attention calculation formula, which is based on is the Q value, Attention processing is performed for K value and V value, and softmax is a normalization function. It is the second intermediate mixed sample representation of the i+1th interactive fusion. LN Representation layer regularization operation, It represents the processing of the feedforward neural network (specifically, a two-layer fully connected network) of the i-th layer of the interactive conversion structure.

[0167] In the above embodiment, the interactive fusion processing method based on the attention mechanism can better fuse the first intermediate mixed representation into the second intermediate mixed representation. Moreover, through this stacked fusion design, the gradient explosion problem can be prevented during model training.

[0168] In some embodiments, referring to FIG4 , the multimodal model processing method includes the following steps:

[0169] Step 402: Obtain a training sample pair, the training sample pair including first sample data and second sample data; the first sample data and the second sample data are sample data of different modalities; through the initial model to be trained, encode based on the first sample data to obtain a first sample data representation; and encode based on the second sample data to obtain a second sample data representation.

[0170] Step 404 : The first sample data representation and the second sample data representation are combined with the initial model to obtain a first mixed sample representation biased toward the first modality and a second mixed sample representation biased toward the second modality.

[0171] Specifically, the initial model performs multiple interactive fusions on the first sample data representation and the second sample data representation to obtain a first mixed sample representation biased towards the first modality and a second mixed sample representation biased towards the second modality. For details on the manner of performing multiple interactive fusions, please refer to the description of the aforementioned related embodiments.

[0172] Step 406: Determine a first prediction weight based on the first mixed sample representation, and determine a second prediction weight based on the second mixed sample representation.

[0173] Specifically, the initial model may perform a linear transformation on the first mixed sample representation and then perform activation processing to obtain a first prediction weight; and perform a linear transformation on the second mixed sample representation and then perform activation processing to obtain a second prediction weight.

[0174] Step 408: The first sample data representation and the first mixed sample representation are fused according to the first prediction weight through the initial model, and decoding is performed based on the fused result to obtain a first prediction feature representation.

[0175] Specifically, the initial model can use the first prediction weight as the coefficient of the first mixed sample representation, and the difference between the value one and the first prediction weight as the coefficient of the first sample data representation; according to the coefficient of the first mixed sample representation and the coefficient of the first sample data representation, the first mixed sample representation and the first sample data representation are weightedly summed, and then decoded based on the result of the weighted summation to obtain the first prediction feature representation.

[0176] Step 410: fuse the second sample data representation and the second mixed sample representation according to the second prediction weight through the initial model, and perform decoding based on the fused result to obtain a second prediction feature representation.

[0177] Specifically, the initial model can use the second prediction weight as the coefficient of the second mixed sample representation, and the difference between the value one and the second prediction weight as the coefficient of the second sample data representation; according to the coefficient of the second mixed sample representation and the coefficient of the second sample data representation, the second mixed sample representation and the second sample data representation are weightedly summed, and then decoded based on the result of the weighted summation to obtain the second prediction feature representation.

[0178] Step 412: Determine a first loss based on the true feature representation and the first predicted feature representation of the first sample data, and the true feature representation and the second predicted feature representation of the second sample data.

[0179] Step 414 : Determine a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weights.

[0180] In step 416 , a target loss function is constructed based on the first loss and the second loss, and the initial model to be trained is trained based on the target loss function, and a multimodal model is obtained after the training is completed.

[0181] In the above embodiment, for the first modality, the initial model will use the first sample modality as the main information, and at the same time integrate the information represented by the second sample data to obtain a first mixed sample representation, and then determine the prediction weight of the first modality side based on the first mixed sample representation, that is, the first prediction weight. In this way, the proportion of single modality information and mixed modality information referenced during decoding on the first modality side can be controlled based on the first prediction weight. For the second modality, the initial model will use the second sample modality as the main information, and at the same time integrate the information represented by the first sample data to obtain a second mixed sample representation, and then determine the prediction weight of the first modality side based on the second mixed sample representation, that is, the second prediction weight. In this way, the proportion of single modality information and mixed modality information referenced during decoding on the second modality side can be controlled based on the second prediction weight. In this way, different modality sides can be more flexible and adaptive when fusing mixed modality information.

[0182] In some embodiments, the mixed sample data representation includes a first mixed sample representation biased toward the first modality and a second mixed sample representation biased toward the second modality, and the prediction weights include a first prediction weight determined based on the first mixed sample representation and a second prediction weight determined based on the second mixed sample representation. Determining a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weights, includes: constructing a first vector based on the alignment labels of the first sample data and the second sample data; determining a target weight based on the mean of the first prediction weight and the second prediction weight, constructing a third vector based on the target weight and the difference between the value one and the target weight; and determining a second loss based on the first vector and the third vector.

[0183] As mentioned above, for different modalities, the initial model will generate a prediction loss that is adapted to the modality. During the model training process, the average of the first prediction loss and the second prediction loss can be used as the target loss, and the second loss can be determined based on the target loss and the alignment label.

[0184] Specifically, the computer device may construct a first vector based on the alignment tags of the first sample data and the second sample data. Specifically, the value of the alignment tag and the difference between 1 and the alignment tag may be combined into a two-dimensional vector, namely, the first vector. For example, when the alignment tag is 1, the constructed first vector is [0, 1]; when the alignment tag is 0, the constructed first vector is [1, 0].

[0185] The computing device may construct a second vector using the target weight and the difference between the value 1 and the target weight. For example, [1-p, p], where p is the target weight. Furthermore, the computing device may determine a second loss based on the cross entropy between the first and second vectors.

[0186] In the above embodiment, determining the second loss based on the alignment labels of the first and second sample data and the average of the first and second prediction weights ensures that the two prediction weights learned by the model during training are positively correlated with the degree of alignment of the input data. In other words, the higher the quality of input data alignment, the larger the two predicted weights, which can dynamically control the increase in the proportion of mixed data representation.

[0187] In some embodiments, the method further includes constructing a third loss, wherein the step of constructing the third loss includes: determining the correlation between the first sample data and the second sample data based on the first sample data representation and the second sample data representation; and determining the third loss based on the alignment labels of the first sample data and the second sample data and the correlation between the first sample data and the second sample data. Constructing a target loss function based on the first loss and the second loss includes: constructing the target loss function based on the first loss, the second loss, and the third loss.

[0188] In some embodiments, in order to align the representations of different modalities, the present application introduces contrastive loss, that is, the third loss. By adding the third loss to the target loss for learning, the model can compensate for the performance impact caused by the data misalignment of training sample pairs during the training process.

[0189] Specifically, the computer device may obtain a first sample data representation output by the first encoder and a second sample data representation output by the second encoder, and then input the first sample data representation into a linear layer and a regularization layer to obtain a first transformed representation, and input the second sample data representation into a linear layer and a regularization layer to obtain a second transformed representation.

[0190] Exemplarily, the computer device may calculate the first conversion representation using the following formula:

[0191] Among them, f NORM represents regularization processing, f x represents linearization processing, represents the first sample data representation, and T1 represents the first transformed representation.

[0192] Exemplarily, the computer device may calculate the second conversion representation using the following formula:

[0193] Among them, f NORM represents regularization processing, f w represents linearization processing, represents the second sample data representation, and T2 represents the second transformed representation.

[0194] Furthermore, the computer device may determine the correlation between the first sample data and the second sample data based on the first transformed representation and the second transformed representation. For example, the computer device may calculate the similarity between the first transformed representation and the second transformed representation, and use the similarity as the correlation between the first sample data and the second sample data. The similarity may specifically be cosine similarity or distance similarity, etc., which is not limited in this embodiment of the present application.

[0195] Next, the computer device may determine a third loss based on the alignment labels of the first sample data and the second sample data, and the difference between the correlations between the first sample data and the second sample data. Specifically, the computer device may calculate the cross entropy between the alignment labels and the correlations, and use the cross entropy as the third loss.

[0196] Furthermore, the computer device may construct a target loss function based on the first loss, the second loss, and the third loss. For example, the computer device may construct the target loss function by summing the first loss, the second loss, and the third loss. For example, the computer device may construct the target loss function by the following formula:

[0197] L=L MIR +L MTR +L ITM +L ITC ;(Formula 16)

[0198] Among them, L MIR Represents the first modal loss in the first loss, L MTR Represents the second modal loss in the first loss, L ITM Represents the second loss, L ITC represents the third loss, and L represents the target loss function.

[0199] In the above embodiment, the contrast loss is introduced during the model training process, which can guide the model to learn the ability to distinguish whether the representations of different modalities are aligned.

[0200] In some embodiments, the correlation between the first sample data and the second sample data includes: a correlation between the first sample data and the second sample data, and a correlation between the second sample data and the first sample data. Determining the correlation between the first sample data and the second sample data based on the first sample data representation and the second sample data representation includes: determining a preset number of third sample data having the same modality as the first sample data, and a preset number of fourth sample data having the same modality as the second sample data; calculating the correlation between the first sample data and the second sample data based on the first sample data representation, the second sample data representation, and the sample data representations of each of the fourth sample data; and calculating the correlation between the second sample data and the first sample data based on the first sample data representation, the second sample data representation, and the sample data representations of each of the third sample data.

[0201] When calculating the correlation between the first sample data and the second sample data, the correlation of the first sample data with respect to the second sample data and the correlation of the second sample data with respect to the first sample data may be calculated respectively.

[0202] Specifically, the computer device performs linear processing and regularization processing on the first sample data representation to obtain a first transformed representation, performs linear processing and regularization processing on the second sample data representation to obtain a second transformed representation, performs linear processing and regularization processing on the third sample data representation to obtain a third transformed representation, and performs linear processing and regularization processing on the fourth sample data representation to obtain a fourth transformed representation.

[0203] Furthermore, the computer device can calculate the correlation of the first sample data relative to the second sample data based on the first conversion representation, the second conversion representation, and each fourth conversion representation; and calculate the correlation of the second sample data relative to the first sample data based on the first conversion representation, the second conversion representation, and each third conversion representation.

[0204] In some embodiments, the correlation of the first sample data with respect to the second sample data is calculated based on the first sample data representation, the second sample data representation, and the sample data representations of each fourth sample data, including: determining a first value based on the first sample data representation and the second sample data representation; determining multiple second values ​​based on the first sample data representation and the sample data representations of each fourth sample data; and using a comparison value of the first value and the sum of the multiple second values ​​as the correlation of the first sample data with respect to the second sample data.

[0205] Specifically, the computer device may determine the first value based on the product of the first converted representation and the second converted representation, for example, by directly using the product as the first value, or by using the quotient of the product and the temperature coefficient as the exponent, and raising the result of the power operation with the natural constant e as the base as the first value. Similarly, the computer device may determine the second value based on the product of the first converted representation and each fourth converted representation. Furthermore, the computer device uses a comparison value between the first value and the sum of the multiple second values ​​as the correlation between the first sample data and the second sample data. The comparison value may specifically be a quotient value, a difference value, or other value that can reflect the difference.

[0206] For example, the computer device may calculate the correlation between the first sample data and the second sample data using the following formula:

[0207] in, represents the correlation between the first sample data and the second sample data; I represents the first transformed representation of the first sample data in the training sample pair, T g ' represents the second transformed representation of the second sample data corresponding to the first sample data in the training sample pair; Mc is the preset number, T j ' represents the fourth conversion representation of the j-th fourth sample data; k represents the temperature coefficient.

[0208] In the above embodiment, when calculating the correlation of the first sample data with respect to the second sample data, the fourth sample data having different content from the second sample data but the same modality is utilized, and the degree of alignment of the sample data pair with respect to other data pairs composed of the first sample data can be calculated, thereby accurately obtaining the correlation of the first sample data with respect to the second sample data.

[0209] In some embodiments, the correlation of the second sample data with respect to the first sample data is calculated based on the first sample data representation, the second sample data representation, and the sample data representations of each third sample data, including: determining a first value based on the first sample data representation and the second sample data representation; determining multiple third values ​​based on the second sample data representation and the sample data representations of each third sample data; and using a comparison value of the first value and the sum of the multiple third values ​​as the correlation of the second sample data with respect to the first sample data.

[0210] Specifically, the computer device may determine the first value based on the product of the first converted representation and the second converted representation, for example, by directly using the product as the first value, or by using the quotient of the product and the temperature coefficient as the exponent, with the natural constant e as the base, as the result of the power operation. Similarly, the computer device may determine the third value based on the product of the second converted representation and each third converted representation. Furthermore, the computer device uses a comparison value between the first value and the sum of the multiple third values ​​as the correlation of the second sample data with the first sample data. The comparison value may specifically be a quotient value, a difference value, or other numerical value that can reflect the difference.

[0211] For example, the computer device may calculate the correlation between the first sample data and the second sample data using the following formula:

[0212] in, Represents the correlation between the second sample data and the first sample data; T represents the second transformed representation of the second sample data in the training sample pair, I g ' represents the first conversion representation of the first sample data corresponding to the second sample data in the training sample pair; Mc is the preset number, I j ' represents the third conversion representation of the j-th third sample data; k represents the temperature coefficient.

[0213] In the above embodiment, when calculating the correlation of the second sample data with respect to the first sample data, third sample data having different content from the first sample data but the same modality as the first sample data is utilized, and the degree of alignment of the sample data pair with respect to other data pairs composed of the second sample data can be calculated, thereby accurately obtaining the correlation of the second sample data with respect to the first sample data.

[0214] Furthermore, the computer device can determine a loss based on the alignment labels of the first sample data and the second sample data, and the difference between the correlation of the first sample data and the second sample data, and determine another loss based on the alignment labels and the difference between the correlation of the second sample data and the first sample data, and then average the two losses to obtain a third loss.

[0215] Exemplarily, the computer device may calculate the third loss by the following formula:

[0216] Among them, for a set of training sample pairs, y i2t (I) and y t2i (T) is the same, which represents the alignment label of the first sample data and the second sample data in the training sample pair. i2t (I) represents the correlation between the first sample data and the second sample data, pt2i (T) represents the correlation between the second sample data and the first sample data. CE is the cross entropy loss function. L ITC Indicates the third loss.

[0217] In the above embodiment, the fourth sample data having different content from the second sample data but the same modality can be used to accurately calculate the correlation of the first sample data with respect to the second sample data; the third sample data having different content from the first sample data but the same modality can be used to accurately obtain the correlation of the second sample data with respect to the first sample data, thereby obtaining the correlation in two dimensions, which can more accurately measure the degree of alignment between the representations of the two modalities after processing by the encoder in the multimodal model.

[0218] In some embodiments, the modality includes an image modality and a text modality, and the method further includes a text retrieval step, specifically including: obtaining a text retrieval task, extracting the image data specified in the text retrieval task, and forming a first data pair to be processed by combining each text data in the text library with the specified image data; processing the first data pair through a multimodal model to obtain two feature representations, and determining the matching degree of the first data pair based on the two feature representations; and determining the target text that matches the specified image data according to the matching degree of each first data pair.

[0219] The multimodal model trained through the above embodiment can be applied in image-text retrieval scenarios, for example, it can be used to retrieve text that matches an image. In actual applications, if a text retrieval task is obtained, the computer device can extract the image data specified in the text retrieval task, and then form multiple first data pairs with each text data in the text library and the specified image data. Thus, any first data pair is processed by the trained multimodal model, and two feature representations corresponding to the first data pair are output. Thus, the computer device can determine the matching degree of the first data pair based on the two feature representations, and according to the matching degree of each first data pair, the text data in the first data whose matching degree meets the push condition is used as the target text that matches the specified image data and is pushed to the terminal for display, that is, the text retrieval result of the specified image data is obtained. Among them, the matching degree meeting the push condition can specifically be the highest matching degree, or the matching degree is greater than a preset threshold, or the top N with the highest matching degree, etc., which is not limited in the embodiment of the present application.

[0220] In some embodiments, the specific method for determining the matching degree of the first data pair can be any of the following methods: calculating the feature similarity of the two feature representations and using the feature similarity as the matching degree; inputting the two feature representations into a classification layer, having the classification layer output a classification result, and using the classification result as the matching degree. The feature similarity can be represented by cosine similarity, Euclidean distance, Manhattan distance, etc.

[0221] In the above embodiment, the trained multimodal model can be used to perform image and text matching calculations, thereby helping to improve the quality and efficiency of text retrieval.

[0222] In some embodiments, the modality includes images and text, and the method also includes an image retrieval step, specifically including: obtaining an image retrieval task, extracting the text data specified in the image retrieval task, and forming a second data pair to be processed by respectively combining each image data in the image library with the specified text data; processing the second data pair through a multimodal model to obtain two feature representations, and determining the matching degree of the second data pair based on the two feature representations; and determining the target image that matches the specified text data according to the matching degree of each second data pair.

[0223] The multimodal model obtained through training in the above embodiment can be applied in image and text retrieval scenarios, for example, it can be used to retrieve images that match text. In actual applications, if an image retrieval task is obtained, the computer device can extract the text data specified in the image retrieval task, and then form multiple second data pairs with each image data in the image library and the specified text data. Thus, any second data pair is processed by the trained multimodal model, and two feature representations corresponding to the second data pair are output. Thus, the computer device can determine the matching degree of the second data pair based on the two feature representations, and according to the matching degree of each second data pair, the image data in the second data whose matching degree meets the push condition is used as the target image that matches the specified text data and is pushed to the terminal for display, that is, the image retrieval result of the specified text data is obtained. Among them, the matching degree meeting the push condition can specifically be the highest matching degree, or the matching degree is greater than a preset threshold, or the top N with the highest matching degree, etc., and the embodiment of the present application does not limit this.

[0224] In some embodiments, the specific method for determining the matching degree of the second data pair can be any of the following methods: calculating the feature similarity of the two feature representations and using the feature similarity as the matching degree; inputting the two feature representations into a classification layer, having the classification layer output a classification result, and using the classification result as the matching degree. The feature similarity can be represented by cosine similarity, Euclidean distance, Manhattan distance, etc.

[0225] In the above embodiment, the trained multimodal model can be used to perform image and text matching calculations, thereby helping to improve the quality and efficiency of image retrieval.

[0226] In an exemplary embodiment, as shown in FIG5 , a data matching method is provided, which is described by taking the method applied to a computer device (such as the terminal or server in FIG1 ) as an example, and includes the following steps. Among them:

[0227] Step 502: obtain first data, encode the first data, and obtain a first data representation; obtain second data, encode the second data, and obtain a second data representation; the first data and the second data are data of different modalities.

[0228] The computer device may encode the first data and the second data respectively to obtain a first data representation and a second data representation, where the first data and the second data are data of different modalities.

[0229] The present application can encode based on the first data through a first encoder in a multimodal model to obtain a first data representation; and encode based on the second data through a second encoder in the multimodal model to obtain a second data representation.

[0230] In some embodiments, the computer device may directly encode the first data to obtain the first data representation, and directly encode the second data to obtain the second data representation. In other embodiments, the computer device may mask the first data, and then encode the masked first data using a first encoder in the multimodal model to obtain the first data representation. The computer device may mask the second data, and then encode the masked second data using a second encoder in the multimodal model to obtain the second data representation.

[0231] For ease of description, the data input to the first encoder is referred to as the first input sequence, and the data input to the second encoder is referred to as the second input sequence. That is, the first input sequence can be the first data or the masked first data; the second input sequence can be the second data or the masked second data.

[0232] In some embodiments, the multimodal model encodes the first input sequence multiple times to obtain a first data representation, and the multimodal model encodes the second input sequence multiple times to obtain a second data representation.

[0233] In some embodiments, the first encoder and the second encoder have the same network structure, so the encoding process performed on the first input sequence is the same as the encoding process performed on the second input sequence. Of course, in other embodiments, the first encoder and the second encoder can have different network structures, and the encoding process performed on the first input sequence is different from the encoding process performed on the second input sequence. This embodiment of the present application is not limited to this.

[0234] Taking the example of the first encoder and the second encoder having the same network structure, the following is explained: for any encoding, such as the l+1th encoding, the output of the lth encoding can be obtained first. A multi-head self-attention layer operation can be performed based on the output of the lth encoding. The result of the multi-head self-attention layer operation and the output of the lth encoding are combined and then layer regularization is performed to obtain intermediate processing features; the intermediate processing features are processed by a forward neural network, and the result of the forward neural network processing and the intermediate processing features are combined and then layer regularization is performed to obtain the output of the l+1th encoding; l+1 is used as the new l, and the output of the lth encoding is returned to continue execution until the last encoding process is completed, and the output of the last encoding is used as the first data representation or the second data representation. Where l is a natural number greater than or equal to 0. When l is 0, the output of the lth encoding obtained is the first input sequence or the second input sequence.

[0235] In some embodiments, the computer device may encode each input data in the first input sequence multiple times to obtain an output corresponding to each input data, and then concatenate the outputs corresponding to each input data in the last encoding to obtain the first data representation. Similarly, the computer device may encode each input data in the second input sequence multiple times to obtain an output corresponding to each input data, and then concatenate the outputs corresponding to each input data in the last encoding to obtain the second data representation.

[0236] Exemplarily, the computer device may determine the output of the (l+1)th encoding by using formula (1) in step 202.

[0237] Step 504: Combine the first data representation and the second data representation to obtain a mixed data representation, and determine a fusion weight based on the mixed data representation.

[0238] This application designs a gated interaction mechanism to control the ratio of reference single data representation and mixed data representation during decoding, that is, dynamically adjust the reference ratio through fusion weights. The value range of fusion weights can be [0, 1] or (0, 1).

[0239] Specifically, the multimodal model can fuse the first data representation and the second data representation through a cross transformer structure to obtain a hybrid data representation, and then perform a numerical transformation on the hybrid data representation to obtain a fusion weight.

[0240] In some embodiments, the computer device may directly perform addition or multiplication operations on the first data representation and the second modality to obtain a mixed data representation. In other embodiments, the computer device may also perform an attention operation on the first data representation and the second data representation to obtain a mixed data representation. In other embodiments, the computer device may also use more complex operations, such as performing an attention mechanism, then performing matrix addition or multiplication, and then performing linear / nonlinear transformations to obtain a mixed data representation. This application does not limit the specific fusion method.

[0241] Furthermore, after combining the first and second data representations to obtain a hybrid data representation, the computer device can perform a linear transformation on the hybrid data representation and then perform activation processing to obtain an adaptive fusion weight. In other words, the hybrid modality is converted into a numerical value, which is also the fusion weight. In this way, the adaptive fusion weight for that particular time can be obtained based on the relationship between the first and second data at that time.

[0242] Step 506: According to the fusion weight, the first data representation and the mixed data representation are fused to obtain a first fused representation, and the second data representation and the mixed data representation are fused to obtain a second fused representation.

[0243] Specifically, the computer device may perform weighted processing on the first data representation and the mixed data representation according to the fusion weight to obtain a first fused representation; and perform weighted processing on the second data representation and the mixed data representation according to the fusion weight to obtain a second fused representation.

[0244] In some embodiments, the computer device may use the fusion weight as the weight of any one of the first data representation and the mixed data representation, and use the difference between the value one and the fusion weight as the weight of the other representation, and perform weighted summation processing on the first data representation and the mixed data representation.

[0245] In some embodiments, the computer device may use the fusion weight as the weight of any one of the second data representation and the mixed data representation, and use the difference between the value one and the fusion weight as the weight of the other representation, and perform weighted summation processing on the second data representation and the mixed data representation.

[0246] Step 508: Determine the matching degree between the first data and the second data according to the first fused representation and the second fused representation.

[0247] In some embodiments, the computer device can directly calculate the matching degree between the first fused representation and the second fused representation, and use the matching degree as the matching degree between the first data and the second data. In other words, in the use scenario of the multimodal model, the first encoder, the second encoder, and the interactive transformation structure can be directly used.

[0248] In other embodiments, determining the degree of matching between the first data and the second data based on the first fused representation and the second fused representation includes: decoding according to the first fused representation to obtain a first feature representation of the first data; decoding according to the second fused representation to obtain a second feature representation of the second data; and determining the degree of matching between the first data and the second data based on the first feature representation and the second feature representation.

[0249] Among them, please refer to the decoding method in the processing process of the multimodal model in the aforementioned embodiment for the specific decoding method. The specific method for determining the matching degree between the first data and the second data can be any of the following methods: calculating the feature similarity between the two representations (the first fusion representation and the second fusion representation, or the first feature representation and the second feature representation), and using the feature similarity as the matching degree; inputting the two representations (the first fusion representation and the second fusion representation, or the first feature representation and the second feature representation) into the classification layer, outputting the classification result through the classification layer, and using the classification result as the matching degree. Among them, the feature similarity can be represented by cosine similarity, Euclidean distance, Manhattan distance, etc.

[0250] The above-mentioned data matching method encodes the first data and the second data belonging to different modalities respectively to obtain the first data representation and the second data representation. Then, the first data representation and the second data representation are combined to obtain a mixed data representation. According to the mixed data representation, a fusion weight can be obtained. The fusion weight is used to control the proportion of the first data representation and the second data representation when they are fused with the mixed data representation respectively, that is, information can be intelligently extracted from the single modality / mixed modality to obtain the first fusion representation and the second fusion representation. The first fusion representation and the second fusion representation obtained in this way can better highlight the relevant information of the first data and the second data, so they can be used to more accurately determine the matching degree between the first data and the second data, greatly improving the data processing effect.

[0251] Especially when there is a missing between the first data and the second data, by intelligently extracting information from the mixed modality and intelligently selecting the proportion of the extracted mixed modal information, a first fused representation of the first data and a second fused representation of the second data are obtained through multimodal information reconstruction, which makes up for the missing information, greatly improves the matching accuracy, and improves the data processing performance.

[0252] In some embodiments, combining the first data representation and the second data representation to obtain a mixed data representation includes: performing multiple interactive fusions based on the first data representation and the second data representation to obtain a first mixed representation biased towards the first modal side and a second mixed representation biased towards the second modal side; combining the first mixed representation and the second mixed representation to obtain a mixed data representation.

[0253] Specifically, in the process of fusing the first data representation and the second data representation, the initial model can perform multiple fusions through a multi-layer neural network to obtain a first mixed representation biased towards the first modality and a second mixed representation biased towards the second modality. Each time the interactive fusion is performed, the second intermediate mixed representation can be fused into the first intermediate mixed representation based on the attention mechanism to obtain a first intermediate mixed representation biased towards the first modality. The first intermediate mixed representation can be fused into the second intermediate mixed representation based on the attention mechanism to obtain a second intermediate mixed representation biased towards the second modality. In this way, the first intermediate mixed representation output by the last layer of the neural network is the first mixed representation, and the second intermediate mixed representation output by the last layer of the neural network is the second mixed representation.

[0254] Furthermore, the initial model may fuse the first mixed representation and the second mixed representation to obtain a mixed data representation, for example, by weighted summation, or by contact, or by concatenation, etc., which is not limited in the present embodiment.

[0255] In the above embodiment, the first data representation and the second data representation are interactively fused multiple times through the initial model, so that a first mixed representation biased towards the first modal side and a second mixed representation biased towards the second modal side can be obtained, and the relationship information between the first data and the second data can be fully extracted. Then, the first mixed representation and the second mixed representation are combined to obtain a mixed data representation.

[0256] In some embodiments, multiple interactive fusions are performed based on the first data representation and the second data representation to obtain a first mixed representation biased towards the first modal side and a second mixed representation biased towards the second modal side, including: when performing the i+1th interactive fusion, obtaining the first intermediate mixed representation and the second intermediate mixed representation output by the i-th interactive fusion; fusing the second intermediate mixed representation into the first intermediate mixed representation to obtain the first intermediate mixed representation of the i+1th interactive fusion; fusing the first intermediate mixed representation into the second intermediate mixed representation to obtain the second intermediate mixed representation of the i+1th interactive fusion; taking i+1 as the new i, and returning to the first intermediate mixed representation and the second intermediate mixed representation obtained by the i-th interactive fusion when performing the i+1th interactive fusion, and continuing to execute until the stopping condition is reached; taking the first intermediate mixed representation output by the last interactive fusion as the first mixed representation, and taking the second intermediate mixed representation output by the last interactive fusion as the second mixed representation; wherein i is a natural number greater than or equal to 0, when i is 0, the first intermediate mixed representation output by the i-th interactive fusion obtained is the first data representation, and the second intermediate mixed representation output by the i-th interactive fusion obtained is the second data representation.

[0257] Specifically, for the first interactive fusion, the initial model can fuse the second data representation into the first data representation based on the attention mechanism to obtain a first intermediate mixed representation, and fuse the first data representation into the second data representation based on the attention mechanism to obtain a second intermediate mixed representation.

[0258] For the second interactive fusion, the initial model can fuse the second intermediate mixed representation obtained from the first interactive fusion into the first intermediate mixed representation based on the attention mechanism to obtain a new first intermediate mixed representation. The initial model can fuse the first intermediate mixed representation obtained from the first interactive fusion into the second intermediate mixed representation based on the attention mechanism to obtain a new second intermediate mixed representation.

[0259] In this way, each interactive fusion process continuously proceeds based on the output of the previous interactive fusion, until the final interactive fusion. The initial model can directly output the first intermediate mixed representation and the second intermediate mixed representation obtained from the last interactive fusion. The first intermediate mixed representation outputted this time is the first mixed representation, and the second intermediate mixed representation outputted this time is the second mixed representation.

[0260] In the above embodiment, after the first data representation and the second data representation are fused, the next interactive fusion is continuously performed based on the previous output through an iterative cycle, which can better utilize the information obtained from each fusion and make the information fusion more complete.

[0261] In some embodiments, the second intermediate mixed representation is fused into the first intermediate mixed representation to obtain the first intermediate mixed representation of the i+1th interactive fusion, including: performing an attention operation based on the first intermediate mixed representation and the second intermediate mixed representation output by the i-th interactive fusion, combining the result of the attention operation with the first intermediate mixed representation output by the i-th interactive fusion, and then performing a layer regularization operation to obtain a first intermediate processing feature; performing a forward neural network processing on the first intermediate processing feature, combining the forward neural network processing result with the first intermediate processing feature, and then performing a layer regularization operation to obtain the first intermediate mixed representation of the i+1th interactive fusion.

[0262] Specifically, when performing a layer regularization operation after combining the result of the attention operation with the first intermediate mixed representation of the output of the i-th interactive fusion, the way in which the result of the attention operation and the first intermediate mixed representation of the output of the i-th interactive fusion are combined can be addition, multiplication, or splicing, etc., and this embodiment of the present application does not limit this. When performing a layer regularization operation after combining the result of the forward neural network processing with the first intermediate processing feature, the way in which the forward neural network processing result and the first intermediate processing feature are combined can be addition, multiplication, or splicing, etc., and this embodiment of the present application does not limit this.

[0263] Similar to the processing method in the model training process, the computer device can determine the first intermediate mixed representation of the (i+1)th interactive fusion through the above formulas 10 and 11.

[0264] In the above embodiment, the interactive fusion processing method based on the attention mechanism can better integrate the second intermediate mixed representation into the first intermediate mixed representation. Moreover, this stacked fusion design can prevent the gradient explosion problem during model training.

[0265] In some embodiments, the first intermediate mixed representation is fused into the second intermediate mixed representation to obtain the second intermediate mixed representation of the i+1th interactive fusion, including: performing an attention operation based on the second intermediate mixed representation output by the i-th interactive fusion and the first intermediate mixed representation, combining the result of the attention operation with the second intermediate mixed representation output by the i-th interactive fusion and performing a layer regularization operation to obtain a second intermediate processing feature; performing a forward neural network processing on the second intermediate processing feature, combining the forward neural network processing result with the second intermediate processing feature and performing a layer regularization operation to obtain the second intermediate mixed representation of the i+1th interactive fusion.

[0266] Specifically, when performing a layer regularization operation after combining the result of the attention operation with the second intermediate mixed representation output by the i-th interactive fusion, the way in which the result of the attention operation and the second intermediate mixed representation output by the i-th interactive fusion are combined can be addition, multiplication, or splicing, etc., which is not limited in this embodiment of the present application. When performing a layer regularization operation after combining the result of the feedforward neural network processing with the second intermediate processing feature, the way in which the result of the feedforward neural network processing with the second intermediate processing feature are combined can be addition, multiplication, or splicing, etc., which is not limited in this embodiment of the present application.

[0267] Similar to the processing method in the model training process, the computer device can determine the second intermediate mixed representation of the (i+1)th interactive fusion through Formula 12 and Formula 13 in the aforementioned embodiment.

[0268] In the above embodiment, the interactive fusion processing method based on the attention mechanism can better fuse the first intermediate mixed representation into the second intermediate mixed representation. Moreover, through this stacked fusion design, the gradient explosion problem can be prevented during model training.

[0269] In some embodiments, the first data representation and the mixed data representation are fused according to the fusion weight to obtain a first fused representation, including: using the fusion weight as the coefficient of the mixed data representation, and using the difference between the value one and the fusion weight as the coefficient of the first data representation; according to the coefficient of the mixed data representation and the coefficient of the first data representation, the mixed data representation and the first data representation are weighted to obtain the first fused representation.

[0270] Specifically, the initial model may perform weighted summation of the mixed data representation and the first data representation according to the coefficients of the mixed data representation and the coefficients of the first data representation to obtain the first fused representation.

[0271] In the above embodiment, by fusing the mixed data representation and the first modality through fusion weights, the ratio of the single data representation and the mixed data representation can be flexibly referred to according to the alignment quality of the two modal data, so that the fused representation can better express the information of the single modal data.

[0272] In some embodiments, the second data representation and the mixed data representation are fused according to the fusion weight to obtain the second fused representation, including: using the fusion weight as the coefficient of the mixed data representation, and using the difference between the value one and the fusion weight as the coefficient of the second data representation; according to the coefficient of the mixed data representation and the coefficient of the second data representation, the mixed data representation and the second data representation are weightedly processed to obtain the first fused representation.

[0273] Specifically, the initial model may perform weighted summation on the mixed data representation and the second data representation according to the coefficients of the mixed data representation and the coefficients of the second data representation to obtain the second fused representation.

[0274] In the above embodiment, by fusing the mixed data representation and the second modality through fusion weights, the ratio of the single data representation and the mixed data representation can be flexibly referred to according to the alignment quality of the two modal data, so that the fused representation can better express the information of the single modal data.

[0275] In some embodiments, referring to FIG6 , the data matching method includes the following steps:

[0276] Step 602: obtain first data, encode the first data, and obtain a first data representation; obtain second data, encode the second data, and obtain a second data representation. The first data and the second data are data of different modalities.

[0277] Step 604 : Combine the first data representation and the second data representation to obtain a first mixed representation biased toward the first modality and a second mixed representation biased toward the second modality.

[0278] Step 606: Determine an adaptive first fusion weight according to the first mixed data representation, and determine an adaptive second fusion weight according to the second mixed data representation.

[0279] Specifically, the multimodal model may perform a linear transformation on the first mixed data representation and then perform activation processing to obtain a first fusion weight; and perform a linear transformation on the second mixed data representation and then perform activation processing to obtain a second fusion weight.

[0280] Step 608: According to the first fusion weight, the first data representation and the first mixed representation are fused to obtain a first fused representation; according to the second fusion weight, the second data representation and the second mixed representation are fused to obtain a second fused representation.

[0281] Specifically, the multimodal model can use the first fusion weight as the coefficient of the first mixed representation, and the difference between the value one and the first fusion weight as the coefficient of the first data representation; according to the coefficient of the first mixed representation and the coefficient of the first data representation, the first mixed representation and the first data representation are weightedly summed, and then decoded based on the result of the weighted summation to obtain the first predicted feature representation.

[0282] The multimodal model can use the second fusion weight as the coefficient of the second mixed representation, and the difference between the value one and the second fusion weight as the coefficient of the second data representation; according to the coefficient of the second mixed representation and the coefficient of the second data representation, the second mixed representation and the second data representation are weightedly summed, and then decoded based on the result of the weighted summation to obtain the second predicted feature representation.

[0283] Step 610: Determine the matching degree between the first data and the second data according to the first fused representation and the second fused representation.

[0284] In the above embodiment, for the first modality, the multimodal model will use the first modality as the main information, and at the same time integrate the information represented by the second data to obtain a first mixed representation, and then determine the fusion weight of the first modality side according to the first mixed representation, that is, the first fusion weight. In this way, the proportion of single modality information and mixed modality information referenced when decoding the first modality side can be controlled according to the first fusion weight. For the second modality, the initial model will use the second modality as the main information, and at the same time integrate the information represented by the first data to obtain a second mixed representation, and then determine the fusion weight of the first modality side according to the second mixed representation, that is, the second fusion weight. In this way, the proportion of single modality information and mixed modality information referenced when decoding the second modality side can be controlled according to the second fusion weight. In this way, different modality sides can be more flexible and adaptive when fusing mixed modality information.

[0285] In some embodiments, the first data includes image data, the second data includes text data, and the method further includes: determining the text retrieval results of the specified image data based on the matching degree between the specified image data and multiple different text data; and / or determining the image retrieval results of the specified text data based on the matching degree between the specified text data and multiple different image data.

[0286] Specifically, the specific method of determining text retrieval results based on matching can be to use text data whose matching degree meets the push conditions as text retrieval results. The specific method of determining image retrieval results based on matching can be to use image data whose matching degree meets the push conditions as image retrieval results.

[0287] Among them, the matching degree that meets the push conditions can specifically be the highest matching degree, or the matching degree is greater than a preset threshold, or the top N with the highest matching degree, etc., and this embodiment of the present application does not limit this.

[0288] In the above embodiment, the trained multimodal model can be used to perform image and text matching calculations, thereby helping to improve the quality and efficiency of image and text retrieval.

[0289] In a specific embodiment, the first modality is an image modality and the second modality is a text modality as an example for explanation: please refer to Figure 7, which is a structural block diagram of a multimodal model in one embodiment.

[0290] Take the model training stage as an example: As shown in Figure 7, LightVLP uses two separate encoders to model different modal information. For one of the modalities, such as image input, the model randomly masks a certain proportion (such as 50%) of patches. Then, the remaining patches form a sequence and are input into the image transformer to obtain an image representation (in the training stage, corresponding to the first sample data representation in the aforementioned embodiment). Correspondingly, for text input, the model randomly masks a certain proportion of characters. Then, the remaining characters form a sequence and are input into the text transformer to obtain a text representation (in the training stage, corresponding to the second sample data representation in the aforementioned embodiment).

[0291] The image and text representations are input into Cross Transformers, which process them and output a mixed-modal sample representation. A fusion weight p is then determined based on this mixed data representation. The image and mixed-modal sample representations are fused using the fusion weight p to produce a first fused sample representation. The text and mixed-modal sample representations are then fused using the fusion weight p to produce a second fused sample representation.

[0292] Furthermore, based on the first fused sample representation and the mask features on the image side, decoding is performed to obtain a representation of the masked block, and based on the second fused sample representation and the mask features on the text side, decoding is performed to obtain a representation of the masked character. This application introduces two traditional image / text unsupervised training tasks, namely the image reconstruction task and the text reconstruction task, to construct the first modality loss and the second modality loss. An image-text matching task is also introduced to construct the second loss. A contrastive learning task is also introduced to construct the third loss. Model training is achieved through joint task learning.

[0293] In some application scenarios, a trained multimodal model can be used as a vision-language pre-training model, or a large model. Based on the requirements of specific downstream tasks, the multimodal model can be fine-tuned to obtain a task model suitable for the task. Downstream tasks include generating text from images, generating images from text, generating image captions, and so on.

[0294] In some application scenarios, this multimodal model can be used to recover original data. For example, the input data is a first data and a second data in two modalities, where the first data is an image with missing data, and the second data is text with missing data. The multimodal model can then output the missing image and text.

[0295] In some embodiments, the present application also conducts an effect test on the multimodal model obtained by the above training. Please refer to Table 1. From the results presented in Table 1, it can be seen that the model of the present application has a good improvement in efficiency and effect compared with the baseline model.

[0296] Table 1 Experimental results of different models on operation speed and recall rate

[0297] Among them, TR refers to text retrieval tasks and IR refers to image retrieval tasks.

[0298] In addition, an ablation experiment was conducted to verify the effectiveness of the gated interaction strategy (gate part) in the multimodal model in this application. Please refer to Table 2 below for the experimental data:

[0299] Table 2 Ablation experiment results

[0300] It can be clearly seen from Table 2 above that the model of this application with the gated interaction strategy has a significant improvement in recall rate.

[0301] Finally, we conducted detailed tests on different mask ratios, and the conclusions are shown in Table 3:

[0302] Table 3 Test results of different mask ratios

[0303] It can be clearly seen from the above test results that the multimodal model provided in this application has advantages in flexibility and efficiency compared with the traditional VLP model, and the dual-coding structure is more suitable for image-text retrieval tasks.

[0304] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0305] Based on the same inventive concept, the present application also provides a data matching device for implementing the aforementioned data matching method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more data matching device embodiments provided below can be found in the above-mentioned limitations on the data matching method and will not be further elaborated here.

[0306] In an exemplary embodiment, as shown in FIG8 , a data matching device 800 is provided, comprising: an encoding module 801 , a determination module 802 , a fusion module 803 and a matching module 804 , wherein:

[0307] The encoding module is used to obtain first data, encode the first data, and obtain a first data representation; obtain second data, encode the second data, and obtain a second data representation, wherein the first data and the second data are data of different modes.

[0308] The determination module is configured to combine the first data representation and the second data representation to obtain a mixed data representation, and determine a fusion weight according to the mixed data representation.

[0309] The fusion module is used to fuse the first data representation and the mixed data representation according to the fusion weight to obtain a first fused representation, and to fuse the second data representation and the mixed data representation to obtain a second fused representation.

[0310] The matching module is used to determine the matching degree between the first data and the second data according to the first fused representation and the second fused representation.

[0311] In some embodiments, the determination module is further used to perform multiple interactive fusions based on the first data representation and the second data representation to obtain a first mixed representation biased towards the first modal side and a second mixed representation biased towards the second modal side; and combine the first mixed representation and the second mixed representation to obtain a mixed data representation.

[0312] In some embodiments, the determination module is also used to obtain the first intermediate mixed representation and the second intermediate mixed representation output of the i-th interactive fusion when performing the i+1-th interactive fusion; fuse the second intermediate mixed representation into the first intermediate mixed representation to obtain the first intermediate mixed representation of the i+1-th interactive fusion; fuse the first intermediate mixed representation into the second intermediate mixed representation to obtain the second intermediate mixed representation of the i+1-th interactive fusion; use i+1 as the new i, and return to the first intermediate mixed representation and the second intermediate mixed representation obtained by the i-th interactive fusion when performing the i+1-th interactive fusion and continue to execute until the stopping condition is reached; use the first intermediate mixed representation output of the last interactive fusion as the first mixed representation, and use the second intermediate mixed representation output of the last interactive fusion as the second mixed representation; wherein, i is a natural number greater than or equal to 0, when i is 0, the first intermediate mixed representation output of the i-th interactive fusion obtained is the first data representation, and the second intermediate mixed representation output of the i-th interactive fusion obtained is the second data representation.

[0313] In some embodiments, the determination module is further used to perform an attention operation based on the first intermediate mixed representation and the second intermediate mixed representation output by the i-th interactive fusion, and perform a layer regularization operation after combining the attention operation result with the first intermediate mixed representation output by the i-th interactive fusion to obtain a first intermediate processing feature; perform a forward neural network processing on the first intermediate processing feature, and perform a layer regularization operation after combining the forward neural network processing result with the first intermediate processing feature to obtain the first intermediate mixed representation of the i+1-th interactive fusion.

[0314] In some embodiments, the determination module is further used to perform an attention operation based on the second intermediate mixed representation output by the i-th interactive fusion and the first intermediate mixed representation, and perform a layer regularization operation after combining the attention operation result and the second intermediate mixed representation output by the i-th interactive fusion to obtain a second intermediate processing feature; perform a forward neural network processing on the second intermediate processing feature, and perform a layer regularization operation after combining the forward neural network processing result and the second intermediate processing feature to obtain a second intermediate mixed representation of the i+1-th interactive fusion.

[0315] In some embodiments, the determination module is further configured to perform linear transformation on the mixed data representation and then perform activation processing to obtain a fusion weight.

[0316] In some embodiments, the fusion module is also used to use the fusion weight as the coefficient of the mixed data representation, and the difference between the value one and the fusion weight as the coefficient of the first data representation; according to the coefficient of the mixed data representation and the coefficient of the first data representation, the mixed data representation and the first data representation are weighted to obtain a first fused representation.

[0317] In some embodiments, the fusion module is also used to use the fusion weight as the coefficient of the mixed data representation, and the difference between the value one and the fusion weight as the coefficient of the second data representation; according to the coefficient of the mixed data representation and the coefficient of the second data representation, the mixed data representation and the second data representation are weighted to obtain a first fusion representation.

[0318] In some embodiments, the mixed data representation includes a first mixed representation biased towards the first modality side and a second mixed representation biased towards the second modality side, and the fusion weight includes a first fusion weight determined based on the first mixed data representation and a second fusion weight determined based on the second mixed data representation; the fusion module is further used to fuse the first data representation and the first mixed representation according to the first fusion weight to obtain a first fused representation; and to fuse the second data representation and the second mixed representation according to the second fusion weight to obtain a second fused representation.

[0319] In some embodiments, the matching module is further used to decode according to the first fused representation to obtain a first feature representation of the first data; decode according to the second fused representation to obtain a second feature representation of the second data; and determine the degree of matching between the first data and the second data based on the first feature representation and the second feature representation.

[0320] In some embodiments, the first data includes image data, the second data includes text data, and the device also includes a retrieval module for determining the text retrieval results of the specified image data based on the matching degree between the specified image data and multiple different text data; and / or determining the image retrieval results of the specified text data based on the matching degree between the specified text data and multiple different image data.

[0321] Based on the same inventive concept, the present application also provides a multimodal model processing device for implementing the multimodal model processing method mentioned above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more multimodal model processing device embodiments provided below can be found in the above-mentioned limitations of the multimodal model processing method and will not be repeated here.

[0322] In an exemplary embodiment, as shown in FIG9 , a multimodal model processing apparatus 900 is provided, comprising an encoding module 901 , a determination module 902 , a fusion module 903 , a construction module 904 , and a training module 905 , wherein:

[0323] The encoding module is used to encode the first sample data and the second sample data in the training sample pair through the initial model to be trained to obtain the first sample data representation and the second sample data representation; the first sample data and the second sample data are sample data of different modalities.

[0324] The determination module is used to obtain a mixed sample data representation by combining the first sample data representation and the second sample data representation through an initial model, and to determine a prediction weight according to the mixed sample data representation.

[0325] The fusion module is used to fuse the first sample data representation and the mixed sample data representation according to the prediction weight through the initial model, and decode based on the fused result to obtain the first prediction feature representation.

[0326] The fusion module is also used to fuse the second sample data representation and the mixed sample data representation according to the prediction weight through the initial model, and decode based on the fused result to obtain the second prediction feature representation.

[0327] A construction module is used to determine a first loss based on the true feature representation and the first predicted feature representation of the first sample data, and the true feature representation and the second predicted feature representation of the second sample data; determine a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weights; and construct a target loss function based on the first loss and the second loss.

[0328] The training module is used to train the initial model to be trained based on the target loss function, and obtain the multimodal model after the training is completed.

[0329] In some embodiments, the apparatus further includes a masking module configured to obtain a training sample pair comprising first sample data and second sample data; split the first sample data into a plurality of first data blocks, split the second sample data into a plurality of second data blocks, extract a portion of the plurality of first data blocks to form a first input sequence, and extract a portion of the plurality of second data blocks to form a second input sequence. An encoding module is configured to perform multiple encoding operations on the first input sequence to obtain a first data representation, and perform multiple encoding operations on the second input sequence to obtain a second data representation.

[0330] In some embodiments, the determination module is further used to perform multiple interactive fusions of the first sample data representation and the second sample data representation through the initial model to obtain a first mixed sample representation biased towards the first modal side and a second mixed sample representation biased towards the second modal side; and combine the first mixed sample representation and the second mixed sample representation to obtain a mixed sample data representation.

[0331] In some embodiments, the construction module is further used to construct a first vector based on the alignment labels of the first sample data and the second sample data; construct a second vector based on the predicted weight and the difference between the numerical value one and the predicted weight; and determine the second loss based on the first vector and the second vector.

[0332] In some embodiments, the mixed sample data representation includes a first mixed sample representation biased toward the first modality and a second mixed sample representation biased toward the second modality, and the prediction weights include a first prediction weight determined based on the first mixed sample representation and a second prediction weight determined based on the second mixed sample representation. The fusion module is further configured to fuse the first sample data representation and the first mixed sample representation using the initial model according to the first prediction weight; and to fuse the second sample data representation and the second mixed sample representation using the initial model according to the second prediction weight.

[0333] In some embodiments, the mixed sample data representation includes a first mixed sample representation biased toward the first modal side and a second mixed sample representation biased toward the second modal side, and the prediction weight includes a first prediction weight determined based on the first mixed sample representation and a second prediction weight determined based on the second mixed sample representation; the construction module is also used to construct a first vector based on the alignment labels of the first sample data and the second sample data; determine the target weight based on the mean of the first prediction weight and the second prediction weight, and construct a third vector based on the target weight and the difference between the numerical value one and the target weight; determine the second loss based on the first vector and the third vector.

[0334] In some embodiments, the construction module is further used to determine the correlation between the first sample data and the second sample data based on the first sample data representation and the second sample data representation; determine the third loss based on the alignment labels of the first sample data and the second sample data, and the correlation between the first sample data and the second sample data; and construct a target loss function based on the first loss, the second loss, and the third loss.

[0335] In some embodiments, the correlation between the first sample data and the second sample data includes: the correlation of the first sample data relative to the second sample data, and the correlation of the second sample data relative to the first sample data; the construction module is further used to determine a preset number of third sample data with the same modality as the first sample data, and a preset number of fourth sample data with the same modality as the second sample data; calculate the correlation of the first sample data with the second sample data based on the first sample data representation, the second sample data representation, and the sample data representation of each fourth sample data; calculate the correlation of the second sample data with the first sample data based on the first sample data representation, the second sample data representation, and the sample data representation of each third sample data.

[0336] In some embodiments, the construction module is further used to determine a first value based on the first sample data representation and the second sample data representation; determine multiple second values ​​based on the first sample data representation and the sample data representation of each fourth sample data; and use the comparison value of the first value and the sum of the multiple second values ​​as the correlation of the second sample data relative to the first sample data.

[0337] In some embodiments, the device also includes a retrieval module for obtaining a text retrieval task, extracting the image data specified in the text retrieval task, and forming a first data pair to be processed by combining each text data in the text library with the specified image data; processing the first data pair through a multimodal model to obtain two feature representations, and determining the matching degree of the first data pair based on the two feature representations; and determining the target text that matches the specified image data based on the matching degree of each first data pair.

[0338] In some embodiments, the device also includes a retrieval module for obtaining an image retrieval task, extracting the text data specified in the image retrieval task, and forming a second data pair to be processed by respectively combining each image data in the image library with the specified text data; processing the second data pair through a multimodal model to obtain two feature representations, and determining the matching degree of the second data pair based on the two feature representations; and determining the target image that matches the specified text data according to the matching degree of each second data pair.

[0339] Each module in the above-mentioned apparatus may be implemented in whole or in part by software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.

[0340] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be shown in FIG10 . The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store graphic data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a data matching method and / or a multimodal model processing method is implemented.

[0341] Those skilled in the art will understand that the structure shown in FIG10 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0342] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0343] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0344] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0345] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0346] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0347] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0348] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A data matching method, executed by a computer device, comprising: Acquire first data, encode the first data, and obtain a first data representation; Acquire second data, encode the second data, and obtain a second data representation, wherein the first data and the second data have different modalities; combining the first data representation and the second data representation to obtain a mixed data representation, and determining a fusion weight according to the mixed data representation; According to the fusion weight, the first data representation and the mixed data representation are fused to obtain a first fused representation, and the second data representation and the mixed data representation are fused to obtain a second fused representation; and A matching degree between the first data and the second data is determined according to the first fused representation and the second fused representation.

2. The method according to claim 1, wherein combining the first data representation and the second data representation to obtain a hybrid data representation comprises: Performing multiple interactive fusions based on the first data representation and the second data representation to obtain a first mixed representation biased towards the first modality and a second mixed representation biased towards the second modality; The first mixed representation and the second mixed representation are combined to obtain a mixed data representation.

3. The method according to claim 2, wherein the performing multiple interactive fusions based on the first data representation and the second data representation to obtain a first mixed representation biased towards the first modality and a second mixed representation biased towards the second modality comprises: When performing the (i+1)th interactive fusion, obtaining a first intermediate mixed representation and a second intermediate mixed representation outputted by the (i)th interactive fusion; fusing the second intermediate mixed representation into the first intermediate mixed representation to obtain an (i+1)th interactively fused first intermediate mixed representation; Fusing the first intermediate mixed representation into the second intermediate mixed representation to obtain an (i+1)th interactively fused second intermediate mixed representation; i+1 is used as the new i, and the first intermediate mixed representation and the second intermediate mixed representation obtained by the i-th interactive fusion are returned when the i+1-th interactive fusion is performed, and the execution continues until the stopping condition is reached; The first intermediate mixed representation output from the last interactive fusion is used as the first mixed representation, and the second intermediate mixed representation output from the last interactive fusion is used as the second mixed representation; Among them, i is a natural number greater than or equal to 0. When i is 0, the first intermediate mixed representation obtained from the i-th interactive fusion output is the first data representation, and the second intermediate mixed representation obtained from the i-th interactive fusion output is the second data representation.

4. The method according to claim 3, wherein fusing the second intermediate mixed representation into the first intermediate mixed representation to obtain the (i+1)th interactively fused first intermediate mixed representation comprises: Perform an attention operation on the first intermediate mixed representation and the second intermediate mixed representation output by the i-th interactive fusion, combine the attention operation result with the first intermediate mixed representation output by the i-th interactive fusion, and then perform a layer regularization operation to obtain a first intermediate processing feature; The first intermediate processing feature is processed by a feedforward neural network, and a layer regularization operation is performed after combining the feedforward neural network processing result and the first intermediate processing feature to obtain a first intermediate mixed representation of the i+1th interactive fusion.

5. The method according to any one of claims 1 to 4, wherein determining a fusion weight according to the mixed data representation comprises: The mixed data representation is linearly transformed and then activated to obtain a fusion weight.

6. The method according to any one of claims 1 to 5, wherein the mixed data representation comprises a first mixed representation biased toward a first modality side and a second mixed representation biased toward a second modality side, and the fusion weight comprises a first fusion weight determined according to the first mixed data representation and a second fusion weight determined according to the second mixed data representation; The step of fusing the first data representation and the mixed data representation to obtain a first fused representation according to the fusion weight, and fusing the second data representation and the mixed data representation to obtain a second fused representation, comprises: According to the first fusion weight, the first data representation and the first mixed representation are fused to obtain a first fused representation; According to the second fusion weight, the second data representation and the second mixed representation are fused to obtain a second fused representation.

7. The method according to any one of claims 1 to 6, wherein determining the matching degree between the first data and the second data according to the first fused representation and the second fused representation comprises: Decoding according to the first fused representation to obtain a first feature representation of the first data; Decoding according to the second fused representation to obtain a second feature representation of the second data; Based on the first feature representation and the second feature representation, a matching degree between the first data and the second data is determined.

8. The method according to any one of claims 1 to 7, wherein the first data comprises image data, and the method further comprises: Based on the matching degrees between the designated image data and a plurality of different text data, a text retrieval result of the designated image data is determined.

9. The method according to any one of claims 1 to 8, wherein the second data comprises text data, and the method further comprises: Based on the matching degrees between the designated text data and a plurality of different image data, an image retrieval result of the designated text data is determined.

10. A multimodal model processing method, executed by a computer device, the method comprising: Acquire a training sample pair, wherein the training sample pair includes first sample data and second sample data; The first sample data and the second sample data are sample data of different modalities; Encoding based on the first sample data using the initial model to be trained to obtain a first sample data representation; Encoding is performed based on the second sample data to obtain a second sample data representation; Obtaining a mixed sample data representation by combining the first sample data representation and the second sample data representation through the initial model, and determining a prediction weight according to the mixed sample data representation; fusing the first sample data representation and the mixed sample data representation according to the prediction weight through the initial model, and performing decoding based on the fused result to obtain a first prediction feature representation; fusing the second sample data representation and the mixed sample data representation according to the prediction weight through the initial model, and performing decoding based on the fused result to obtain a second prediction feature representation; Determine a first loss based on the true feature representation of the first sample data and the first predicted feature representation, and the true feature representation of the second sample data and the second predicted feature representation; Determining a second loss based on the alignment labels of the first sample data and the second sample data, and the prediction weight; and A target loss function is constructed according to the first loss and the second loss, and the initial model to be trained is trained based on the target loss function, and a multimodal model is obtained after the training is completed.

11. The method according to claim 10, before encoding the first sample data and the second sample data in the training sample pair, the method further comprises: Acquire a training sample pair, wherein the training sample pair includes first sample data and second sample data; The first sample data is split into a plurality of first data blocks, and the second sample data is split into a plurality of second data blocks, Extracting some data blocks from the plurality of first data blocks to form a first input sequence, and extracting some data blocks from the plurality of second data blocks to form a second input sequence; encoding based on the first sample data to obtain a first sample data representation; Based on the second sample number The second sample data is encoded to obtain a second sample data representation, including: Performing multiple encodings on the first input sequence to obtain a first data representation; The second input sequence is encoded multiple times to obtain a second data representation.

12. The method according to claim 10 or 11, wherein determining the second loss based on the alignment labels of the first sample data and the second sample data and the prediction weight comprises: Constructing a first vector based on the alignment labels of the first sample data and the second sample data; constructing a second vector according to the prediction weight and the difference between the value one and the prediction weight; A second loss is determined based on the first vector and the second vector.

13. The method according to any one of claims 10 to 12, wherein the mixed sample data representation comprises a first mixed sample representation biased toward a first modality and a second mixed sample representation biased toward a second modality, and the prediction weight comprises a first prediction weight determined according to the first mixed sample representation and a second prediction weight determined according to the second mixed sample representation; The fusing the first sample data representation and the mixed sample data representation according to the prediction weight through the initial model includes: fusing the first sample data representation and the first mixed sample representation through the initial model according to the first prediction weight; The fusing the second sample data representation and the mixed sample data representation according to the prediction weight through the initial model includes: The second sample data representation and the second mixed sample representation are fused according to the second prediction weight through the initial model.

14. The method according to any one of claims 10 to 13, further comprising: determining a correlation between the first sample data and the second sample data according to the first sample data representation and the second sample data representation; determining a third loss based on the alignment labels of the first sample data and the second sample data, and the correlation between the first sample data and the second sample data; The constructing a target loss function according to the first loss and the second loss includes: An objective loss function is constructed according to the first loss, the second loss, and the third loss.

15. The method according to claim 14, wherein the correlation between the first sample data and the second sample data comprises: a correlation degree of the first sample data with respect to the second sample data, and a correlation degree of the second sample data with respect to the first sample data; Determining the correlation between the first sample data and the second sample data according to the first sample data representation and the second sample data representation includes: Determine a preset number of third sample data having the same modality as the first sample data, and a preset number of fourth sample data having the same modality as the second sample data; Calculating the correlation of the first sample data with respect to the second sample data according to the first sample data representation, the second sample data representation, and the sample data representations of each fourth sample data; The correlation of the second sample data with respect to the first sample data is calculated based on the first sample data representation, the second sample data representation, and the sample data representations of each third sample data.

16. The method according to claim 15, wherein calculating the correlation between the first sample data and the second sample data according to the first sample data representation, the second sample data representation, and the sample data representations of each fourth sample data comprises: determining a first value based on the first sample data representation and the second sample data representation; determining a plurality of second values ​​according to the first sample data representation and the sample data representations of each fourth sample data; A comparison value between the first value and the sum of the plurality of second values ​​is used as a correlation degree between the first sample data and the second sample data.

17. The method according to any one of claims 10 to 16, wherein the modality comprises an image modality and a text modality, and the method further comprises: Acquire a text retrieval task, extract the image data specified in the text retrieval task, and form a first data pair to be processed with each text data in the text library and the specified image data; Processing the first data pair by using the multimodal model to obtain two feature representations, and determining a matching degree of the first data pair based on the two feature representations; According to the matching degree of each first data pair, a target text matching the specified image data is determined.

18. The method according to any one of claims 10 to 17, wherein the modality comprises an image and text, and the method further comprises: Acquire an image retrieval task, extract text data specified in the image retrieval task, and form a second data pair to be processed with each image data in the image library and the specified text data; Processing the second data pair by using the multimodal model to obtain two feature representations, and determining a matching degree of the second data pair based on the two feature representations; According to the matching degree of each second data pair, a target image matching the specified text data is determined.

19. A data matching device, comprising: An encoding module, used for acquiring first data, encoding the first data, and obtaining a first data representation; Acquire second data, encode the second data, and obtain a second data representation, wherein the first data and the second data are data of different modalities; a determination module, configured to combine the first data representation and the second data representation to obtain a mixed data representation, and determine a fusion weight according to the mixed data representation; a fusion module, configured to fuse the first data representation and the mixed data representation to obtain a first fused representation, and to fuse the second data representation and the mixed data representation to obtain a second fused representation according to the fusion weight; and A matching module is used to determine a matching degree between the first data and the second data according to the first fusion representation and the second fusion representation.

20. A multimodal model processing device, the device comprising: An encoding module, used to obtain a training sample pair, wherein the training sample pair includes a first sample data and a second sample data; The first sample data and the second sample data are sample data of different modalities; encoding is performed based on the first sample data through the initial model to be trained to obtain a first sample data representation; Encoding is performed based on the second sample data to obtain a second sample data representation; a determination module, configured to obtain a mixed sample data representation by combining the first sample data representation and the second sample data representation through the initial model, and determine a prediction weight according to the mixed sample data representation; A fusion module, configured to fuse the first sample data representation and the mixed sample data representation according to the prediction weight through the initial model, and perform decoding based on the fused result to obtain a first prediction feature representation; The fusion module is further used to fuse the second sample data representation and the mixed sample data representation according to the prediction weight through the initial model, and perform decoding based on the fused result to obtain a second prediction feature representation; A construction module is used to determine a first loss based on the real feature representation of the first sample data and the first predicted feature representation, and the real feature representation of the second sample data and the second predicted feature representation; determine a second loss based on the alignment label of the first sample data and the second sample data, and the prediction weight; and construct a target loss function according to the first loss and the second loss; and The training module is used to train the initial model to be trained based on the target loss function, and obtain a multimodal model after the training is completed.

21. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 18 when executing the computer program.

22. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 18.

23. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 18 are implemented.

Citation Information

Patent Citations

  • Cross-modal information retrieval method and device and storage medium

    CN109816039A

  • Cross-modal information retrieval method and device

    CN113032614A

  • Image-text retrieval method based on text generation and iterative matching

    CN116842201A

  • Image analysis method and system based on multi-modal information

    CN116994069A

  • Text Based Image Search

    US20220343626A1