Methods, apparatuses, devices, and media for multi-modal data processing

By alternately deploying feature extraction model architectures with cross-modal coding and visual coding components, the problem of modal semantic alignment in image-text matching is solved, achieving more efficient feature extraction and more accurate matching results.

CN115982596BActive Publication Date: 2026-07-21FACE CUTE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FACE CUTE CO LTD
Filing Date
2023-01-04
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the semantic alignment problem between different modalities in image-text matching tasks, especially the asynchronous semantic alignment caused by the difference in information density between image and text data.

Method used

A feature extraction model architecture with alternating deployment of cross-modal coding and visual coding components is adopted. By processing image and text data through alternating cross-modal coding and visual coding components, asynchronous cross-modal semantic alignment is achieved, preserving dense feature coding of image data while reducing dense coding of text data.

Benefits of technology

It improves performance in image-text matching tasks, enhances the accuracy and efficiency of feature extraction, reduces computational overhead, and strengthens the model's generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115982596B_ABST
    Figure CN115982596B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide methods, apparatuses, devices and media for multi-modal data processing. The method comprises: obtaining image data and text data; and extracting target visual features of the image data and target text features of the text data using a feature extraction model. The feature extraction model comprises alternately deployed cross-modal encoding parts and visual encoding parts. The extracting comprises: performing cross-modal feature encoding on first intermediate visual features and first intermediate text features using a first cross-modal encoding part to obtain second intermediate visual features and second intermediate text features; performing visual modal feature encoding on the second intermediate visual features using a first visual encoding part to obtain third intermediate visual features. Such feature extraction can capture non-synchronous semantic alignment of image and text modalities, enabling the model to learn and achieve accurate feature extraction of image and text modalities more quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments of this disclosure generally relate to machine learning, and more particularly to methods, apparatuses, devices, and computer-readable storage media for multimodal data processing. Background Technology

[0002] Image-text matching is a typical task in the fields of vision and language, involving the processing of data from different modalities. Image data can include dynamic images, such as videos, as well as static images, such as single images. Image-text matching can be used, for example, to retrieve images from text or text from images. The main challenge of this task is aligning semantics across different modalities. In recent years, pre-training or training models from large-scale video and text content has become a trend. The modeling process can uncover sufficient cross-modal cues to achieve specific tasks. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for multimodal data processing is provided. The method includes: acquiring image data and text data; and extracting target visual features from the image data and target text features from the text data using a feature extraction model, the feature extraction model including alternately deployed cross-modal coding and visual coding parts. The extraction includes: performing cross-modal feature coding on a first intermediate visual feature of the image data and a first intermediate text feature of the text data using a first cross-modal coding part of the feature extraction model to obtain a second intermediate visual feature and a second intermediate text feature; performing visual modal feature coding on the second intermediate visual feature using a first visual coding part of the feature extraction model to obtain a third intermediate visual feature; performing cross-modal feature coding on the third intermediate visual feature and the second intermediate text feature using a second cross-modal coding part of the feature extraction model to obtain a fourth intermediate visual feature and a third intermediate text feature; and determining target visual features and target text features based on the fourth intermediate visual feature and the third intermediate text feature.

[0004] In a second aspect of this disclosure, an apparatus for multimodal data processing is provided. The apparatus includes: an acquisition module configured to acquire image data and text data; and an extraction module configured to extract target visual features from the image data and target text features from the text data using a feature extraction model, the feature extraction model including alternately deployed cross-modal coding and visual coding portions. The extraction module includes: a first cross-modal coding module configured to perform cross-modal feature coding on a first intermediate visual feature of the image data and a first intermediate text feature of the text data using the first cross-modal coding portion of the feature extraction model to obtain a second intermediate visual feature and a second intermediate text feature; a first visual coding module configured to perform visual modal feature coding on the second intermediate visual feature using the first visual coding portion of the feature extraction model to obtain a third intermediate visual feature; a second cross-modal coding module configured to perform cross-modal feature coding on the third intermediate visual feature and the second intermediate text feature using the second cross-modal coding portion of the feature extraction model to obtain a fourth intermediate visual feature and a third intermediate text feature; and a target feature determination module configured to determine target visual features and target text features based on the fourth intermediate visual feature and the third intermediate text feature.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The medium stores a computer program that, when executed by a processor, implements the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram of an example data processing environment in which embodiments of the present disclosure can be implemented is shown;

[0010] Figure 2 A schematic diagram illustrates a model training and application environment in which embodiments of the present disclosure can be implemented;

[0011] Figures 3A to 3C A schematic diagram of an example model architecture for multimodal feature extraction is shown;

[0012] Figure 4 A schematic diagram illustrating an example of the information density difference between image data and text data;

[0013] Figure 5 A schematic diagram of an example structure of a feature extraction model according to some embodiments of the present disclosure is shown;

[0014] Figure 6 A simplified architecture of a feature extraction model according to some embodiments of the present disclosure is shown in the diagram;

[0015] Figure 7 The diagram illustrates some example deployments of different coding portions in a feature extraction model according to some embodiments of the present disclosure;

[0016] Figure 8 A schematic diagram illustrating an example masking method for sample image data according to some embodiments of the present disclosure is shown;

[0017] Figure 9 A flowchart illustrating a process for multimodal data processing according to some implementations of this disclosure is shown;

[0018] Figure 10 A block diagram of an apparatus for multimodal data processing according to some implementations of this disclosure is shown; and

[0019] Figure 11 An electronic device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0020] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the degree of matching between various data. For example, the aforementioned degree of matching can be obtained based on a variety of currently known and / or future-developed technical solutions.

[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0024] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0025] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0027] As used in this paper, the term "model" refers to a model that learns the degree of matching between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0028] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0029] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0030] Figure 1 A schematic diagram of an example data processing environment 100 in which embodiments of the present disclosure can be implemented is shown. Figure 1 In environment 100, the cross-modal data processing system 110 processes pairs of image data 102 and text data 104. Here, image data 102 can include dynamic image data or still image data. Dynamic image data is, for example, video, where each video frame can be considered a single image. Still image data is a single image. Figure 1In the examples below, image data 102 is shown as video data, which includes multiple video frames. However, it should be understood from the following description that embodiments of this disclosure are equally applicable to still image data.

[0031] Many applications require processing both image and text modal data. For example, some applications involve image and text matching tasks. Such tasks include retrieving images from text, retrieving text from images, video / image question answering (finding answers to questions from videos or images), and so on. To accomplish tasks involving multimodal data, machine learning techniques can be relied upon to provide feature extraction models to extract features from both image and text data. Features can be vectors with a specific dimension, also known as feature representations, feature vectors, feature codes, etc., these terms are used interchangeably in this paper. The extracted features can represent the corresponding data in a specific dimensional space.

[0032] exist Figure 1 In this system, the cross-modal data processing system 110 utilizes a feature extraction model 120 to extract target visual features from image data 102 and target text features from text data 104. Depending on the specific task, the target visual features and target text features are provided to the output layer 130 to determine the task output. In a matching task, the output layer 130 is configured to determine whether image data 102 and text data 104 match each other based on the target visual features and target text features. Matching means that text data 104 accurately describes the information expressed by image data 102. For example, in... Figure 1 In the example, text data 104 is the English sentence "Yellow clown fish dart through coral," which accurately describes the information presented by image data 102 in video form. If feature extraction model 120 can extract visual and text features that are aligned or matched with each other, then the correct matching result can be determined based on these two features.

[0033] Notice, Figure 1 The example image and text data provided are for illustrative purposes only and should not be construed as limiting the scope of embodiments of this disclosure to any particular form of text and image data.

[0034] As can be seen, extracting feature representations that accurately characterize each modality of data from image and text data is a crucial task in cross-modal data processing. The architecture of the feature extraction model affects its feature extraction capabilities. Embodiments of this disclosure propose an improved architecture for the feature extraction model, which enhances feature extraction from multimodal data and improves performance in various tasks involving multimodal data.

[0035] In some embodiments, the training process of the feature extraction model 120 may include a pre-training process and a fine-tuning process. Large-scale pre-trained models typically possess strong generalization capabilities and efficient utilization of large-scale data. After pre-training the model on large-scale data, the pre-trained model can be fine-tuned using a small amount of data based on the specific needs of different downstream tasks. This can significantly improve the overall model learning efficiency and reduce the need for labeled data for specific downstream tasks. The trained feature extraction model 120 can then be provided for use in specific application scenarios.

[0036] Figure 2 A schematic diagram is shown of a model training and application environment 200 in which embodiments of the present disclosure can be implemented. Figure 2 The environment 200 illustrates three distinct phases of the model: a pre-training phase 202, a fine-tuning phase 204, and an application phase 206. A testing phase, not shown in the figure, may also occur after the pre-training or fine-tuning phases.

[0037] In the pre-training phase 202, the model pre-training system 210 is configured to pre-train the feature extraction model 120. At the start of pre-training, the feature extraction model 120 may have initial parameter values. The pre-training process involves updating the parameter values ​​of the feature extraction model 120 to the desired values ​​based on the training data.

[0038] The training data used in pre-training includes sample image data 212 and sample text data 214, and may also include annotation information 216. Annotation information 216 can be used to indicate whether the sample image data 212 and sample text data 214 input to the feature extraction model 120 match. Although a pair of sample images and text is shown, a large number of sample images and text may be used for training during the pre-training phase. During pre-training, one or more pre-training tasks 207-1, 207-2, etc., can be designed. Pre-training tasks are used to help update the parameters of the feature extraction model 120. Some pre-training tasks may perform parameter updates based on annotation information 216.

[0039] During the pre-training phase 202, the feature extraction model 120 can learn strong generalization capabilities through a large amount of training data. After pre-training, the parameter values ​​of the feature extraction model 120 have been updated and have the pre-trained parameter values. The pre-trained feature extraction model 120 can extract feature representations of the input data relatively accurately.

[0040] The pre-trained feature extraction model 120 can be provided to the fine-tuning stage 204, where it is fine-tuned by the model fine-tuning system 220 for different downstream tasks. In some embodiments, depending on the downstream task, the pre-trained feature extraction model 120 can be connected to a corresponding task-specific output layer 227 to construct a downstream task model 225. This is because the required output may differ for different downstream tasks.

[0041] In the fine-tuning phase 204, the parameter values ​​of the feature extraction model 120 are further adjusted using the training data. If necessary, the parameters of the task feature output layer 227 may also be adjusted. The training data used in the fine-tuning phase includes sample image data 222 and sample text data 224, and may also include annotation information 226. The annotation information 226 can be used to indicate whether the sample image data 222 and sample text data 224 input to the feature extraction model 120 match. Although a pair of sample images and text is shown, the fine-tuning phase may use a certain amount of sample images and text for training. The feature extraction model 120 can perform feature representation extraction on the input image data and text data and provide it to the task-specific output layer 227 to provide the corresponding task output.

[0042] During fine-tuning, the corresponding training algorithm is also used to update and adjust the parameters of the overall model. Since the feature extraction model 120 has learned a lot of knowledge from the training data in the pre-training stage, a downstream task model that meets expectations can be obtained using a small amount of training data in the fine-tuning stage 204.

[0043] In some embodiments, during the pre-training phase 202, one or more task-specific output layers may be constructed based on the objective of the pre-training task to pre-train the feature extraction model 120 across multiple downstream tasks. In this case, if the task-specific output layer used in the downstream task is the same as the task-specific output layer constructed during pre-training, the pre-trained feature extraction model 120 and the task-specific output layer can be directly used to form the corresponding downstream task model. In this case, the downstream task model may not require fine-tuning, or may only require fine-tuning with a small amount of training data.

[0044] In application phase 206, the obtained downstream task model 225, with trained parameter values, can be provided to the model application system 230 for use. In application phase 206, the downstream task model 225 can be used to process corresponding inputs in the real-world scenario and provide corresponding outputs. For example, the feature extraction model 120 in the downstream task model 225 receives input target image data 232 and target text data 234 to extract corresponding target visual features and target text features. The extracted target visual features and target text features are provided to the task feature output layer 227 to determine the output of the corresponding task. Typically, this output can be summarized as determining whether the target image data 232 and the target text data 234 match or the degree of matching.

[0045] exist Figure 1 and Figure 2 In this system, the cross-modal data processing system 110, the model pre-training system 210, the model fine-tuning system 220, and the model application system 230 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0046] It should be understood that Figure 1 and Figure 2 The components and arrangements shown in environments 100 and 200 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, the model pre-training system 210, the model fine-tuning system 220, and the model application system 230 may be integrated in the same system or device. Implementations of this disclosure are not limited in this respect.

[0047] In some embodiments, the training phase of the feature extraction model 120 may not be divided into... Figure 2 Instead of the pre-training and fine-tuning stages shown, it can directly build downstream task models based on the task and use a large amount of training data to train the feature extraction model.

[0048] The above discussion covered some example environments for feature extraction from image and text modalities. To handle data of different modalities, the architecture of the feature extraction model needs to be specifically designed. Several approaches have proposed example architectures for feature extraction models.

[0049] Figures 3A to 3C A schematic diagram of an example model architecture for multimodal feature extraction is shown. Figure 3A The dual-stream processing-based architecture 301 shown includes independently parallel visual encoding sections 310 and text encoding sections 320. The number of visual encoding sections 310 and text encoding sections 320 is set to N, where N can be an integer greater than or equal to 1. Each visual encoding section 310 includes multiple visual encoding units for extracting visual features from image data. Each text encoding section 320 includes multiple text encoding units for processing text features from text data.

[0050] Figure 3B The hybrid-stream-based architecture 302 shown includes N single-modal coding sections 330, which can be similar to... Figure 3A Architecture 301 includes independent and parallel visual encoding and text encoding parts for independently encoding visual and text features. Architecture 302 also includes N' cross-modal encoding parts 340 for unified fusion, where N' and N can be the same or different. The cross-modal encoding parts 340 are used to perform cross-modal encoding on the text and visual features received from the separate modal encoding parts 330.

[0051] Figure 3C The unified stream processing-based architecture 303 shown includes N cross-modal coding parts 350, which are similar to Figure 3B The cross-modal coding section 340. Image data and text data are directly input into the cross-modal coding section 350 to perform cross-modal coding, outputting visual features and text features.

[0052] Dual-stream processing architectures can independently encode each modality, but their ability to achieve cross-modal semantic alignment is limited. Hybrid-stream processing architectures add unified fusion to independent dual-stream processing to fuse information from the two modalities, thereby achieving cross-modal alignment capabilities. However, hybrid-stream processing architectures have excessively high computational overhead. Unified-stream processing architectures rely on a single processing stream to jointly encode the two modalities, which effectively achieves training convergence and increases cross-modal alignment capabilities. Although unified-stream processing architectures are relatively lightweight, they still require significant computational overhead when training on large-scale data, and the model generalizes poorly, performing poorly on some tasks. The inventors discovered through research and analysis that this may be due to the modeling of dense cross-modal interactions between image and text data.

[0053] The purpose of cross-modal coding is to correlate semantic information between two modalities. However, the problem lies in the fact that image data and text data have different information densities. Generally, image data, including video and still data, typically has a natural, continuous signal and high spatial redundancy. Furthermore, video data also has high temporal and spatial redundancy. Therefore, image data usually needs to be encoded into hierarchical features using a large model. However, text, especially natural language text, is discrete and has refined semantics. Based on this fact, the inventors discovered that the feature extraction process for these two modalities will exhibit significant differences.

[0054] Figure 4 This diagram illustrates an example of the information density difference between image data and text data. Assume a visual coding model 400 is used to extract feature information from image data 102. The visual coding model 400 is constructed with multiple processing layers from low to high levels. Each layer processes the input and provides the extracted intermediate feature representations to the next layer for further processing. Observation reveals that the lower processing layers of the visual coding model 400 can perceive shallow information in image data 102, such as the color information "Yellow". As processing deepens, the visual coding model 400 can perceive information such as contours in image data 102, such as identifying the object "fish". In higher-level processing layers, through temporal exploration of multiple video frames of the dynamic image data 102, the visual coding model 400 can perceive the temporal information of image data 102, namely the motion information of objects in the image data, "dart through".

[0055] Therefore, due to the temporal and spatial redundancy of image data, the process of constructing visual features from image data has discrete levels (from low to high). However, textual features of text data are highly abstract, and there is no such low-to-high discretization extraction. Therefore, during the feature extraction process, the semantic granularity evolution of each modality is not synchronous, which is called asynchronous semantic alignment.

[0056] Considering this asynchronous semantic alignment, embodiments of this disclosure propose a feature extraction model architecture for asynchronous cross-modal semantic alignment to perform feature extraction from image and text data. Specifically, this feature extraction model is constructed to include a sparse cross-modal coding component. Furthermore, this feature extraction model preserves the denser feature encoding for image data while truncating the dense encoding for text data. This feature extraction process can better capture the asynchronous semantic alignment of image and text modalities, enabling the model to learn and achieve accurate feature extraction for both image and text modalities more quickly. The extracted features can accurately represent the corresponding image and text data, and therefore can be applied to various downstream tasks related to images and text.

[0057] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0058] Figure 5 A schematic diagram of an example structure of a feature extraction model 120 according to some embodiments of the present disclosure is shown. For illustrative purposes only, it is still assumed that the feature extraction model 120 is to process the image data 102 and text data 104 shown in the figure. Figure 2 As can be understood from the example environment, the image and text data input to the feature extraction model 120 may differ at different stages of pre-training, fine-tuning, and application. In some embodiments, the training phase of the feature extraction model 120 may be integrated and does not need to be divided into... Figure 2 The pre-training and fine-tuning phases are shown.

[0059] As analyzed above, due to the different information densities of image and text data, dense cross-modal interactions (e.g., architecture 303 based on unified stream processing) not only interfere with the semantic alignment of the two modalities but also increase training overhead, including pre-training overhead. In embodiments of this disclosure, an improved feature extraction model architecture is proposed. Figure 5 As shown, the feature extraction model 120 is constructed as alternating deployments of a cross-modal coding part 510 and a visual coding part 520 to enable asynchronous interaction between images and text.

[0060] "Alternating deployment" means that one cross-modal coding section 510 is connected to a visual coding section 520, which in turn is connected to the next cross-modal coding section 510, and so on. For example, in Figure 5 In this process, the cross-modal coding section 510-1 is connected to the visual coding section 520-1, which is then connected to another cross-modal coding section 510-2. The cross-modal coding section 510-2 is connected to the visual coding section 520-2, which is then connected to another cross-modal coding section 510-3, and so on, until the final cross-modal coding section 510-4.

[0061] Each cross-modal coding section 510 (cross-modal coding sections 510-1, 510-2, 510-3, ..., 510-4, etc.) is configured to perform cross-modal feature encoding on image and text data. Each visual coding section 520 (cross-modal coding sections 520-1, 520-2, 520-3, ...) is configured to perform single-modal feature encoding on the visual modality of the image data. The feature extraction model 120 may not include separate text modality encoding. In this way, the feature encoding result of the previous cross-modal coding section 510 on the text data is directly input into the next cross-modal coding section 510 for further processing. Overall, in the feature extraction process of the feature extraction model 120, image data is encoded more densely, while the cross-modal interaction between images and text is relatively sparse. This asynchronous coding method can solve the feature extraction differences caused by the difference in information density between image and text data.

[0062] Specifically, assuming that image data 102 is dynamic image data, it is represented as Where T represents the duration, such as the number of video frames; H and W represent the height and width of a single video frame; 3 represents the number of color channels in the video frame, which depends on the color space used and can also take other values. Text data 104 is represented as... Where L represents the text length.

[0063] In the initial stage, image data 102 can be converted into initial visual features 502, and text data 104 can be converted into initial text features 504, thus transforming the data into a multi-dimensional vector form for model processing. Initial visual features 502 and initial text features 504 are also called feature embeddings or embedding representations. This feature transformation is represented as follows:

[0064]

[0065] in () represents the feature transformation process of image data. () indicates the feature transformation process of text data. This can be achieved using an image embedding model. In some embodiments, the individual images (or video frames) in the image data 102 can be divided into multiple visual blocks, and each visual block can be converted into a feature embedding form. For example, the image data V can be divided into HW / P visual blocks with a resolution of PxP. In such an implementation, It can be implemented as a visual block embedding model. This can be achieved through word embedding models.

[0066] In some embodiments, the dimensions of the initial visual feature 502 and the initial text feature 504 can be configured to be the same. For example, the transformed initial visual feature 502 is represented as... The initial text feature 504 is represented as

[0067] Initial visual features 502 and initial text features 504, as representations of image data 102 and text data 104, are input into feature extraction model 120 for further processing. As mentioned above, in feature extraction model 120, some processing parts (i.e., cross-modal coding part 510) are configured to perform cross-modal feature coding of image and text modalities, while other parts (i.e., visual coding part 520) are configured to perform visual modal feature coding of image modalities.

[0068] In this paper, the feature encoding results of the cross-modal coding section 510 for the image modality are referred to as intermediate visual features, and the feature encoding results for the text modality are referred to as intermediate text features. Similarly, the feature encoding results of the visual coding section 520 for the image modality are also referred to as intermediate visual features. Thus, the intermediate visual features extracted by the current cross-modal coding section 510 are provided to the connected visual coding section 520 for further visual modality feature encoding to obtain additional intermediate visual features; the intermediate text features extracted by the current visual coding section 520 and the intermediate visual features output by the visual coding section 520 are provided to the next visual coding section 520 for processing. For the first processing section of the feature extraction model 120, its input is the initial visual feature 502 (and the initial text feature 504, if the first processing section is a cross-modal coding section). The aforementioned process is repeated iteratively until the last section of the feature extraction model 120 is reached. The features output by the last section are considered as the target visual features of the image data 102 and the target text features of the text data 104.

[0069] For the cross-modal coding part 510, its inputs (e.g., intermediate visual features and intermediate visual features, or initial visual features and initial visual features) are concatenated as Used for processing.

[0070] In some embodiments, the cross-modal encoding portion 510 may include one or more network layers, and the visual encoding portion 520 may also include one or more network layers. In some embodiments, the cross-modal encoding portion 510 and / or the visual encoding portion 520 may include transformer layers. The transformer layers may include multi-head self-attention (MSA) blocks and feedforward network (FFN) blocks. Of course, only some example implementations of the cross-modal encoding portion 510 and the visual encoding portion 520 are given here. In practical applications, the network layers used by the cross-modal encoding portion 510 and / or the visual encoding portion 520 can be configured according to actual needs. The cross-modal encoding portion 510 and the visual encoding portion 520 can also be configured with different types of network layers. Furthermore, different cross-modal encoding portions 510 can also use different types of network layers to implement cross-modal feature encoding. Similarly, the visual encoding portion 520 can also use different types of network layers to implement visual modal feature encoding.

[0071] Based on the deployment method of feature extraction model 210 described above, the processing of each network layer in feature extraction model 210 can be defined as follows:

[0072]

[0073] Where 1≤n≤N, and N represents the total number of network layers in the feature extraction model 210; [·,·] represents the feature concatenation of two modalities; Represents the intermediate visual features at the nth network layer; Φ represents the intermediate text features at the nth network layer; n This indicates the feature encoding processing of the nth network layer. From equation (2) above, it can be seen that if the nth network layer in the feature extraction model 210 belongs to the cross-modal coding part 510, this network layer processes the cascaded intermediate visual features and intermediate text features from the previous layer. If the nth network layer in the feature extraction model 210 belongs to the visual encoding part 520, this network layer only processes intermediate visual features from the previous layer. And the features of intermediate text No action will be taken.

[0074] The alternating deployment of the cross-modal coding section 510 and the visual coding section 520 can also be varied. Although Figure 5 The diagram shows a cross-modal coding section 510 deployed as the first processing section, but visual coding section 520 can also be deployed as the first processing section. In some embodiments, the last processing section of feature extraction model 120 can be deployed as cross-modal coding section 510, although such a constraint is not required in other embodiments.

[0075] The cross-modal coding part 510 and the visual coding part 520 can each be configured to have one or more network layers, i.e., different network depths. The network layers in the cross-modal coding part 510 can be called cross-modal coding layers, and the network layers in the visual coding part 520 can be called visual coding layers.

[0076] As mentioned above, in the feature extraction model 120, the cross-modal coding portion 510 and the visual coding portion 520 are deployed alternately. In some embodiments, the feature extraction model 120 may include multiple pairs of alternating cross-modal coding portions 510 and visual coding portions 520. Figure 6 A simplified architecture diagram of a feature extraction model 120 according to some embodiments of the present disclosure is shown. Figure 6 As shown, the feature extraction model 120 may include multiple processing parts 610, each processing part 610 including one or more (s) visual coding layers 622 and one or more cross-modal coding layers 612.

[0077] In order to Figures 3A to 3C Compared to the feature extraction architecture shown, it can be assumed that, based on achieving the same depth of feature extraction on image data, feature extraction model 120 can be deployed with There are 610 processing parts, and s can be set to be much smaller than N. Of course, in practical applications, the alternating deployment of the various cross-modal coding parts 510 and visual coding parts 520 in the feature extraction model 120 can also have other methods.

[0078] Assuming the feature extraction model 120 has N network layers in total, some of these layers can be configured as cross-modal coding layers, while others can be configured as visual coding layers. Compared to visual coding layers, cross-modal coding layers require more parameters, resulting in relatively higher training and application costs.

[0079] Figure 7 The diagram illustrates some example deployments of different coding portions in a feature extraction model 120 according to some embodiments of the present disclosure. In some embodiments, some network layers may be randomly selected and configured as cross-modal coding layers based on a random scheme. Figure 7As shown, based on the random scheme 701, network layers 1 and 2 in the feature extraction model 120 are selected as cross-modal coding layers 712 (forming one cross-modal coding part); network layer 5 is selected as a cross-modal coding layer 712 (forming another cross-modal coding part); network layers 7 and 8 are selected as cross-modal coding layers 712 (forming yet another cross-modal coding part); and network layer 10 is selected as a cross-modal coding layer 712 (forming yet another cross-modal coding part). In the remaining network layers, network layers 3 and 4 are deployed as visual coding layers 722 (forming one visual coding part), network layer 6 is deployed as a visual coding layer 722 (forming another visual coding part), and network layer 9 is also deployed as a visual coding layer 722 (forming yet another visual coding part).

[0080] In some embodiments, in order to obtain high-quality text and visual feature extraction with low overhead, in addition to random schemes, cross-modal coding parts and visual coding layers can be deployed according to some predetermined criteria.

[0081] In some embodiments, cross-modal coding layers can be deployed at predetermined intervals in feature extraction model 120 based on a unified scheme. The predetermined interval can be quantized as a predetermined number of visual coding layers. That is, based on the unified scheme, the visual coding portion between two adjacent cross-modal coding portions can include a predetermined number of visual coding layers. The predetermined number can be 1, 2, 3, or any other suitable number. Assume that cross-modal coding portions are deployed at predetermined intervals w starting from the s1 network layer. Based on the unified scheme deployment, the following network layers in feature extraction model 120 can be selected as cross-modal coding layers:

[0082] Where s i =s i-1 +w (3) where 1≤s1≤s M ≤N, w is an integer and 1≤w≤N-1.

[0083] According to equation (3) above, the nth network layer in feature extraction model 120 The cross-modal coding layers are deployed as cross-modal coding components (assuming each cross-modal coding component includes a single cross-modal coding layer), and there are a total of M cross-modal coding layers in feature extraction model 120. i-1 There are always w visual coding layers between each network layer and the s-th network layer.

[0084] like Figure 7As shown, based on the random scheme 702, the network layers in the feature extraction model 120 are alternately deployed as cross-modal coding layers 712 and visual coding layers 722. For example, the 1st, 3rd, 5th, 7th and 9th network layers are deployed as cross-modal coding layers 712, the 2nd, 4th, 6th, 8th and 10th network layers are deployed as visual coding layers 722, and so on.

[0085] exist Figure 7 In the example, w is configured to 1, and s1 = 1, meaning the first network layer is a cross-modal coding layer 712. If the parameter w is set to a larger value, the cost of cross-modal feature encoding is lower. If w is too large, i.e., the cross-modal feature encoding is too sparse, it may also affect the generalization ability of the extracted features. In practical applications, a trade-off can be struck between computational cost and model representation.

[0086] In some embodiments, the cross-modal coding portion can also be deployed based on a progressive scheme. For example, instead of a uniform scheme, the cross-modal extraction portion can be set up with progressive intervals, starting from a beginning position. Based on a progressive scheme, the following network layers in the feature extraction model 120 can be selected as cross-modal coding layers:

[0087] Where s i =s i-1 +(wi*k) (4) where 1≤s1≤s M ≤N, and (wi*k) is the asymptotic interval. w and k are both integers, and 1≤w≤N-1, M*k≤w. According to the above equation (4), the nth network layer in the feature extraction model 120 The cross-modal coding layers are deployed as cross-modal coding parts (assuming each cross-modal coding part includes a single cross-modal coding layer), and there are a total of M cross-modal coding layers in the feature extraction model 120.

[0088] In equation (4) above, if k > 0, the interval between adjacent cross-modal coding parts (i.e., the number of visual coding layers deployed therein) gradually increases, exhibiting a cross-modal interaction from dense to sparse. Figure 7 As shown, based on the progressive scheme a 703, network layers 1, 3, 6, and 10 are deployed as cross-modal coding layers 712. For network layers 1 and 3, these two adjacent cross-modal coding parts are separated by a visual coding layer 722; for network layers 3 and 6, these two adjacent cross-modal coding parts are separated by two visual coding layers 722; for network layers 6 and 10, these two adjacent cross-modal coding parts are separated by four visual coding layers 722, and so on. From the lower to the higher layers of the overall model, the cross-modal feature coding exhibits a change from dense to sparse.

[0089] In equation (4) above, if k < 0, the interval between adjacent cross-modal coding parts (i.e., the number of visual coding layers deployed therein) gradually decreases, exhibiting a cross-modal interaction from sparse to dense. Figure 7 As shown, based on the progressive scheme b 704, network layers 1, 5, 8, and 10 are deployed as cross-modal coding layers 712. For network layers 1 and 5, these two adjacent cross-modal coding parts are separated by four visual coding layers 722; for network layers 5 and 8, these two adjacent cross-modal coding parts are separated by two visual coding layers 722; for network layers 8 and 10, these two adjacent cross-modal coding parts are separated by one visual coding layer 722, and so on. From the lower to the higher layers of the overall model, the cross-modal feature encoding exhibits a change from sparse to dense.

[0090] Despite Figure 7 The schemes shown assume that each cross-modal coding part includes a single cross-modal coding layer, but the cross-modal coding part can also be configured to include multiple cross-modal coding layers.

[0091] The sparse-to-dense cross-modal interaction strategy is essentially equivalent to reducing cross-modal interactions at low-level features while maintaining dense interactions at high-level features. The dense-to-sparse cross-modal interaction strategy presents the opposite approach. In practical applications, the choice can be made based on the specific needs.

[0092] It should be understood that Figure 7 The feature extraction model 120 shown in the various schemes is only an example. The specific number and arrangement of the cross-modal coding layer and the visual coding layer can be selected according to the actual application.

[0093] As analyzed above, images and text have different semantic densities. Frequent alignment and interaction between highly semantic text and highly redundant image data is not only unnecessary but also limits the learning of feature representations of visual information. Therefore, the feature extraction model architecture proposed in this disclosure cuts off a large amount of unnecessary dense cross-modal interaction and text modeling, while still retaining slightly denser modeling of image data. This significantly improves the representational power of the extracted features.

[0094] In addition to causing asynchronous semantic alignment between text and image data in feature extraction, the inventors have also found that the large number of visual blocks in image data is redundant for cross-modal alignment. For example, multiple regions in a single image may represent the same meaning, and multiple consecutive video frames may contain even more redundant regions. Considering such redundancy, in order to save training overhead during the training phase, in some embodiments, the image data 102 is sparsely sampled and masked before being input into the feature extraction model 120 for processing.

[0095] In some embodiments, sparse block sampling and masking can be applied to the training data of the feature extraction model 120, particularly the image data in the training data during the pre-training phase. Thus, during training, the target visual features and target text features extracted from the image data 102 and text data 104 are used to perform parameter updates to the feature extraction model 120. The amount of training data in the pre-training phase is often very large. Masking the image data can significantly reduce the amount of data processing without affecting the model's learning efficiency.

[0096] Specifically, suppose the image data 102 and text data 104 to be input to the feature extraction model 102 are sample image data and sample text data used for model training. Suppose the image data 102 includes multiple video frames from a video segment, for example, T video frames, where T is greater than or equal to 1. At least one visual block of at least one of the T video frames can be masked in both the temporal and spatial domains to obtain T masked video frames. In some embodiments, it is assumed that each video frame is divided into multiple visual blocks. Mask images are applied to each of the T video frames. in This represents the mask map applied to the t-th video frame, indicating whether each visual block in the video frame should be masked.

[0097] For a masked visual patch, its corresponding features can be masked (e.g., ...). Figure 5 (The masked portion of the initial visual features 502 shown). Accordingly, during model extraction, the corresponding processing units in the cross-modal coding and visual coding parts can be cut off, eliminating the need to process this part of the data.

[0098] Various methods can be used to select visual blocks to be masked from video frames. In some embodiments, at least one visual block can be randomly masked from the various video frames included in image data 102 based on a random masking scheme. In some embodiments, the visual blocks to be masked in each video frame can be selected at a given ratio. Such random masking schemes are time-independent. Such masking schemes focus only on removing potentially redundant content in each frame, without regard to temporal correlations between frames.

[0099] Figure 8 A schematic diagram illustrating an example masking method for sample image data according to some embodiments of the present disclosure is shown. Figure 8 In the example, assume T = 3, meaning the image data includes video frames 810, 820, and 830. According to the random masking scheme 801, for each video frame, visual blocks to be masked are randomly selected at a given ratio (e.g., 1 / 3). The resulting masked video frames are provided to feature extraction model 120 for training.

[0100] In some embodiments, a fixed masking scheme can be used, utilizing a predetermined mask image m * From each video frame in the masked image data 102, multiple masked video frames are obtained. A predetermined mask image m is then generated. * At least one visual block at a predetermined position in a video frame is to be masked. In some embodiments, m * It can be randomly generated, and can specify which visual blocks should be masked at a given ratio. Such a fixed masking scheme can eliminate potentially redundant content through coarse temporal and spatial consistency. Figure 8 As shown, according to the fixed masking scheme 802, visual blocks at fixed positions in video frames 810, 820, and 830 are masked. In some embodiments, different predetermined mask maps may be applied for different sample image data (e.g., different video segments).

[0101] In some embodiments, at least one visual block can be selected from each of a plurality of video frames for masking based on a temporally complementary masking scheme, and the positions of the masked visual blocks are different from each other in the plurality of video frames. In some embodiments, the visual blocks selected from the plurality of video frames for masking are complementary to each other to form a complete "video frame". Figure 8 As shown, according to the temporal complementary masking scheme 803, in each of video frames 810, 820, and 830, one-third of the nine visual blocks (i.e., three visual blocks) are selected for masking. The positions of the masked visual blocks in the three video frames are different from each other. The nine visual blocks selected from all three video frames for masking will form the complete video frame. Compared with the previous two schemes, the temporal complementary masking scheme captures more redundant but more diverse visual information from multiple video frames.

[0102] It should be understood that Figure 8 The number of visual block divisions shown, and the visual blocks selected for masking in each video frame under each scheme, are examples. Other masking results may occur in different scenarios.

[0103] By using sparse visual patch sampling and masking, computational overhead, especially during training, can be reduced while maintaining model performance.

[0104] As mentioned earlier, the feature extraction model 120 can be pre-trained to learn better representations of image and text modal data from a large amount of data during pre-training. During the pre-training phase, a pre-training task can be constructed to achieve the pre-training objective. In some embodiments, the pre-training task may include an image and text matching task. Under this task, for a given pair of matching sample image data and sample text data, the sample image data can be randomly replaced with other sample image data with a certain probability (e.g., a probability of 0.5). Then, the feature extraction model 120 extracts the target visual features and target text features for each pair of sample image data and sample text data. The extracted target visual features and target text features are input into the output layer 130 (see reference 120) for pre-training purposes. Figure 5 The output layer 130 determines whether the input sample image data and sample text data match based on the target visual features and target text features. The goal of the pre-training task is to continuously update the parameter values ​​of the feature extraction model 120 so that the target visual features and target text features output by the feature extraction model 120 can be used to accurately determine whether the input sample image data and sample text data match.

[0105] In some embodiments, the pre-training task may further include a masked language modeling (MLM) task. For this pre-training task, a portion of the text in the sample text data is masked, and the masked sample text data and the corresponding sample image data are input into the feature extraction model 120. Then, based on the target visual features and target text features output by the feature extraction model 120, an attempt is made to predict the masked portion of the text in the sample text data. The goal of this pre-training task is to correctly predict the masked portion of the text. To achieve this task, the target visual features and target text features output by the feature extraction model 120 can be provided to an output layer 130, which is used to predict the masked portion of the text in the sample text data.

[0106] The above discussion covers some implementations of the pre-training phase for the feature extraction model 120. During the fine-tuning and model application phases, the feature extraction model 120 can be combined with downstream task-specific output layers according to the actual task requirements, which will not be elaborated upon here.

[0107] Figure 9 A flowchart of a process 900 for multimodal data processing according to some implementations of this disclosure is shown. Process 900 can be performed in... Figure 1 The cross-modal data processing system 110 may include... Figure 2 The model pre-training system 210, the model fine-tuning system 220, and / or the model application system 230. For ease of discussion, reference will be made to... Figure 1The environment 100 is used to describe the process 900.

[0108] In box 910, the cross-modal data processing system 110 acquires image data and text data.

[0109] In box 920, the cross-modal data processing system 110 uses a feature extraction model to extract target visual features from image data and target text features from text data. The feature extraction model includes alternately deployed cross-modal coding and visual coding parts. The extraction in box 920 includes: in box 922, performing cross-modal feature coding on a first intermediate visual feature of image data and a first intermediate text feature of text data using the first cross-modal coding part of the feature extraction model to obtain a second intermediate visual feature and a second intermediate text feature; in box 924, performing visual modal feature coding on the second intermediate visual feature using the first visual coding part of the feature extraction model to obtain a third intermediate visual feature; in box 926, performing cross-modal feature coding on the third intermediate visual feature and the second intermediate text feature using the second cross-modal coding part of the feature extraction model to obtain a fourth intermediate visual feature and a third intermediate text feature; and in box 928, determining target visual features and target text features based on the fourth intermediate visual feature and the third intermediate text feature.

[0110] In some embodiments, process 900 further includes: determining the degree of matching between image data and text data based on target visual features and target text features.

[0111] In some embodiments, determining the target visual features and target text features based on the fourth intermediate visual features and the third intermediate text features includes: performing visual modal feature encoding on the fourth intermediate visual features using the second visual encoding part of the feature extraction model to obtain the fifth intermediate visual features; and determining the target visual features and target text features based on the fifth intermediate visual features and the third intermediate text features.

[0112] In some embodiments, the feature extraction model includes multiple pairs of alternately deployed cross-modal coding portions and visual coding portions, wherein the visual coding portion deployed between two adjacent cross-modal coding portions includes a predetermined number of visual coding layers.

[0113] In some embodiments, the feature extraction model includes multiple pairs of alternating cross-modal coding portions and visual coding portions, wherein the visual coding portion deployed between a first pair of adjacent cross-modal coding portions includes a first number of visual coding layers, and the visual coding portion deployed between a second pair of adjacent cross-modal coding portions includes a second number of visual coding layers, the first number being different from the second number.

[0114] In some embodiments, image data and text data are included in the training data of the feature extraction model, wherein the image data includes multiple video frames in a video clip. In some embodiments, extracting target visual features and text features includes: generating multiple masked video frames by masking at least one visual block of at least one video frame in the multiple video frames; and using the feature extraction model to extract target visual features from the multiple masked video frames and target text features from the text data.

[0115] In some embodiments, process 900 further includes: performing parameter updates on the feature extraction model based on target visual features and target text features.

[0116] In some embodiments, generating multiple masked video frames includes: randomly masking at least one visual block from each of the multiple video frames to obtain multiple masked video frames.

[0117] In some embodiments, generating a plurality of masked video frames includes: masking each video frame in a plurality of video frames using a predetermined mask image to obtain a plurality of masked video frames, wherein the predetermined mask image indicates that at least one visual block at a predetermined position in the video frame is to be masked.

[0118] In some embodiments, generating multiple masked video frames includes: selecting at least one visual block from each of the multiple video frames for masking, thereby obtaining multiple masked video frames, wherein the positions of the masked visual blocks in the multiple video frames are different from each other.

[0119] Figure 10 A block diagram of an apparatus 1000 for training a contrastive learning model according to some implementations of this disclosure is shown. The apparatus 1000 may, for example, be implemented in or included in... Figure 1 The cross-modal data processing system 110 may include... Figure 2 The device 1000 includes a model pre-training system 210, a model fine-tuning system 220, and / or a model application system 230. Each module / component in the device 1000 can be implemented using hardware, software, firmware, or any combination thereof.

[0120] As shown in the figure, the device 1000 includes an acquisition module 1010 configured to acquire image data and text data. The device 1000 also includes an extraction module 1020 configured to extract target visual features from the image data and target text features from the text data using a feature extraction model. The feature extraction model includes alternating deployments of cross-modal coding and visual coding components. The extraction module 1020 includes: a first cross-modal coding module 1022, configured to perform cross-modal feature coding on the first intermediate visual features of image data and the first intermediate text features of text data using the first cross-modal coding part of the feature extraction model to obtain second intermediate visual features and second intermediate text features; a first visual coding module 1024, configured to perform visual modal feature coding on the second intermediate visual features using the first visual coding part of the feature extraction model to obtain third intermediate visual features; a second cross-modal coding module 1026, configured to perform cross-modal feature coding on the third intermediate visual features and the second intermediate text features using the second cross-modal coding part of the feature extraction model to obtain fourth intermediate visual features and third intermediate text features; and a target feature determination module 1028, configured to determine target visual features and target text features based on the fourth intermediate visual features and the third intermediate text features.

[0121] In some embodiments, the apparatus 1000 further includes a matching degree determination module, configured to determine the matching degree between image data and text data based on target visual features and target text features.

[0122] In some embodiments, the target feature determination module 1028 includes: a second visual encoding module configured to perform visual modal feature encoding on a fourth intermediate visual feature using the second visual encoding part of a feature extraction model to obtain a fifth intermediate visual feature; and a further target feature determination module configured to determine target visual features and target text features based on the fifth intermediate visual feature and the third intermediate text feature.

[0123] In some embodiments, the feature extraction model includes multiple pairs of alternately deployed cross-modal coding portions and visual coding portions, wherein the visual coding portion deployed between two adjacent cross-modal coding portions includes a predetermined number of visual coding layers.

[0124] In some embodiments, the feature extraction model includes multiple pairs of alternating cross-modal coding portions and visual coding portions, wherein the visual coding portion deployed between a first pair of adjacent cross-modal coding portions includes a first number of visual coding layers, and the visual coding portion deployed between a second pair of adjacent cross-modal coding portions includes a second number of visual coding layers, the first number being different from the second number.

[0125] In some embodiments, image data and text data are included in the training data of the feature extraction model, and the image data includes multiple video frames in a video segment. The extraction module 1020 includes: a masking module configured to generate multiple masked video frames by masking at least one visual block of at least one of the multiple video frames; and a mask-based extraction module configured to extract target visual features from the multiple masked video frames and target text features from the text data using the feature extraction model.

[0126] In some embodiments, the apparatus 1000 further includes a parameter update module configured to perform parameter updates on the feature extraction model based on target visual features and target text features.

[0127] In some embodiments, the masking module includes a random masking module configured to randomly mask at least one visual block from each of a plurality of video frames to obtain a plurality of masked video frames.

[0128] In some embodiments, the masking module includes: a fixed masking module configured to mask each of a plurality of video frames using a predetermined mask image to obtain a plurality of masked video frames, wherein the predetermined mask image indicates that at least one visual block at a predetermined position in the video frame is to be masked.

[0129] In some embodiments, the masking module includes a complementary masking module configured to select at least one visual block from each of a plurality of video frames for masking, thereby obtaining a plurality of masked video frames, wherein the positions of the masked visual blocks in the plurality of video frames are different from each other.

[0130] Figure 11 A block diagram of an electronic device 1100 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 11 The electronic device 1100 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 1100 can, for example, be used to implement... Figure 1 The cross-modal data processing system 110 may include... Figure 2 The model pre-training system 210, the model fine-tuning system 220, and / or the model application system 230. Electronic device 1100 can also be used to implement... Figure 10 Device 1000.

[0131] like Figure 11As shown, electronic device 1100 is in the form of a general-purpose computing device. Components of electronic device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 1100.

[0132] Electronic device 1100 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1130 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 1100.

[0133] Electronic device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 11 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1120 may include computer program product 1125 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0134] Communication unit 1140 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 1100 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 1100 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0135] Input device 1150 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1160 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1100 can also communicate with one or more external devices (not shown) via communication unit 1140 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1100, or with any device that enables electronic device 1100 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0136] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0137] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0138] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0139] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0141] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for multimodal data processing, comprising: Acquire image and text data; as well as The target visual features of the image data and the target text features of the text data are extracted using a feature extraction model. The feature extraction model includes alternating cross-modal coding and visual coding components. The extraction includes: The first cross-modal coding part of the feature extraction model is used to perform cross-modal feature coding on the first intermediate visual feature of the image data and the first intermediate text feature of the text data to obtain the second intermediate visual feature and the second intermediate text feature; The first visual encoding part of the feature extraction model is used to perform visual modal feature encoding on the second intermediate visual feature to obtain the third intermediate visual feature; The second cross-modal encoding part of the feature extraction model is used to perform cross-modal feature encoding on the third intermediate visual feature and the second intermediate text feature to obtain the fourth intermediate visual feature and the third intermediate text feature; and The target visual feature and the target text feature are determined based on the fourth intermediate visual feature and the third intermediate text feature.

2. The method according to claim 1, further comprising: The matching degree between the image data and the text data is determined based on the target visual features and the target text features.

3. The method according to claim 1, wherein determining the target visual feature and the target text feature based on the fourth intermediate visual feature and the third intermediate text feature comprises: The fourth intermediate visual feature is encoded using the second visual encoding part of the feature extraction model to obtain the fifth intermediate visual feature; as well as The target visual feature and the target text feature are determined based on the fifth intermediate visual feature and the third intermediate text feature.

4. The method of claim 1, wherein the feature extraction model comprises multiple pairs of alternately deployed cross-modal coding parts and visual coding parts, and wherein the visual coding part deployed between two adjacent cross-modal coding parts comprises a predetermined number of visual coding layers.

5. The method of claim 1, wherein the feature extraction model comprises multiple pairs of alternating cross-modal coding portions and visual coding portions, and wherein the visual coding portions deployed between the first pair of adjacent cross-modal coding portions comprise a first number of visual coding layers, and the visual coding portions deployed between the second pair of adjacent cross-modal coding portions comprise a second number of visual coding layers, the first number being different from the second number.

6. The method of claim 1, wherein the image data and the text data are included in the training data of the feature extraction model, and wherein the image data includes multiple video frames from a video segment, wherein extracting the target visual features and the text features comprises: Multiple masked video frames are generated by masking at least one visual block of at least one of the multiple video frames. as well as The feature extraction model is used to extract the target visual features of the multiple masked video frames and the target text features of the text data.

7. The method according to claim 6, further comprising: The parameters of the feature extraction model are updated based on the target visual features and the target text features.

8. The method of claim 6, wherein generating the plurality of masked video frames comprises: At least one visual block is randomly masked from each of the plurality of video frames to obtain the plurality of masked video frames.

9. The method of claim 6, wherein generating the plurality of masked video frames comprises: Each video frame in the plurality of video frames is masked using a predetermined mask image to obtain the plurality of masked video frames, wherein the predetermined mask image indicates that at least one visual block at a predetermined position in the video frame is to be masked.

10. The method of claim 6, wherein generating the plurality of masked video frames comprises: At least one visual block is selected from each of the plurality of video frames for masking to obtain the plurality of masked video frames, wherein the positions of the masked visual blocks in the plurality of video frames are different from each other.

11. An apparatus for multimodal data processing, comprising: The acquisition module is configured to acquire image data and text data; as well as An extraction module is configured to extract target visual features from the image data and target text features from the text data using a feature extraction model, wherein the feature extraction model includes alternately deployed cross-modal coding and visual coding parts, and the extraction module includes: The first cross-modal coding module is configured to perform cross-modal feature coding on the first intermediate visual feature of the image data and the first intermediate text feature of the text data using the first cross-modal coding part of the feature extraction model to obtain the second intermediate visual feature and the second intermediate text feature. The first visual encoding module is configured to perform visual modal feature encoding on the second intermediate visual feature using the first visual encoding part of the feature extraction model to obtain the third intermediate visual feature; The second cross-modal coding module is configured to perform cross-modal feature coding on the third intermediate visual feature and the second intermediate text feature using the second cross-modal coding part of the feature extraction model to obtain the fourth intermediate visual feature and the third intermediate text feature; and The target feature determination module is configured to determine the target visual features and the target text features based on the fourth intermediate visual features and the third intermediate text features.

12. The apparatus of claim 11, wherein the target feature determination module comprises: The second visual encoding module is configured to perform visual modal feature encoding on the fourth intermediate visual feature using the second visual encoding part of the feature extraction model to obtain the fifth intermediate visual feature; The further target feature determination module is configured to determine the target visual feature and the target text feature based on the fifth intermediate visual feature and the third intermediate text feature.

13. The apparatus of claim 11, wherein the feature extraction model comprises multiple pairs of alternately deployed cross-modal coding portions and visual coding portions, and wherein the visual coding portion deployed between two adjacent cross-modal coding portions comprises a predetermined number of visual coding layers.

14. The apparatus of claim 11, wherein the feature extraction model comprises multiple pairs of alternating cross-modal coding portions and visual coding portions, and wherein the visual coding portions deployed between the first pair of adjacent cross-modal coding portions comprise a first number of visual coding layers, and the visual coding portions deployed between the second pair of adjacent cross-modal coding portions comprise a second number of visual coding layers, the first number being different from the second number.

15. The apparatus of claim 11, wherein the image data and the text data are included in the training data of the feature extraction model, and wherein the image data includes multiple video frames from a video segment, wherein the extraction module comprises: The masking module is configured to generate a plurality of masked video frames by masking at least one visual block of at least one of the plurality of video frames; as well as The mask-based extraction module is configured to use the feature extraction model to extract the target visual features of the plurality of masked video frames and the target text features of the text data.

16. The apparatus of claim 15, wherein the masking module comprises: The random masking module is configured to randomly mask at least one visual block from each of the plurality of video frames to obtain the plurality of masked video frames.

17. The apparatus of claim 15, wherein the masking module comprises: A fixed masking module is configured to mask each of the plurality of video frames using a predetermined mask image to obtain the plurality of masked video frames, wherein the predetermined mask image indicates that at least one visual block at a predetermined position in the video frame is to be masked.

18. The apparatus of claim 15, wherein the mask module comprises: A complementary masking module is configured to select at least one visual block from each of the plurality of video frames for masking, thereby obtaining the plurality of masked video frames, wherein the positions of the masked visual blocks in the plurality of video frames are different from each other.

19. An electronic device comprising: At least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

20. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 10.