Multi-modal recognition model training method, multi-modal data recognition method and related devices

Through the dual-channel feature extraction with different particle sizes and the dynamic gating mechanism of feature-level hybrid models, combined with the combined loss optimization of multi-particle size mixed output and feature-level hybrid output, the problem that the existing multi-modal pre-training method cannot fully utilize the multi-particle size alignment characteristics is solved, and the identification accuracy of multi-modal data and the modeling ability of the model to model complex modal relationships is improved.

CN120068973BActive Publication Date: 2025-07-01PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510535064.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-01
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing multimodal pre-training methods cannot fully utilize the multi-grained alignment characteristics of visual language, resulting in low recognition accuracy of multimodal recognition models in recognition tasks.

Method used

Through dual-channel feature extraction with different particle sizes, multi-level correlation across multimodal data is captured, and key features are screened through dynamic gating mechanisms using feature-level hybrid models, combining the combined loss optimization of multi-grained hybrid output and feature-level hybrid output to independently balance the modal alignment targets of different particle sizes.

Benefits of technology

It improves the identification accuracy of multimodal data, enhances the model's ability to model complex modal relationships, and achieves a more comprehensive understanding of the implicit relationship between images and text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068973B_ABST
    Figure CN120068973B_ABST
Patent Text Reader

Abstract

The multi-modal recognition model training method, multi-modal data recognition method and related devices proposed in the embodiments of the present application. The method includes: obtaining multiple multi-modal training sample pairs; based on the multi-granularity output model in the multi-modal recognition model, performing first feature extraction and second feature extraction on each multi-modal training sample pair to obtain first feature data and second feature data with different processing granularities, and mixing each first feature data and the corresponding second feature data to obtain a multi-granularity mixed output; based on the feature-level mixing model in the multi-modal recognition model, performing feature-level mixing processing on each multi-modal training sample pair to obtain a feature-level mixed output corresponding to each multi-modal training sample; calculating a multi-modal loss based on the multi-granularity mixed output and the feature-level mixed output, and updating the multi-modal recognition model based on the multi-modal loss, effectively improving the recognition accuracy of multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a multi-modal recognition model training method, a multi-modal data recognition method, and related devices. Background Art

[0002] In recent years, with the continuous development of deep learning technology, significant progress has been made in fields such as computer vision and natural language processing. In many tasks, the fusion of multi-modal information (such as images and text) has become an important research direction. Large-scale pre-trained models (such as BERT, GPT, etc.) have significantly improved the performance of natural language processing tasks through unsupervised pre-training on a large amount of text data. Similarly, large-scale multi-modal pre-trained models (such as ViLBERT, etc.) have achieved significant improvements in cross-modal task performance through joint pre-training on image and text data. However, existing multi-modal pre-training methods usually only target the visual-language correlation at a single granularity, making it impossible to fully utilize the multi-granularity alignment characteristics of visual language for training. As a result, when the trained multi-modal recognition model performs multi-modal recognition tasks, its recognition accuracy is relatively low. Summary of the Invention

[0003] Embodiments of this application provide a multi-modal recognition model training method, a multi-modal data recognition method, and related devices, which can improve the recognition accuracy of multi-modal data.

[0004] To achieve the above object, a first aspect of the embodiments of this application proposes a multi-modal recognition model training method, the method including:

[0005] Obtain multiple multi-modal training sample pairs, where each multi-modal training sample pair includes an image training sample and a corresponding text training sample;

[0006] Based on the multi-granularity output model in the multi-modal recognition model, perform first feature extraction on each multi-modal training sample pair to obtain first feature data, perform second feature extraction on each multi-modal training sample pair to obtain second feature data, and perform hybrid processing on each first feature data and the corresponding second feature data to obtain a multi-granularity hybrid output, where the processing granularities of the first feature extraction and the second feature extraction are different;

[0007] Based on the feature-level hybrid model in the multi-modal recognition model, perform feature-level hybrid processing on each multi-modal training sample pair to obtain a feature-level hybrid output corresponding to each multi-modal training sample;

[0008] Calculate a multi-modal loss based on the multi-granularity hybrid output and the feature-level hybrid output, and update the multi-modal recognition model based on the multi-modal loss.

[0009] In some embodiments, the process of mixing each of the first feature data and the corresponding second feature data to obtain a multi-granularity mixed output includes:

[0010] Obtain an expert selection model, input the multi-modal training sample pairs into the expert selection model for image-text association processing, and obtain a first feature weight corresponding to the first feature extraction and a second feature weight corresponding to the second feature extraction;

[0011] Based on the product of the first feature weight and the output of the first feature extraction, plus the product of the second feature weight and the output of the second feature extraction, obtain the multi-granularity mixed output.

[0012] In some embodiments, the process of performing first feature extraction on each multi-modal training sample pair to obtain first feature data includes:

[0013] Perform position encoding processing on the image training sample pair to obtain a position-encoded image;

[0014] Based on a first processing granularity, divide the position-encoded image into multiple granularity-divided images;

[0015] Perform feature extraction on each of the granularity-divided images one by one to obtain granularity-divided image features;

[0016] Based on all the granularity-divided image features and the text training sample, obtain the granularity-divided image features, and based on all the granularity-divided image features, obtain the first feature data.

[0017] In some embodiments, the multi-layer perceptron of the multi-modal recognition model is the multi-layer perceptron of the mixture-of-experts module. The process of performing feature-level mixing processing on each multi-modal training sample pair to obtain a feature-level mixed output corresponding to each multi-modal training sample includes:

[0018] Input the multi-modal training sample pair into the feature-level mixing model for recognition processing to obtain multiple feature expert model outputs, and at least two of the feature expert model outputs are different;

[0019] Obtain a gating weight matrix, and based on the product of the gating weight matrix and the multiple feature expert model outputs, perform normalization processing to obtain gating parameters;

[0020] Accumulate the products of all the feature expert model outputs and the corresponding gating parameters to obtain the feature-level mixed output.

[0021] In some embodiments, calculating the multimodal loss based on the multi-granularity hybrid output and the feature-level hybrid output includes:

[0022] Calculating the image-text matching loss value of the multi-granularity hybrid output and the feature-level hybrid output based on the image-text matching training task;

[0023] Calculating the text prediction loss value of the multi-granularity hybrid output and the feature-level hybrid output based on the text prediction training task;

[0024] Calculating the multimodal contrast loss value of the multi-granularity hybrid output and the feature-level hybrid output based on the multimodal contrast training task;

[0025] Obtaining the multimodal loss based on the image-text matching loss value, the text prediction loss value, and the multimodal contrast loss value.

[0026] In some embodiments, calculating the image-text matching loss value of the multi-granularity hybrid output and the feature-level hybrid output based on the image-text matching training task includes:

[0027] Calculating the image-text matching score of the multi-granularity hybrid output and the feature-level hybrid output based on the image-text matching training task by using a matching score function;

[0028] Calculating the image-text matching loss value based on the image-text matching score.

[0029] In some embodiments, calculating the text prediction loss value of the multi-granularity hybrid output and the feature-level hybrid output based on the text prediction training task includes:

[0030] Obtaining a partial text occlusion sample corresponding to the text training sample in the multimodal training sample pair;

[0031] Based on the text prediction training task, using the partial text occlusion sample and the image training sample in the multimodal training sample as a text prediction training sample pair, and calculating the multi-granularity hybrid output and the feature-level hybrid output of the text prediction training sample pair;

[0032] Using a text prediction function to calculate an occluded predicted text based on the multi-granularity hybrid output, the feature-level hybrid output, and the partial text occlusion sample of the text prediction training sample pair;

[0033] Calculating the text prediction loss value based on the occluded predicted text and the corresponding text training sample.

[0034] In some embodiments, calculating the multi-modal contrast loss value of the multi-granularity mixed output and the feature-level mixed output based on the multi-modal contrast training task includes:

[0035] Obtain the incorrect training text corresponding to the text training sample in the multi-modal training sample pair;

[0036] Based on the multi-modal contrast training task, use the incorrect training text and the image training sample in the multi-modal training sample as a text contrast sample pair, and calculate the contrast multi-granularity mixed output and the contrast feature-level mixed output of the text contrast sample pair;

[0037] Use a similarity function to calculate the positive multi-modal contrast similarity of the multi-granularity mixed output and the feature-level mixed output, and calculate the negative multi-modal contrast similarity of the contrast multi-granularity mixed output and the contrast feature-level mixed output;

[0038] Based on the positive multi-modal contrast similarity and the negative multi-modal contrast similarity, obtain the multi-modal contrast loss value.

[0039] In some embodiments, obtaining the multi-modal contrast loss value based on the positive multi-modal contrast similarity and the negative multi-modal contrast similarity includes:

[0040] Use a probability distribution function to obtain the positive probability corresponding to the positive multi-modal contrast similarity and the negative probability corresponding to the negative multi-modal contrast similarity;

[0041] Accumulate the positive probability and the negative probability to obtain a multi-modal probability value;

[0042] Based on the ratio of the positive probability and the multi-modal probability value, and then perform logarithmic processing to obtain the multi-modal contrast loss value.

[0043] In some embodiments, obtaining the multi-modal loss based on the text-image matching loss value, the text prediction loss value, and the multi-modal contrast loss value includes:

[0044] Obtain a text prediction weight and a multi-modal contrast weight;

[0045] Based on the product of the text prediction weight and the text prediction loss value, obtain a text prediction weighted loss value;

[0046] Based on the product of the multi-modal contrast weight and the multi-modal contrast loss value, obtain a multi-modal contrast weighted loss value;

[0047] Accumulate the graphic-text matching loss value, the weighted text prediction loss value, and the weighted multi-modal contrast loss value to obtain the multi-modal loss.

[0048] To achieve the above object, a second aspect of the embodiments of the present application provides a multi-modal data recognition method, the method comprising:

[0049] In response to a multi-modal recognition processing task, obtain multi-modal processing data;

[0050] Based on the multi-modal recognition processing task, input the multi-modal processing data into a multi-modal recognition model trained by the multi-modal recognition model training method described in the first aspect for data recognition processing to obtain target recognition data.

[0051] To achieve the above object, a third aspect of the embodiments of the present application provides a multi-modal recognition model training device, the device comprising:

[0052] A sample pair acquisition module, configured to acquire multiple multi-modal training sample pairs, the multi-modal training sample pairs including image training samples and corresponding text training samples;

[0053] A multi-granularity processing module, configured to perform first feature extraction on each multi-modal training sample pair based on a multi-granularity output model in the multi-modal recognition model to obtain first feature data, perform second feature extraction on each multi-modal training sample pair to obtain second feature data, and perform hybrid processing on each first feature data and the corresponding second feature data to obtain a multi-granularity hybrid output, where the processing granularities of the first feature extraction and the second feature extraction are different;

[0054] A feature hybrid processing module, configured to perform feature-level hybrid processing on each multi-modal training sample pair based on a feature-level hybrid model in the multi-modal recognition model to obtain a feature-level hybrid output corresponding to each multi-modal training sample;

[0055] A model update module, configured to calculate a multi-modal loss based on the multi-granularity hybrid output and the feature-level hybrid output, and update the multi-modal recognition model based on the multi-modal loss.

[0056] To achieve the above object, a fourth aspect of the embodiments of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the multi-modal recognition model training method described in the first aspect or the multi-modal data recognition method described in the second aspect when executing the computer program.

[0057] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-modal recognition model training method described in the first aspect above or the multi-modal data recognition method described in the second aspect.

[0058] The multi-modal recognition model training method, multi-modal data recognition method and related devices provided by the embodiments of the present application include: First, obtain multiple multi-modal training sample pairs, where the multi-modal training sample pairs include image training samples and corresponding text training samples; Second, based on the multi-granularity output model in the multi-modal recognition model, perform first feature extraction on each multi-modal training sample pair to obtain first feature data, perform second feature extraction on each multi-modal training sample pair to obtain second feature data, and perform a mixing process on each first feature data and the corresponding second feature data to obtain a multi-granularity mixed output, where the processing granularities of the first feature extraction and the second feature extraction are different; Next, based on the feature-level mixing model in the multi-modal recognition model, perform feature-level mixing processing on each multi-modal training sample pair to obtain a feature-level mixed output corresponding to each multi-modal training sample; Finally, calculate a multi-modal loss based on the multi-granularity mixed output and the feature-level mixed output, and update the multi-modal recognition model based on the multi-modal loss. The embodiments of the present application capture the multi-level correlation across multi-modal data from the perspectives of fine-grained processing and coarse-grained processing through dual-channel feature extraction with different granularities, enabling the model to adapt to both coarse-grained retrieval and fine-grained reasoning tasks simultaneously. The feature-level mixing model is used to screen key features through a dynamic gating mechanism, enhancing the model's ability to model complex modal relationships. Then, based on the joint loss optimization of the multi-granularity mixed output and the feature-level mixed output, the multi-modal recognition model autonomously balances the modal alignment objectives of different granularities during training, thereby achieving a more comprehensive understanding of the implicit relationship between images and texts. Furthermore, when using the trained multi-modal recognition model to perform the recognition of multi-modal data, the recognition accuracy of multi-modal data is effectively improved.

[0059] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the specification, claims, and drawings. Description of the Drawings

[0060] Figure 1 is a flowchart of a multi-modal recognition model training method provided by an embodiment of the present application.

[0061] Figure 2 is Figure 1 the flowchart of step 102 in

[0062] Figure 3 is Figure 1 Another flowchart of step 102 in

[0063] Figure 4 is Figure 1 The flowchart of step 103 in

[0064] Figure 5 is Figure 1 The flowchart of step 104 in

[0065] Figure 6 is Figure 5 The flowchart of step 501 in

[0066] Figure 7 is Figure 5 The flowchart of step 502 in

[0067] Figure 8 is Figure 5 The flowchart of step 503 in

[0068] Figure 9 is Figure 8 The flowchart of step 804 in

[0069] Figure 10 is Figure 5 The flowchart of step 504 in

[0070] Figure 11 It is a schematic diagram of the training process of a multi-modal recognition model provided by another embodiment of the present application.

[0071] Figure 12 It is a schematic diagram of the training process block of a multi-modal recognition model provided by another embodiment of the present application.

[0072] Figure 13 It is a flowchart of a multi-modal data recognition method provided by another embodiment of the present application.

[0073] Figure 14 It is a schematic diagram of the structure of a multi-modal recognition model training device provided by another embodiment of the present application.

[0074] Figure 15 It is a schematic diagram of the hardware structure of an electronic device provided by another embodiment of the present application. Detailed implementation manners

[0075] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0076] It should be noted that although the functional modules are divided in the schematic diagram of the device and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device or a different sequence from that in the flowchart.

[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0078] In recent years, with the continuous development of deep learning technology, significant progress has been made in fields such as computer vision and natural language processing. In many tasks, the fusion of multimodal information (such as images and text) has become an important research direction. Large-scale pre-trained models (such as BERT, GPT, etc.) have significantly improved the performance of natural language processing tasks through unsupervised pre-training on a large amount of text data. Similarly, large-scale multimodal pre-trained models (such as ViLBERT, etc.) have achieved significant improvements in cross-modal task performance through joint pre-training on image and text data. However, existing multimodal pre-training methods usually only target the visual-language correlation at a single granularity, making it impossible to fully utilize the multi-granularity alignment characteristics of visual language for training. As a result, when the trained multimodal recognition model performs multimodal recognition tasks, its recognition accuracy is relatively low.

[0079] In order to improve the recognition accuracy of multimodal data, the embodiments of this application capture the multi-level correlation across multimodal data from the perspectives of fine-grained processing and coarse-grained processing through dual-channel feature extraction at different granularities, enabling the model to adapt to both coarse-grained retrieval and fine-grained reasoning tasks simultaneously. And a feature-level hybrid model is used to screen key features through a dynamic gating mechanism to enhance the model's ability to model complex modal relationships. Then, based on the joint loss optimization of the multi-granularity hybrid output and the feature-level hybrid output, the multimodal recognition model autonomously balances the modal alignment objectives at different granularities during training, thereby achieving a more comprehensive understanding of the implicit correlation between images and text. Furthermore, when using the trained multimodal recognition model to perform the recognition of multimodal data, the recognition accuracy of multimodal data is effectively improved.

[0080] Next, the multimodal recognition model training method, multimodal data recognition method, and related devices provided by the embodiments of this application will be further described. First, the multimodal recognition model training method in the embodiments of this application will be described. Referring to Figure 1 , which is an optional flowchart of the multimodal recognition model training method provided by the embodiments of this application, Figure 1 The method inFigure 1 The order of steps 101 to 104 is not specifically limited and can be adjusted according to actual needs, or some steps can be reduced or added. The multi-modal recognition model training method provided in the embodiments of this application can be applied to any server or intelligent terminal configured with certain computing power resources, etc.

[0081] Step 101: Obtain multiple multi-modal training sample pairs.

[0082] The following provides a detailed description of step 101.

[0083] In some embodiments, in order to train a multi-modal recognition model that can accurately recognize multi-modal data in actual applications, a large number of multi-modal training sample pairs need to be obtained in advance to facilitate training and updating the multi-modal recognition model using these multi-modal training sample pairs. Among them, multi-modal training sample pairs usually include image training samples and corresponding text training samples, and even audio training texts, etc.

[0084] These multi-modal training sample pairs can be obtained by collecting common image-text data sets on the Internet, processing the data, mainly clearing invalid and dirty data, and then using them as the training set of the model. The data sets corresponding to these multi-modal training sample pairs usually contain approximately 9.6 million multi-modal data pairs in total, including MSCOCO, SBU, OPENIMAGE, GOOGLECC, FLICKER30K, and GQA.

[0085] Step 102: Based on the multi-granularity output model in the multi-modal recognition model, perform first feature extraction on each multi-modal training sample pair to obtain first feature data, perform second feature extraction on each multi-modal training sample pair to obtain second feature data, and perform hybrid processing on each first feature data and the corresponding second feature data to obtain a multi-granularity hybrid output.

[0086] The following provides a detailed description of step 102.

[0087] In some embodiments, after obtaining a large number of multi-modal training sample pairs, next, use the multi-granularity output model in the multi-modal recognition model to perform first feature extraction and second feature extraction with different granularities on each multi-modal training sample pair to obtain the corresponding first feature data and second feature data, and capture the multi-level correlation across multi-modal data from the perspectives of fine-grained processing and coarse-grained processing, so as to facilitate subsequent hybrid processing of the first feature data and the second feature data to obtain a multi-granularity hybrid output considering the characteristics of multi-granularity alignment.

[0088] In some embodiments, the processing granularity corresponding to the first feature extraction is fine-grained, and the processing granularity corresponding to the second feature extraction is coarse-grained. Herein, the fine-grained and coarse-grained refer to the size of the processing granularity when performing feature extraction for the same image training sample. For example, for an image training sample with a dimension of 30*30, the fine-grained processing is to divide the image training sample into 9 image blocks with a dimension of 10*10, and then perform separate feature extraction on each image block and then perform merging processing; while the coarse-grained processing is to directly perform feature extraction on the image training sample with a dimension of 30*30.

[0089] In some embodiments, the first feature extraction corresponding to the fine-grained processing can process the multi-modal training sample pair by inputting it into a fine-grained learning expert (FE) model, that is, FE(I, T), where I is the image feature corresponding to the image training sample in the multi-modal training sample, and T is the text feature corresponding to the text training sample in the multi-modal training sample.

[0090] Correspondingly, the second feature extraction corresponding to the coarse-grained processing can process the multi-modal training sample pair by inputting it into a coarse-grained learning expert (CE) model, that is, CE(I, T), where I is the image feature corresponding to the image training sample in the multi-modal training sample, and T is the text feature corresponding to the text training sample in the multi-modal training sample.

[0091] In some embodiments, there may not be only two-granularity processing, but also multiple-granularity processing. For example, the first feature extraction corresponding to the fine-grained processing, the second feature extraction corresponding to the medium-grained processing, and the third feature extraction corresponding to the coarse-grained processing, etc., can be custom-set according to actual needs. It can be understood that the more granularity processing is performed, the better the multi-modal recognition model can learn the multi-granularity alignment characteristics, thereby improving the accuracy of the multi-modal recognition model in recognizing multi-modal data. However, relatively speaking, the training cost also increases accordingly.

[0092] The specific steps for the first feature extraction corresponding to the fine-grained processing will be further described below.

[0093] Refer to Figure 2 , perform the first feature extraction on each multi-modal training sample pair to obtain the first feature data, including the following steps 201 to step 204.

[0094] Step 201: Perform position encoding processing on the image training sample pair to obtain a position-encoded image.

[0095] Step 202: Based on the first processing granularity, divide the position-encoded image into multiple granularity-divided images.

[0096] Step 203: Feature extraction is performed on each granularity partitioned image one by one to obtain granularity partitioned image features.

[0097] Step 204: Granularity partitioned image features are obtained based on all granularity partitioned image features and text training samples, and first feature data is obtained based on all granularity partitioned image features.

[0098] Steps 201 to 204 will be described in detail below.

[0099] In some embodiments, different from directly performing feature extraction processing on the complete (or dividing the multi-modal training samples into several larger chunks) multi-modal training samples for the second feature extraction of the coarse granularity, in the Fine-grained Expert (FE) model, for the first feature extraction corresponding to the fine granularity, it is usually necessary to perform position encoding processing on each divided block in the image training sample pair to obtain a position encoding image containing multiple position encodings. Then, according to the first processing granularity size corresponding to the fine granularity, the position encoding image is divided into multiple granularity partitioned images. Then, feature extraction is performed on each granularity partitioned image one by one to obtain granularity partitioned image features corresponding to each granularity partitioned image. Then, based on the correlation between each granularity partitioned image feature and the corresponding partial text training sample, granularity partitioned image features are obtained. Then, all the granularity partitioned image features are concatenated or merged to obtain the first feature data corresponding to the fine granularity.

[0100] Through the above steps 201 to 204, by introducing position encoding, the spatial distribution relationship of objects in the perceptual image training sample is improved, local feature confusion is avoided, and through fine-grained partitioning (such as dividing the image into local regions of different sizes), differential processing of local details is realized, enabling the model to capture key local features, and then separately fusing multi-granularity visual features and text features for each local feature, forcing the model to take into account the correlation of each divided block in the fine granularity during cross-modal alignment, and thus paying more attention to the feature information of the fine-grained local area.

[0101] In some embodiments, after performing the first feature extraction processing corresponding to the fine granularity and the second feature extraction processing corresponding to the coarse granularity on the multi-modal training sample pair respectively to obtain the first feature data and the second feature data, it is necessary to perform a mixing process on the first feature data and the second feature data, so as to perform multi-granularity alignment. How to perform the mixing process will be further described below.

[0102] Refer to Figure 3 , and perform a mixing process on each first feature data and the corresponding second feature data to obtain a multi-granularity mixed output, including the following steps 301 to 302.

[0103] Step 301: Obtain an expert selection model, input the multi-modal training sample pair into the expert selection model for image-text association processing, and obtain the first feature weight corresponding to the first feature extraction and the second feature weight corresponding to the second feature extraction.

[0104] Step 302: Based on the product of the first feature weight and the output of the first feature extraction, plus the product of the second feature weight and the output of the second feature extraction, obtain a multi-granularity hybrid output.

[0105] The following provides a detailed description of Steps 301 to 302.

[0106] In some embodiments, an expert selection model (Expert Select Model, ESM) is pre-trained to identify the correlation between image and text modalities. The expert selection model ESM can be a neural network model, such as a multi-layer perceptron (MLP). Then, input the multi-modal training sample (I, T) into the expert selection model ESM to calculate the weight wi = ESM(I, T) for each granularity processing, that is, obtain the first feature weight w1 corresponding to the first feature extraction and the second feature weight w2 corresponding to the second feature extraction.

[0107] In the mixture-of-experts technique, the expert selection model (Expert Select Model, ESM) is responsible for allocating to different expert models according to the features of the input data. In multi-modal pre-training, the expert selection model can assign weights to the visual-language correlations of different granularities according to the relevance between images and texts. In this way, the entire pre-training model can better learn cross-modal alignment relationships, thereby improving performance in downstream tasks.

[0108] Then, based on the product of the first feature weight w1 and the output of the first feature extraction FE(I, T), plus the product of the second feature weight w2 and the output of the second feature extraction CE(I, T), obtain a multi-granularity hybrid output As shown in the following formula (1).

[0109] (1)

[0110] Through the above steps 301 to 302, using the pre-trained Expert Selection Model (ESM), based on the semantic relevance of the image-text pair, the weights of fine-grained and coarse-grained feature extraction are adaptively allocated, enabling the model to flexibly adjust the focus of attention for different input data (for example, strengthening local details in complex scenarios and emphasizing global semantics in overall descriptions). Then, by weighted-fusing the features of the two granularities, the complementarity and collaboration of multi-level information are achieved. It not only preserves the context coherence of the coarse-grained features but also incorporates the precise local alignment ability of the fine-grained features. This dynamic fusion mechanism enables the model to autonomously balance the modality alignment requirements of different granularities, avoiding feature biases caused by fixed weights, thereby more comprehensively capturing the implicit associations between images and texts in cross-modal tasks and enhancing the model's generalization ability for diverse semantic relationships.

[0111] Step 103: Based on the feature-level hybrid model in the multi-modal recognition model, perform feature-level hybrid processing on each multi-modal training sample pair to obtain the feature-level hybrid output corresponding to each multi-modal training sample.

[0112] The following provides a detailed description of Step 103.

[0113] In some embodiments, after obtaining the multi-granularity hybrid output with the characteristics of multi-granularity alignment In addition, to further improve the recognition accuracy of the multi-modal recognition model, each multi-modal training sample pair will also be subjected to feature-level hybrid processing based on the feature-level hybrid model in the multi-modal recognition model to obtain the feature-level hybrid output corresponding to each multi-modal training sample for subsequent co-training and updating of the multi-modal recognition model using the multi-granularity hybrid output and the feature-level hybrid output together.

[0114] The following will further describe how to perform feature-level hybrid processing on multi-modal training samples.

[0115] Referring to Figure 4 performing feature-level hybrid processing on each multi-modal training sample pair to obtain the feature-level hybrid output corresponding to each multi-modal training sample includes the following steps 401 to 403.

[0116] Step 401: Input the multi-modal training sample pair into the feature-level hybrid model for recognition processing to obtain multiple feature expert model outputs.

[0117] Step 402: Obtain the gating weight matrix, and based on the product of the gating weight matrix and the multiple feature expert model outputs, perform normalization processing to obtain the gating parameters.

[0118] Step 403: Accumulate the products of all the outputs of the feature expert models and the corresponding gating parameters to obtain the feature-level hybrid output.

[0119] The following provides a detailed description of Steps 401 to 403.

[0120] To improve the performance and computational efficiency of the model, the embodiment of this application adopts feature-level expert mixture in the feature-level hybrid model. By replacing the multi-layer perceptron (MLP) in the original Transformer module with a mixture of experts (MoE)-based MLP and using a gating mechanism to select the expert models participating in the training.

[0121] The mixture of experts (MoE) technology is a method that uses multiple expert models to jointly learn. These expert models can be trained for different types or granularities of data to improve the generalization ability and adaptability of the model. In multi-modal pre-training, the mixture of experts technology can be used to model the visual-language correlations at different granularities, thereby improving the performance of the model.

[0122] In some embodiments, in this feature-level hybrid model, the multi-modal training sample pairs are input into the feature-level hybrid model for recognition processing by multiple expert models to obtain multiple outputs of the feature expert models .

[0123] Next, based on the product of the gating weight matrix and the multiple outputs of the feature expert models , and through normalization processing softmax(), the gating weights are obtained as shown in the following formula (2).

[0124] (2)

[0125] where the gating weights include the gating parameters corresponding to each expert model .

[0126] Next, accumulate the products of all the outputs of the feature expert models and the corresponding gating parameters to obtain the feature-level hybrid output as shown in the following formula (3).

[0127] (3)

[0128] Through the above steps 401 to 403, the input data is processed in parallel by multiple expert models in the feature-level hybrid model, enabling the model to learn separately for different modal characteristics (such as local image texture, text semantics), avoiding the feature representation bias of a single model. The activation intensity of each expert is dynamically allocated through the gating weight matrix (such as screening key experts using normalized weights), which not only reduces redundant calculations (only activating some expert models), but also ensures that the model can flexibly call the most relevant feature processing modules according to the input characteristics. Then, the weighted accumulation mechanism is further used to fuse the complementary information output by multiple experts, enhancing the synergy of cross-modal features while maintaining the model capacity. This method realizes the efficient modeling of complex multi-modal data through gating-driven expert collaboration, enabling the model to adaptively integrate feature information at different levels, thereby improving the accuracy and robustness of cross-modal alignment.

[0129] Step 104: Calculate the multi-modal loss based on the multi-granularity hybrid output and the feature-level hybrid output, and update the multi-modal recognition model based on the multi-modal loss.

[0130] The following provides a detailed description of step 104.

[0131] In some embodiments, after obtaining the multi-granularity hybrid output and the feature-level hybrid output , a hybrid calculation is performed based on the multi-granularity hybrid output and the feature-level hybrid output to obtain the multi-modal loss , so as to use the multi-modal loss to update the multi-modal recognition model, thereby effectively improving the accuracy of using the multi-modal recognition model to recognize multi-modal data.

[0132] In the pre-training stage of the multi-modal recognition model of the present invention, three pre-training tasks are adopted: ImageText Matching (ITM), Masked Text Learning (MTL), and Cross-modal Contrastive Learning (CCL). These tasks can jointly improve the generalization ability and adaptability of the multi-modal recognition model.

[0133] Next, it will be further described how to obtain the multi-modal loss based on the multi-granularity hybrid output and the feature-level hybrid output based on these three pre-training tasks.

[0134] Referring to Figure 5 , calculating the multi-modal loss based on the multi-granularity hybrid output and the feature-level hybrid output includes the following steps 501 to 504.

[0135] Step 501: Based on the image-text matching training task, calculate the image-text matching loss values for the multi-granularity hybrid output and the feature-level hybrid output.

[0136] The following provides a detailed description of Step 501.

[0137] In some embodiments, the image-text matching (ITM) task aims to train a multi-modal recognition model to correctly match images and texts. Given an image I and a text T, the model needs to predict the matching score between them. The following will further describe how to obtain the image-text matching loss value.

[0138] Referring to Figure 6 , based on the image-text matching training task, calculate the image-text matching loss values for the multi-granularity hybrid output and the feature-level hybrid output, including the following steps 601 to 602.

[0139] Step 601: Based on the image-text matching training task, use the matching score function to calculate the image-text matching scores for the multi-granularity hybrid output and the feature-level hybrid output.

[0140] Step 602: Based on the image-text matching scores, calculate the image-text matching loss values.

[0141] The following provides a detailed description of Steps 601 to 602.

[0142] In some embodiments, based on the image-text matching training task, use the matching score function , calculate the image-text matching scores for the multi-granularity hybrid output and the feature-level hybrid output as shown in the following formula (4).

[0143] (4)

[0144] where the matching score function is based on the matching score calculation method in information theory and is widely used in pattern recognition and computer vision tasks to measure the similarity between two samples.

[0145] Next, based on this image-text matching score , use a common loss function (such as cross-entropy loss, etc.) to obtain the image-text matching loss value corresponding to the image-text matching training task .

[0146] Step 502: Based on the text prediction training task, calculate the text prediction loss values for the multi-granularity hybrid output and the feature-level hybrid output.

[0147] The following provides a detailed description of Step 502.

[0148] In some embodiments, the text prediction (MTL) task aims to train a multi-modal recognition model to learn to predict occluded text, that is, given an occluded text and an image I, the multi-modal recognition model needs to predict the text of the occluded part, which is described as follows.

[0149] Referring to Figure 7 , based on the text prediction training task, the text prediction loss values of the multi-granularity mixed output and the feature-level mixed output are calculated, including the following steps 701 to step 702.

[0150] Step 701: Obtain a partial text occlusion sample corresponding to the text training sample in the multi-modal training sample pair.

[0151] Step 702: Based on the text prediction training task, take the partial text occlusion sample and the image training sample in the multi-modal training sample as a text prediction training sample pair, and calculate the multi-granularity mixed output and the feature-level mixed output of the text prediction training sample pair.

[0152] Step 703: Use the text prediction function to calculate the occluded predicted text based on the multi-granularity mixed output, the feature-level mixed output and the partial text occlusion sample of the text prediction training sample pair.

[0153] Step 704: Calculate the text prediction loss value based on the occluded predicted text and the corresponding text training sample.

[0154] The following is a detailed description of steps 701 to 704.

[0155] In some embodiments, for the multi-modal training sample pair (image training sample I, text training sample T), a partial text occlusion sample corresponding to the text training sample T needs to be generated , that is, part of the information of the text training sample T is occluded to generate the partial text occlusion sample .

[0156] Then, based on the text prediction training task, the partial text occlusion sample and the image training sample in the multi-modal training sample are used as a text prediction training sample pair (image training sample I, partial text occlusion sample ), and the multi-granularity mixed output of the text prediction training sample pair is calculated and the feature-level mixed output .

[0157] Next, use the text prediction function , based on the multi-granularity mixed output of the text prediction training sample pair , the feature-level mixed output and the partial text occlusion sample , the occluded prediction text is calculated as shown in the following formula (5).

[0158] (5)

[0159] where represents the predicted text, and the text prediction function is a common text prediction function in multi-task learning, often used to combine different prediction tasks (such as text classification, regression, etc.) in a network model. This function processes text based on classical deep learning frameworks (such as multi-task neural networks).

[0160] Next, based on the occluded prediction text and the corresponding text training sample T, the text prediction loss value corresponding to the text prediction training task is obtained using a common loss function (such as cross-entropy loss, etc.) .

[0161] Step 503: Based on the multi-modal contrast training task, calculate the multi-modal contrast loss value of the multi-granularity mixed output and the feature-level mixed output.

[0162] The following is a detailed description of step 503.

[0163] In some embodiments, the cross-modal contrast learning (CCL) task aims to train a multi-modal recognition model to learn to extract useful information from different modal data, that is, given a pair of positive example text-image pairs and a pair of negative example text-image pairs, the model needs to maximize the similarity of the positive example text-image pairs while minimizing the similarity of the negative example text-image pairs, as described in detail below.

[0164] Referring to Figure 8 , based on the text prediction training task, calculate the text prediction loss value of the multi-granularity mixed output and the feature-level mixed output, including the following steps 801 to step 802.

[0165] Step 801: Obtain the incorrect training text corresponding to the text training sample in the multi-modal training sample pair.

[0166] Step 802: Based on the multi-modal contrast training task, use the incorrect training text and the image training sample in the multi-modal training sample as a text contrast sample pair, and calculate the contrast multi-granularity mixed output and the contrast feature-level mixed output of the text contrast sample pair.

[0167] Step 803: Use the similarity function to calculate the positive multi-modal contrast similarity of the multi-granularity mixed output and the feature-level mixed output, and calculate the negative multi-modal contrast similarity of the contrast multi-granularity mixed output and the contrast feature-level mixed output.

[0168] Step 804: Obtain a multi-modal contrast loss value based on the positive multi-modal contrast similarity and the negative multi-modal contrast similarity.

[0169] The following provides a detailed description of Steps 801 to 804.

[0170] In some embodiments, for a multi-modal training sample pair (image training sample I, text training sample T), an incorrect training text corresponding to the text training sample T needs to be generated , that is, the image training sample I does not correspond to the incorrect training text .

[0171] Next, take the multi-modal training sample pair (image training sample I, text training sample T) as a positive example sample pair, and denote its corresponding multi-granularity mixed output as and the feature-level mixed output as ; and take the incorrect training text and the image training sample in the multi-modal training sample as a negative example text comparison sample pair (image training sample I, incorrect training text ), and calculate the contrast multi-granularity mixed output and the contrast feature-level mixed output .

[0172] Then, use the similarity function to calculate the positive multi-modal contrast similarity of the multi-granularity mixed output as shown in the following formula (6). As shown in the following formula (6).

[0173] (6)

[0174] And calculate the negative multi-modal contrast similarity of the contrast multi-granularity mixed output and the contrast feature-level mixed output as shown in the following formula (7).

[0175] (7)

[0176] Among them, and respectively represent the similarities of positive and negative example image-text pairs, and the similarity function represents a function for calculating similarity. This function is used to calculate similarity, usually based on methods such as cosine similarity or Euclidean distance, and is widely used in information retrieval and recommendation systems.

[0177] After that, based on the positive multi-modal contrast similarity and negative multi-modal contrast similarity , obtain the multi-modal contrast loss value corresponding to the multi-modal contrast training task , which is described in detail as follows

[0178] Refer to Figure 9 , based on the positive multi-modal contrast similarity and the negative multi-modal contrast similarity, obtain the multi-modal contrast loss value, including the following steps 901 to 903

[0179] Step 901: Use the probability distribution function to obtain the positive probability corresponding to the positive multi-modal contrast similarity and the negative probability corresponding to the negative multi-modal contrast similarity

[0180] Step 902: Accumulate the positive probability and the negative probability to obtain the multi-modal probability value

[0181] Step 903: Based on the ratio of the positive probability and the multi-modal probability value, and then perform logarithmic processing to obtain the multi-modal contrast loss value

[0182] The following is a detailed description of steps 901 to 903

[0183] In some embodiments, after obtaining the positive multi-modal contrast similarity and the negative multi-modal contrast similarity , based on the probability distribution function, obtain the positive probability corresponding to the positive multi-modal contrast similarity and the negative probability corresponding to the negative multi-modal contrast similarity softmax(), and obtain the positive probability corresponding to the positive multi-modal contrast similarity and the negative probability corresponding to the negative multi-modal contrast similarity .

[0184] Next, accumulate the positive probability and the negative probability to obtain the multi-modal probability value, and then based on the ratio of the positive probability and the multi-modal probability value, perform logarithmic processing to obtain the multi-modal contrast loss value As shown in the following formula (8)

[0185] (8)

[0186] Step 504: Based on the image-text matching loss value, the text prediction loss value, and the multi-modal contrast loss value, obtain the multi-modal loss

[0187] The following is a detailed description of step 504

[0188] In some embodiments, during the pre-training process of the multi-modal recognition model, obtain the image-text matching loss value corresponding to the image-text matching training task , the text prediction loss value corresponding to the text prediction training task and the multi-modal contrast loss value corresponding to the multi-modal contrast training task After that, the multi-modal loss will be further combined as follows and is described in detail below

[0189] Referring to Figure 10 based on the image-text matching loss value, the text prediction loss value, and the multi-modal contrast loss value, the multi-modal loss is obtained, including the following steps 1001 to 1004

[0190] Step 1001: Obtain the text prediction weight and the multi-modal contrast weight

[0191] Step 1002: Based on the product of the text prediction weight and the text prediction loss value, obtain the text prediction weighted loss value

[0192] Step 1003: Based on the product of the multi-modal contrast weight and the multi-modal contrast loss value, obtain the multi-modal contrast weighted loss value

[0193] Step 1004: Accumulate the image-text matching loss value, the text prediction weighted loss value, and the multi-modal contrast weighted loss value to obtain the multi-modal loss

[0194] The following provides a detailed description of steps 1001 to 1004

[0195] In some embodiments, based on the text prediction weight corresponding to the text prediction training task and the multi-modal contrast weight corresponding to the multi-modal contrast training task respectively perform weighted multiplication on the text prediction loss value and the multi-modal contrast loss value to obtain the corresponding text prediction weighted loss value and multi-modal contrast weighted loss value, and then accumulate the image-text matching loss value, the text prediction weighted loss value, and the multi-modal contrast weighted loss value to obtain the multi-modal loss as shown in the following formula (9)

[0196] (9)

[0197] Through the above steps 501 to 504, the multi-modal recognition model is forced to learn the overall semantic matching relationship between images and texts using the image-text matching loss, strengthening the global consistency across modalities. The text prediction loss is used to drive the multi-modal recognition model to mine fine-grained image-text associations (such as the correspondence between local objects and descriptions) by predicting occluded texts, enhancing the complementary reasoning ability between modalities. The multi-modal contrast loss is used to improve the model's sensitivity to cross-modal feature differences by distinguishing positive and negative sample pairs, avoiding semantic confusion. When jointly optimizing the three, the multi-modal recognition model needs to simultaneously consider the global matching, local generation, and feature discriminability objectives, promoting the dynamic collaboration of different granularity features (multi-granularity hybrid output) and multi-level modal interactions (feature-level hybrid output) during training. Thus, without manual intervention, it adaptively learns a more robust cross-modal joint representation, ultimately improving the generalization and interpretability of the multi-modal recognition model in complex downstream tasks (such as fine-grained retrieval, visual question answering), and further effectively enhancing the recognition accuracy and reliability of the multi-modal recognition model for multi-modal data in practice.

[0198] In some embodiments, after obtaining the multi-modal loss the multi-modal recognition model can be optimized by Stochastic Gradient Descent (SGD) or other optimization algorithms. During the training process, the multi-modal recognition model will automatically adjust the weights to minimize the joint loss function, thereby improving its performance in downstream tasks.

[0199] Refer to Figure 11 which is a schematic diagram of the training process of a multi-modal recognition model provided by an embodiment of the present application. As Figure 11 shown, it demonstrates the image-text data processing and training process based on the multi-modal recognition model. The process starts with the collection and preprocessing of image-text data, and the feature extraction module is used to obtain the feature representations of images and texts respectively (such as CNN for extracting image features, BERT for generating text embeddings). Subsequently, it enters the expert model training stage: multi-granularity expert mixing (FE+CE) is achieved through a dynamic weight allocation mechanism (such as the ESM model), and model fusion is performed in combination with a feature-level gating mechanism (MOE-based MLP) to solve the problem of insufficient alignment of a single granularity. The pre-training tasks include image-text matching (ITM), text prediction (MTL), and cross-modal contrast learning (CCL), and the learning ability of the multi-modal recognition model for multi-granularity associations is improved by jointly optimizing the loss function. Finally, the model is fine-tuned for specific application scenarios (such as image retrieval, semantic understanding) through downstream tasks. The overall process reflects the collaborative innovation of the multi-granularity alignment and feature-level mixing strategies, taking into account both computational efficiency and accuracy advantages.

[0200] Refer to Figure 12 which is a schematic block diagram of the training process of a multi-modal recognition model provided by an embodiment of the present application. As Figure 12As shown, it presents a data processing flow architecture based on a multi-modal recognition model, which generally includes three core parts: an input layer, a feature expert module, and a multi-granularity output module. The input layer receives multi-modal data such as text, video, and audio through a single-stream structure. After being encoded by the common sense knowledge layer, it is input into the feature expert module. This module contains multiple parallel expert models (such as "certificate expert", "supplementary certificate expert", "generalization expert"), and through a dynamic weight allocation mechanism (such as the ESM model), it performs fine-grained fusion of features between modalities, and introduces a feature-level gating mechanism (MOE-based MLP) to achieve a balance between computational efficiency and expressive ability. The fused features are further input into the pending domain expert network module, and the final result is generated through the multi-granularity relevant feature output module.

[0201] In addition, the embodiments of this application propose a multi-modal data recognition method. Referring to Figure 13 , it is an optional flowchart of the multi-modal data recognition method provided by the embodiments of this application. Figure 13 The method in Figure 13 may include but is not limited to steps 1301 to 1304. At the same time, it can be understood that this embodiment does not specifically limit the order of steps 1301 to 1304 in

[0202] Step 1301: In response to a multi-modal recognition processing task, obtain multi-modal processing data.

[0203] Step 1302: Based on the multi-modal recognition processing task, input the multi-modal processing data into the multi-modal recognition model trained by the multi-modal recognition model training method to perform data recognition processing, and obtain target recognition data.

[0204] The following will describe steps 1301 to 1302 in detail.

[0205] In some embodiments, after obtaining the trained multi-modal recognition model through the above-mentioned training process of the multi-modal recognition model, in actual application, in response to an actual multi-modal recognition processing task, input the multi-modal processing data corresponding to the multi-modal recognition processing task into the multi-modal recognition model to perform corresponding data recognition processing, and obtain accurate target recognition data for the multi-modal recognition processing task.

[0206] Among them, the multi-modal recognition processing task can be cross-modal retrieval, visual reasoning, behavior recognition, etc.

[0207] The multi-modal recognition model training method, multi-modal data recognition method and related devices proposed in the embodiments of the present application. The method includes: First, obtain multiple multi-modal training sample pairs, where the multi-modal training sample pairs include image training samples and corresponding text training samples; Secondly, based on the multi-granularity output model in the multi-modal recognition model, perform position encoding processing on the image training sample pairs to obtain position-encoded images. Based on the first processing granularity, divide the position-encoded images into multiple granularity division images, extract features from each granularity division image one by one to obtain granularity division image features, obtain granularity division image features based on all granularity division image features and text training samples, and obtain first feature data based on all granularity division image features. Perform second feature extraction on each multi-modal training sample pair to obtain second feature data. Obtain an expert selection model, input the multi-modal training sample pairs into the expert selection model for image-text association processing to obtain the first feature weight corresponding to the first feature extraction and the second feature weight corresponding to the second feature extraction. Based on the product of the first feature weight and the output of the first feature extraction, plus the product of the second feature weight and the output of the second feature extraction, obtain a multi-granularity mixed output, where the processing granularities of the first feature extraction and the second feature extraction are different; Next, based on the feature-level mixing model in the multi-modal recognition model, input the multi-modal training sample pairs into the feature-level mixing model for recognition processing to obtain multiple feature expert model outputs, where at least two feature expert model outputs are different. Obtain a gating weight matrix, based on the product of the gating weight matrix and the multiple feature expert model outputs, and perform normalization processing to obtain gating parameters. Accumulate the products of all feature expert model outputs and the corresponding gating parameters to obtain a feature-level mixed output; Finally, based on the image-text matching training task, calculate the image-text matching loss value between the multi-granularity mixed output and the feature-level mixed output. Based on the text prediction training task, calculate the text prediction loss value between the multi-granularity mixed output and the feature-level mixed output. Based on the multi-modal contrast training task, calculate the multi-modal contrast loss value between the multi-granularity mixed output and the feature-level mixed output. Based on the image-text matching loss value, the text prediction loss value and the multi-modal contrast loss value, obtain a multi-modal loss, and update the multi-modal recognition model based on the multi-modal loss.

[0208] Embodiments of the present application capture multi-level correlations across multi-modal data from both fine-grained and coarse-grained processing perspectives through dual-channel feature extraction with different granularities, enabling the model to simultaneously adapt to coarse-grained retrieval and fine-grained inference tasks. The feature-level hybrid model is used to screen key features through a dynamic gating mechanism, enhancing the model's ability to model complex modal relationships. Based on the joint loss optimization of multi-granularity hybrid output and feature-level hybrid output, the multi-modal recognition model autonomously balances the modal alignment objectives at different granularities during training, thereby achieving a more comprehensive understanding of the implicit associations between images and texts. Furthermore, when using the trained multi-modal recognition model to perform the recognition of multi-modal data, the recognition accuracy of multi-modal data is effectively improved. In addition, by introducing position encoding, the spatial distribution relationship of objects in the perceptual image training samples is improved, avoiding local feature confusion. Through fine-grained partitioning (such as segmenting the image into local regions of different sizes), differential processing of local details is achieved, enabling the model to capture key local features and then fuse the multi-granularity visual features and text features for each local feature separately, forcing the model to consider the correlations of each partition block in the fine-grained level during cross-modal alignment, and thus paying more attention to the feature information of the fine-grained local areas. Moreover, the pre-trained Expert Selection Model (ESM) adaptively assigns weights to fine-grained and coarse-grained feature extraction based on the semantic relevance of image-text pairs, enabling the model to flexibly adjust the focus of attention for different input data (for example, strengthening local details in complex scenes and emphasizing global semantics in overall descriptions). Then, by weighted fusion of the features of the two granularities, the complementarity and cooperation of multi-level information are achieved, retaining both the context coherence of the coarse-grained features and integrating the precise local alignment ability of the fine-grained features. This dynamic fusion mechanism enables the model to autonomously balance the modal alignment requirements at different granularities, avoiding feature biases caused by fixed weights, and thus more comprehensively capturing the implicit associations between images and texts in cross-modal tasks and enhancing the model's generalization expression ability for diverse semantic relationships.In addition, the image-text matching loss is used to force the multi-modal recognition model to learn the overall semantic matching relationship between images and texts, strengthening the global consistency across modalities. The text prediction loss drives the multi-modal recognition model to mine fine-grained image-text associations (such as the correspondence between local objects and descriptions) by predicting occluded texts, enhancing the complementary reasoning ability between modalities. The multi-modal contrast loss improves the model's sensitivity to cross-modal feature differences by distinguishing positive and negative sample pairs, avoiding semantic confusion. When jointly optimizing the three, the multi-modal recognition model needs to simultaneously consider the global matching, local generation, and feature discriminability objectives, enabling dynamic collaboration of different granularity features (multi-granularity mixed output) and multi-level modality interactions (feature-level mixed output) during training. Thus, without manual intervention, it adaptively learns a more robust cross-modal joint representation, ultimately enhancing the generalization and interpretability of the multi-modal recognition model in complex downstream tasks (such as fine-grained retrieval and visual question answering), and effectively improving the recognition accuracy and reliability of the multi-modal recognition model for multi-modal data in practice.

[0209] The embodiment of the present application also provides a multi-modal recognition model training device, which can implement the above multi-modal recognition model training method. Referring to Figure 14 , the device 1400 includes:

[0210] A sample pair acquisition module 1410, configured to acquire multiple multi-modal training sample pairs, where the multi-modal training sample pairs include image training samples and corresponding text training samples;

[0211] A multi-granularity processing module 1420, configured to perform first feature extraction on each multi-modal training sample pair based on the multi-granularity output model in the multi-modal recognition model to obtain first feature data, perform second feature extraction on each multi-modal training sample pair to obtain second feature data, and perform hybrid processing on each first feature data and the corresponding second feature data to obtain a multi-granularity hybrid output, where the processing granularities of the first feature extraction and the second feature extraction are different;

[0212] A feature hybrid processing module 1430, configured to perform feature-level hybrid processing on each multi-modal training sample pair based on the feature-level hybrid model in the multi-modal recognition model to obtain a feature-level hybrid output corresponding to each multi-modal training sample;

[0213] A model update module 1440, configured to calculate a multi-modal loss based on the multi-granularity hybrid output and the feature-level hybrid output, and update the multi-modal recognition model based on the multi-modal loss.

[0214] In some embodiments, the multi-granularity processing module 1420 is further configured to:

[0215] Obtain an expert selection model, input a multi-modal training sample pair into the expert selection model for image-text association processing, and obtain a first feature weight corresponding to the first feature extraction and a second feature weight corresponding to the second feature extraction;

[0216] Based on the product of the first feature weight and the output of the first feature extraction, plus the product of the second feature weight and the output of the second feature extraction, obtain a multi-granularity hybrid output.

[0217] In some embodiments, the multi-granularity processing module 1420 is further configured to:

[0218] Perform position encoding processing on the image training sample pair to obtain a position-encoded image;

[0219] Based on the first processing granularity, divide the position-encoded image into multiple granularity-divided images;

[0220] Extract features from each granularity-divided image one by one to obtain granularity-divided image features;

[0221] Based on all granularity-divided image features and text training samples, obtain granularity-divided image features, and based on all granularity-divided image features, obtain first feature data.

[0222] In some embodiments, the feature hybrid processing module 1430 is further configured to:

[0223] Input the multi-modal training sample pair into a feature-level hybrid model for recognition processing to obtain multiple feature expert model outputs, and at least two feature expert model outputs are different;

[0224] Obtain a gating weight matrix, based on the product of the gating weight matrix and multiple feature expert model outputs, and perform normalization processing to obtain gating parameters;

[0225] Accumulate the products of all feature expert model outputs and their corresponding gating parameters to obtain a feature-level hybrid output.

[0226] In some embodiments, the model update module 1440 is further configured to:

[0227] Based on the image-text matching training task, calculate the image-text matching loss value of the multi-granularity hybrid output and the feature-level hybrid output;

[0228] Based on the text prediction training task, calculate the text prediction loss value of the multi-granularity hybrid output and the feature-level hybrid output;

[0229] Based on the multi-modal contrast training task, calculate the multi-modal contrast loss value of the multi-granularity hybrid output and the feature-level hybrid output;

[0230] Based on the image - text matching loss value, the text prediction loss value, and the multi - modal contrast loss value, a multi - modal loss is obtained.

[0231] In some embodiments, the model update module 1440 is further configured to:

[0232] Based on the image - text matching training task, use the matching score function to calculate the image - text matching scores of the multi - granularity mixed output and the feature - level mixed output;

[0233] Based on the image - text matching scores, calculate the image - text matching loss value.

[0234] In some embodiments, the model update module 1440 is further configured to:

[0235] Obtain partial text occlusion samples corresponding to the text training samples in the multi - modal training sample pair;

[0236] Based on the text prediction training task, use the partial text occlusion samples and the image training samples in the multi - modal training sample pair as the text prediction training sample pair, and calculate the multi - granularity mixed output and the feature - level mixed output of the text prediction training sample pair;

[0237] Use the text prediction function to calculate the occluded predicted text based on the multi - granularity mixed output, the feature - level mixed output, and the partial text occlusion samples of the text prediction training sample pair;

[0238] Based on the occluded predicted text and the corresponding text training samples, calculate the text prediction loss value.

[0239] In some embodiments, the model update module 1440 is further configured to:

[0240] Obtain incorrect training texts corresponding to the text training samples in the multi - modal training sample pair;

[0241] Based on the multi - modal contrast training task, use the incorrect training texts and the image training samples in the multi - modal training sample pair as the text contrast sample pair, and calculate the contrast multi - granularity mixed output and the contrast feature - level mixed output of the text contrast sample pair;

[0242] Use the similarity function to calculate the positive multi - modal contrast similarity of the multi - granularity mixed output and the feature - level mixed output, and calculate the negative multi - modal contrast similarity of the contrast multi - granularity mixed output and the contrast feature - level mixed output;

[0243] Based on the positive multi - modal contrast similarity and the negative multi - modal contrast similarity, obtain the multi - modal contrast loss value.

[0244] In some embodiments, the model update module 1440 is further configured to:

[0245] Using the probability distribution function, obtain the positive probability corresponding to the positive multi-modal contrast similarity and the negative probability corresponding to the negative multi-modal contrast similarity;

[0246] Accumulate the positive probability and the negative probability to obtain the multi-modal probability value;

[0247] Based on the ratio of the positive probability to the multi-modal probability value, and then perform logarithmic processing to obtain the multi-modal contrast loss value.

[0248] In some embodiments, the model update module 1440 is further configured to:

[0249] Obtain the text prediction weight and the multi-modal contrast weight;

[0250] Based on the product of the text prediction weight and the text prediction loss value, obtain the text prediction weighted loss value;

[0251] Based on the product of the multi-modal contrast weight and the multi-modal contrast loss value, obtain the multi-modal contrast weighted loss value;

[0252] Accumulate the graphic-text matching loss value, the text prediction weighted loss value, and the multi-modal contrast weighted loss value to obtain the multi-modal loss.

[0253] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, the specific implementation manners of the multi-modal recognition model training device are basically the same as those of the above multi-modal recognition model training method, and will not be elaborated here.

[0254] In the embodiments of the present application, the multi-modal recognition model training device captures the multi-level correlation across multi-modal data from the perspectives of fine-grained processing and coarse-grained processing through dual-channel feature extraction with different granularities, enabling the model to adapt to both coarse-grained retrieval and fine-grained reasoning tasks simultaneously. The feature-level hybrid model is used to screen key features through a dynamic gating mechanism, enhancing the model's ability to model complex modal relationships. Based on the joint loss optimization of the multi-granularity hybrid output and the feature-level hybrid output, the multi-modal recognition model autonomously balances the modal alignment objectives at different granularities during training, thereby achieving a more comprehensive understanding of the implicit correlation between images and texts. Furthermore, when using the trained multi-modal recognition model to perform the recognition of multi-modal data, the recognition accuracy of multi-modal data is effectively improved; in addition, by introducing position encoding, the spatial distribution relationship of objects in the perceptual image training samples is improved, avoiding local feature confusion, and through fine-grained partitioning (such as dividing the image into local regions of different sizes), differential processing of local details is achieved, enabling the model to capture key local features and then fuse the multi-granularity visual features and text features for each local feature separately, forcing the model to consider the correlation of each partition block in the fine-grained level during cross-modal alignment, and thus paying more attention to the feature information of the fine-grained local area; and, using the pre-trained expert selection model (ESM) to adaptively allocate the weights of fine-grained and coarse-grained feature extraction based on the semantic correlation of image-text pairs, enabling the model to flexibly adjust the focus of attention for different input data (for example, strengthening local details in complex scenarios and emphasizing global semantics in overall descriptions), and then using weighted fusion of the two granularity features to achieve the complementarity and collaboration of multi-level information, retaining both the context coherence of the coarse-grained features and integrating the precise local alignment ability of the fine-grained features. This dynamic fusion mechanism enables the model to autonomously balance the modal alignment requirements at different granularities, avoiding feature biases caused by fixed weights, and thus more comprehensively capturing the implicit correlation between images and texts in cross-modal tasks and enhancing the model's generalization expression ability for diverse semantic relationships;In addition, the image-text matching loss is used to force the multi-modal recognition model to learn the overall semantic matching relationship between images and texts, strengthening the global consistency across modalities. The text prediction loss is used to drive the multi-modal recognition model to mine fine-grained image-text associations (such as the correspondence between local objects and descriptions) by predicting occluded texts, enhancing the complementary reasoning ability between modalities. The multi-modal contrast loss is used to improve the model's sensitivity to cross-modal feature differences and avoid semantic confusion by distinguishing positive and negative sample pairs. When jointly optimizing the three, the multi-modal recognition model needs to simultaneously consider the global matching, local generation, and feature discriminability objectives, enabling the dynamic cooperation of different granularity features (multi-granularity hybrid output) and multi-level modal interactions (feature-level hybrid output) during training. Thus, without manual intervention, it can adaptively learn a more robust cross-modal joint representation, ultimately improving the generalization and interpretability of the multi-modal recognition model in complex downstream tasks (such as fine-grained retrieval and visual question answering), and effectively enhancing the recognition accuracy and reliability of the multi-modal recognition model for multi-modal data in practice.

[0255] The embodiment of the present application also provides an electronic device, including:

[0256] At least one memory;

[0257] At least one processor;

[0258] At least one program;

[0259] The program is stored in the memory, and the processor executes at least one program to implement the multi-modal recognition model training method described above in the present application. This electronic device can be any intelligent terminal including mobile phones, tablets, personal digital assistants (PDAs for short), in-vehicle computers, etc.

[0260] Please refer to Figure 15 , Figure 15 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0261] A processor 1501, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0262] The memory 1502 can be implemented in the form of a ROM (ReadOnly Memory), a static storage device, a dynamic storage device, or a RAM (RandomAccess Memory), etc. The memory 1502 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1502 and are called by the processor 1501 to execute the multi-modal recognition model training method of the embodiments of this application;

[0263] The input / output interface 1503 is used to implement information input and output;

[0264] The communication interface 1504 is used to implement communication and interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0265] The bus 1505 transmits information between various components of the device (such as the processor 1501, the memory 1502, the input / output interface 1503, and the communication interface 1504);

[0266] Among them, the processor 1501, the memory 1502, the input / output interface 1503, and the communication interface 1504 achieve communication connections with each other inside the device through the bus 1505.

[0267] The embodiments of this application also provide a storage medium. The storage medium is a computer-readable storage medium. This storage medium stores a computer program, and when this computer program is executed by a processor, it implements the above-mentioned multi-modal recognition model training method.

[0268] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0269] The embodiments described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0270] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.

[0271] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0272] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0273] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0274] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0275] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0276] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0277] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0278] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0279] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A multimodal recognition model training method, characterized in that: The method comprises: Acquire a plurality of multimodal training sample pairs, wherein the multimodal training sample pairs include image training samples and corresponding text training samples; Based on the multi-granularity output model in the multi-modal recognition model, a first feature extraction is performed on each of the multi-modal training sample pairs to obtain first feature data, a second feature extraction is performed on each of the multi-modal training sample pairs to obtain second feature data, and each of the first feature data and the corresponding second feature data are mixed to obtain a multi-granularity mixed output, wherein the processing granularity of the first feature extraction and the second feature extraction is different; Based on the feature-level mixing model in the multimodal recognition model, each of the multimodal training samples is subjected to feature-level mixing processing to obtain a feature-level mixing output corresponding to each of the multimodal training samples; Calculating a multimodal loss based on the multi-granularity mixed output and the feature-level mixed output, and updating a multimodal recognition model based on the multimodal loss; The mixing of each of the first feature data and the corresponding second feature data to obtain a multi-granularity mixed output includes: Acquire an expert selection model, input the multimodal training sample pair into the expert selection model for image-text association processing, and obtain a first feature weight corresponding to the first feature extraction and a second feature weight corresponding to the second feature extraction; Based on the product of the first feature weight and the first feature extraction output, plus the product of the second feature weight and the second feature extraction output, the multi-granularity mixed output is obtained; The multi-layer perceptron of the multi-modal recognition model is a multi-layer perceptron of a hybrid expert module, and the feature-level hybrid processing is performed on each of the multi-modal training samples to obtain a feature-level hybrid output corresponding to each of the multi-modal training samples, including: Input the multimodal training samples into the feature-level hybrid model for recognition processing to obtain multiple feature expert model outputs, at least two of the feature expert model outputs are different; Obtaining a gating weight matrix, based on the product of the gating weight matrix and the outputs of the plurality of feature expert models, and performing normalization processing to obtain a gating parameter; The products of all the feature expert model outputs and the corresponding gating parameters are accumulated to obtain the feature-level mixed output.

2. The multimodal recognition model training method according to claim 1, characterized in that: The step of extracting a first feature from each of the multimodal training samples to obtain first feature data includes: Performing position coding processing on the image training sample pairs to obtain position coded images; Based on a first processing granularity, dividing the position-encoded image into a plurality of granularity-divided images; Extracting features from each of the granularity-divided images one by one to obtain granularity-divided image features; The granularity-divided image features are obtained based on all the granularity-divided image features and the text training samples, and the first feature data are obtained based on all the granularity-divided image features.

3. The multimodal recognition model training method according to claim 1, characterized in that: The calculating the multimodal loss based on the multi-granularity mixed output and the feature-level mixed output includes: Based on the image-text matching training task, calculating the image-text matching loss value of the multi-granularity mixed output and the feature-level mixed output; Based on the text prediction training task, calculating the text prediction loss value of the multi-granularity mixed output and the feature-level mixed output; Based on the multimodal contrast training task, a multimodal contrast loss value of the multi-granularity mixed output and the feature-level mixed output is calculated; The multimodal loss is obtained based on the image-text matching loss value, the text prediction loss value and the multimodal contrast loss value.

4. The multimodal recognition model training method according to claim 3, characterized in that: The calculating of the image-text matching loss values ​​of the multi-granularity mixed output and the feature-level mixed output based on the image-text matching training task includes: Based on the image-text matching training task, using a matching score function, calculating the image-text matching scores of the multi-granularity mixed output and the feature-level mixed output; The image-text matching loss value is calculated based on the image-text matching score.

5. The multimodal recognition model training method according to claim 3, characterized in that: The step of calculating the text prediction loss value of the multi-granularity mixed output and the feature-level mixed output based on the text prediction training task includes: Acquire a portion of text occlusion samples corresponding to the text training samples in the multimodal training sample pair; Based on the text prediction training task, the partial text occlusion sample and the image training sample in the multimodal training sample are used as a text prediction training sample pair, and the multi-granularity mixed output and the feature-level mixed output of the text prediction training sample pair are calculated; Using a text prediction function, based on the multi-granularity mixed output of the text prediction training sample pair, the feature-level mixed output and the partial text occlusion sample, calculate the occlusion prediction text; The text prediction loss value is calculated based on the occluded predicted text and the corresponding text training sample.

6. The multimodal recognition model training method according to claim 3, characterized in that: The multimodal contrast loss value of the multi-granularity mixed output and the feature-level mixed output is calculated based on the multimodal contrast training task, including: Acquire an erroneous training text corresponding to the text training sample in the multimodal training sample pair; Based on the multimodal contrast training task, the erroneous training text and the image training sample in the multimodal training sample are used as a text contrast sample pair, and a contrast multi-granularity mixed output and a contrast feature-level mixed output of the text contrast sample pair are calculated; Using a similarity function, calculating a positive multimodal comparison similarity between the multi-granularity mixed output and the feature-level mixed output, and calculating a negative multimodal comparison similarity between the comparison multi-granularity mixed output and the comparison feature-level mixed output; The multimodal contrast loss value is obtained based on the positive multimodal contrast similarity and the negative multimodal contrast similarity.

7. The multimodal recognition model training method according to claim 6, characterized in that: The obtaining the multimodal contrast loss value based on the positive multimodal contrast similarity and the negative multimodal contrast similarity includes: Using a probability distribution function, obtaining a positive probability corresponding to the positive multimodal comparison similarity and a negative probability corresponding to the negative multimodal comparison similarity; Accumulating the positive probability and the negative probability to obtain a multimodal probability value; Based on the ratio of the forward probability and the multimodal probability value, logarithmic processing is performed to obtain the multimodal contrast loss value.

8. The multimodal recognition model training method according to claim 3, characterized in that: The obtaining the multimodal loss based on the image-text matching loss value, the text prediction loss value and the multimodal contrast loss value includes: Get text prediction weights and multimodal comparison weights; Obtaining a text prediction weighted loss value based on the product of the text prediction weight and the text prediction loss value; Obtaining a multimodal contrast weighted loss value based on the product of the multimodal contrast weight and the multimodal contrast loss value; The multimodal loss is obtained by accumulating the image-text matching loss value, the text prediction weighted loss value, and the multimodal comparison weighted loss value.

9. A multimodal data recognition method, characterized in that: The method comprises: In response to the multimodal recognition processing task, obtaining multimodal processing data; Based on the multimodal recognition processing task, the multimodal processing data is input into the multimodal recognition model trained by the multimodal recognition model training method as described in claim 1 to perform data recognition processing to obtain target recognition data.

10. A multimodal recognition model training device, characterized in that: The device comprises: A sample pair acquisition module, used to acquire a plurality of multimodal training sample pairs, wherein the multimodal training sample pairs include image training samples and corresponding text training samples; A multi-granularity processing module, for performing a first feature extraction on each of the multi-modal training sample pairs based on a multi-granularity output model in a multi-modal recognition model to obtain first feature data, performing a second feature extraction on each of the multi-modal training sample pairs to obtain second feature data, and performing a mixed processing on each of the first feature data and the corresponding second feature data to obtain a multi-granularity mixed output, wherein the processing granularities of the first feature extraction and the second feature extraction are different; A feature mixing processing module, used to perform feature-level mixing processing on each of the multimodal training samples based on the feature-level mixing model in the multimodal recognition model to obtain a feature-level mixing output corresponding to each of the multimodal training samples; A model updating module, configured to calculate a multimodal loss based on the multi-granularity mixed output and the feature-level mixed output, and update a multimodal recognition model based on the multimodal loss; The mixing of each of the first feature data and the corresponding second feature data to obtain a multi-granularity mixed output includes: Acquire an expert selection model, input the multimodal training sample pair into the expert selection model for image-text association processing, and obtain a first feature weight corresponding to the first feature extraction and a second feature weight corresponding to the second feature extraction; Based on the product of the first feature weight and the first feature extraction output, plus the product of the second feature weight and the second feature extraction output, the multi-granularity mixed output is obtained; The multi-layer perceptron of the multi-modal recognition model is a multi-layer perceptron of a hybrid expert module, and the feature-level hybrid processing is performed on each of the multi-modal training samples to obtain a feature-level hybrid output corresponding to each of the multi-modal training samples, including: Input the multimodal training samples into the feature-level hybrid model for recognition processing to obtain multiple feature expert model outputs, at least two of the feature expert model outputs are different; Obtaining a gating weight matrix, based on the product of the gating weight matrix and the outputs of the plurality of feature expert models, and performing normalization processing to obtain a gating parameter; The products of all the feature expert model outputs and the corresponding gating parameters are accumulated to obtain the feature-level mixed output.

11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, it implements the multimodal recognition model training method described in any one of claims 1 to 8 or the multimodal data recognition method described in claim 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multimodal recognition model training method as described in any one of claims 1 to 8 or the multimodal data recognition method as described in claim 9.

Citation Information

Patent Citations

  • Multi-granularity and multi-mode fused artwork image description generation method

    CN115082693A

  • Multi-granularity image-text matching method and system based on deep fusion

    CN117093692A