Multi-modal learning model training method, task processing method, equipment and medium

By introducing hybrid expert connectors in multimodal learning for bidirectional alignment, the problem of knowledge representation splitting in the existing technology is solved, the alignment effect and expression ability of cross-modal tasks are improved, and the symmetric flow of information and the mining of complementary information is realized.

CN120164060APending Publication Date: 2025-06-17ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235193.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing multimodal learning methods often rely on connectors in specific modal and specific directions, resulting in knowledge representation splits, limiting the model's alignment effect and expression ability for cross-modal tasks.

Method used

A multimodal learning model training method is proposed. By obtaining the training data of the first mode and the second mode, using the pre-trained single-modal model to generate eigenvectors, and input these eigenvectors into the hybrid expert connector to be trained for training, a multimodal learning model that can be bidirectionally aligned is obtained.

Benefits of technology

Through the two-way alignment mechanism, the alignment effect and expression ability of the multimodal learning model for cross-modal tasks is improved, and the symmetric flow of information between the two modes is realized, which can more comprehensively explore the complementary information between the modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164060A_ABST
    Figure CN120164060A_ABST
Patent Text Reader

Abstract

The invention provides a model training method for multi-modal learning, a task processing method, equipment and a medium, and belongs to the technical field of artificial intelligence, and the model training method comprises the steps: obtaining training data of a first modal and training data of a second modal; inputting the training data of the first modal into a pre-trained first modal model to obtain a first feature vector; inputting the training data of the second modal into a pre-trained second modal model to obtain a second feature vector; inputting the feature vector into a to-be-trained hybrid expert connector to train the hybrid expert connector to obtain a multi-modal learning model; wherein the multi-modal learning model comprises a first modal model, a second modal model and a trained hybrid expert connector, and the trained hybrid expert connector is used for performing alignment from a first modal to a second modal and performing alignment from the second modal to the first modal. According to the method, the alignment effect and the expression capability of the model for the cross-modal task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method, a task processing method, a device, and a medium for multi-modal learning. Background Art

[0002] With the continuous development of deep learning, multi-modal learning, as a technology capable of integrating different types of data, has received extensive attention in the industry. In related technologies, there have been methods for multi-modal learning by aligning the source modality to the target modality. However, this method usually relies on connectors for specific modalities and specific directions, so it can only achieve one-way alignment from the source modality to the target modality, which easily leads to the negative effect of fragmented knowledge representation in the model's multi-modal learning, thus limiting the alignment effect and expression ability of the model for cross-modal tasks. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a model training method, a task processing method, a device, and a medium for multi-modal learning, aiming to improve the ability of an artificial intelligence model to form a unified multi-modal knowledge representation in multi-modal learning, thereby improving the alignment effect and expression ability of the model for cross-modal tasks.

[0004] To achieve the above object, in the first aspect of the embodiments of this application, a model training method for multi-modal learning is proposed. The method includes:

[0005] Obtain training data of a first modality and training data of a second modality;

[0006] Input the training data of the first modality into a pre-trained first modality model to obtain a first feature vector; and input the training data of the second modality into a pre-trained second modality model to obtain a second feature vector;

[0007] Input the first feature vector and the second feature vector into a to-be-trained mixture-of-experts connector to train the mixture-of-experts connector, and obtain a multi-modal learning model;

[0008] Wherein, the multi-modal learning model includes the first modality model, the second modality model, and the trained mixture-of-experts connector. The trained mixture-of-experts connector is respectively connected to the first modality model and the second modality model, and the trained mixture-of-experts connector is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0009] In some embodiments, the step of inputting the first feature vector and the second feature vector into a to-be-trained mixture-of-experts connector to train the mixture-of-experts connector includes:

[0010] Input the target feature vector into the to-be-trained mixture-of-experts connector; wherein, the target feature vector is either the first feature vector or the second feature vector.

[0011] Based on the mixture-of-experts connector, perform a prediction task and a contrast task on the target feature vector to align the first modality to the second modality and to align the second modality to the first modality.

[0012] Wherein, when the target feature vector is the first feature vector, the prediction task is to generate a predicted feature vector in the second modality based on the first feature vector, and the contrast task is to compare the predicted feature vector in the second modality with the second feature vector; when the target feature vector is the second feature vector, the prediction task is to generate a predicted feature vector in the first modality based on the second feature vector, and the contrast task is to compare the predicted feature vector in the first modality with the first feature vector.

[0013] In some embodiments, the performing the prediction task and the contrast task on the target feature vector based on the mixture-of-experts connector includes:

[0014] At the t-th time step, compare the predicted feature vector in the second modality with the second feature vector based on the mixture-of-experts connector; wherein, the mixture-of-experts connector generates the predicted feature vector in the second modality based on the first feature vector.

[0015] At the (t + 1)-th time step, compare the predicted feature vector in the first modality with the first feature vector based on the mixture-of-experts connector; wherein, the mixture-of-experts connector generates the predicted feature vector in the first modality based on the second feature vector.

[0016] In some embodiments, the method further includes:

[0017] Calculate the multimodal alignment loss between the predicted feature vector in the second modality and the second feature vector; wherein, the multimodal alignment loss includes a contrast loss and an L2 squared difference loss.

[0018] Based on the contrast loss and the L2 squared difference loss, optimize the mixture-of-experts connector for aligning the first modality to the second modality and optimize the mixture-of-experts connector for aligning the second modality to the first modality.

[0019] In some embodiments, the training data of the first modality includes first image data in the image modality, and the training data of the second modality includes first text data in the text modality, and the first image data is associated with the first text data; the first modality model includes a visual model, and the second modality model includes a language model; the first feature vector includes a first image feature vector obtained by the visual model learning the first image data, and the second feature vector includes a first text feature vector obtained by the language model learning the first text data;

[0020] The training of the hybrid expert connector to be trained by inputting the first feature vector and the second feature vector into the hybrid expert connector to be trained includes:

[0021] Input the first image feature vector and the first text feature vector into the hybrid expert connector to be trained;

[0022] Based on the hybrid expert connector, process the first image feature vector and the first text feature vector respectively to perform alignment from the image modality to the text modality and alignment from the text modality to the image modality.

[0023] To achieve the above object, a second aspect of the embodiments of the present application proposes a task processing method for multi-modal learning, and the method includes:

[0024] Obtain the data to be processed for the multi-modal task; wherein, the data to be processed includes the processed data of the first modality and the processed data of the second modality;

[0025] Input the processed data of the first modality and the processed data of the second modality into a preset multi-modal learning model to obtain the processing result of the multi-modal task;

[0026] Wherein, the first modality model in the multi-modal learning model learns the processed data of the first modality to obtain a first feature vector; the second modality model in the multi-modal learning model learns the processed data of the second modality to obtain a second feature vector; the hybrid expert connector in the multi-modal learning model is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0027] In some embodiments,

[0028] The processed data of the first modality includes second image data of the image modality, and the processed data of the second modality includes second text data of the text modality. The second image data is associated with the second text data. The first modality model includes a visual model, and the second modality model includes a language model. The first feature vector includes a second image feature vector obtained by the visual model learning the second image data, and the second feature vector includes a second text feature vector obtained by the language model learning the second text data.

[0029] The method further includes:

[0030] Inputting the second image feature vector and the second text feature vector into a mixture-of-experts connector in the multi-modal learning model;

[0031] Based on the mixture-of-experts connector, generating a predicted feature vector of the second text feature vector based on the second image feature vector, and comparing the predicted feature vector of the second text feature vector with the second text feature vector based on the mixture-of-experts connector to perform alignment from the image modality to the text modality;

[0032] Based on the mixture-of-experts connector, generating a predicted feature vector of the second image feature vector based on the second text feature vector, and comparing the predicted feature vector of the second image feature vector with the second image feature vector to perform alignment from the text modality to the image modality.

[0033] To achieve the above object, a third aspect of the embodiments of the present application provides a model training device for multi-modal learning. The device includes:

[0034] A first acquisition module, configured to acquire training data of the first modality and training data of the second modality;

[0035] A single-modal learning module, configured to input the training data of the first modality into a pre-trained first modality model to obtain a first feature vector; and input the training data of the second modality into a pre-trained second modality model to obtain a second feature vector;

[0036] A multi-modal alignment module, configured to input the first feature vector and the second feature vector into a to-be-trained mixture-of-experts connector to train the mixture-of-experts connector to obtain a multi-modal learning model;

[0037] Among them, the multimodal learning model includes the first modal model, the second modal model, and a trained mixture-of-experts connector. The trained mixture-of-experts connector is respectively connected to the first modal model and the second modal model, and the trained mixture-of-experts connector is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0038] To achieve the above object, a fourth aspect of the embodiments of the present application provides a task processing device for multimodal learning. The device includes:

[0039] A second acquisition module, configured to acquire data to be processed for a multimodal task; among them, the data to be processed includes processed data of a first modality and processed data of a second modality;

[0040] A multimodal task processing module, configured to input the processed data of the first modality and the processed data of the second modality into a preset multimodal learning model to obtain a processing result of the multimodal task;

[0041] Among them, the first modal model in the multimodal learning model learns the processed data of the first modality to obtain a first feature vector; the second modal model in the multimodal learning model learns the processed data of the second modality to obtain a second feature vector; the mixture-of-experts connector in the multimodal learning model is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0042] To achieve the above object, a fifth aspect of the embodiments of the present application provides a multimodal learning device. The multimodal learning device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method described in the first aspect above, and / or implements the method described in the second aspect above.

[0043] To achieve the above object, a sixth aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect above, and / or implements the method described in the second aspect above.

[0044] To achieve the above object, a seventh aspect of the embodiments of the present application provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method provided in the first aspect above, and / or implements the method provided in the second aspect above.

[0045] In the embodiments of the present application, by obtaining the training data of the first modality and the training data of the second modality, then inputting the training data of the first modality into the pre-trained first modality model to obtain the first feature vector, and inputting the training data of the second modality into the pre-trained second modality model to obtain the second feature vector. After that, by inputting the first feature vector and the second feature vector into the hybrid expert connector to be trained to train the hybrid expert connector, a multi-modal learning model including the first modality model, the second modality model, and the trained hybrid expert connector is obtained. Among them, the trained hybrid expert connector is used to perform alignment from the first modality to the second modality, and alignment from the second modality to the first modality.

[0046] In this way, compared with the training method of the multi-modal learning model in the related art by aligning the source modality to the target modality (unidirectional alignment), the embodiments of the present application perform bidirectional alignment processing between the first modality and the second modality by introducing a hybrid expert connector, which can further capture the bidirectional semantic relationship and complex interaction between different modalities compared with unidirectional alignment, thereby effectively improving the multi-modal alignment effect of the multi-modal learning model. Moreover, the embodiments of the present application can also achieve symmetric information flow between the two modalities through the bidirectional alignment of the hybrid expert connector, which helps the multi-modal learning model to more comprehensively mine the complementary information between modalities. More importantly, the embodiments of the present application perform deep fusion between multi-modalities based on bidirectional alignment, improving the alignment effect and expression ability of the multi-modal learning model for cross-modal tasks. Especially when processing multi-modal data, the multi-modal learning model can better consider the reverse influence of the target modality on the source modality.

[0047] In addition, the embodiments of the present application also process the training data through pre-trained single-modal models (the first modality model and the second modality model) to obtain the corresponding feature vectors, which can not only reduce the high computing cost and large-scale data requirements for training the model from scratch, but also retain the rich information in the training data under each modality, avoiding the possible loss of local details when directly modeling multi-modal alignment. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic flowchart of the steps in some embodiments of the model training method for multi-modal learning provided by the embodiments of the present application;

[0049] Figure 2 It is a schematic diagram of the network structure of the hybrid expert model involved in the model training method for multi-modal learning provided by the embodiments of the present application;

[0050] Figure 3 It is a schematic diagram of the model architecture involved in the model training method for multi-modal learning provided by the embodiments of the present application;

[0051] Figure 4 For Figure 1 the detailed step flow diagram of step S103 in

[0052] Figure 5 For Figure 4 the detailed step flow diagram of step S402 in

[0053] Figure 6 the schematic diagram of the training strategy based on alternating gradient descent involved in the model training method for multi-modal learning provided by the embodiments of the present application in some embodiments;

[0054] Figure 7 the step flow diagram of the model training method for multi-modal learning provided by the embodiments of the present application in some other embodiments;

[0055] Figure 8 the schematic diagram of the alternating alignment algorithm for pairing image and text features involved in the model training method for multi-modal learning provided by the embodiments of the present application in some embodiments;

[0056] Figure 9 the schematic diagram of a model verification result involved in the model training method for multi-modal learning provided by the embodiments of the present application in some embodiments;

[0057] Figure 10 the schematic diagram of another model verification result involved in the model training method for multi-modal learning provided by the embodiments of the present application in some embodiments;

[0058] Figure 11 the step flow diagram of the task processing method for multi-modal learning provided by the embodiments of the present application in some embodiments;

[0059] Figure 12 the step flow diagram of the task processing method for multi-modal learning provided by the embodiments of the present application in some other embodiments;

[0060] Figure 13 the structural schematic diagram of the model training device for multi-modal learning provided by the embodiments of the present application;

[0061] Figure 14 the structural schematic diagram of the task processing device for multi-modal learning provided by the embodiments of the present application;

[0062] Figure 15 is the hardware structural schematic diagram of the multi-modal learning device provided by the embodiments of the present application. Specific Embodiments

[0063] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division from that in the device or a different order from that in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0066] First, the overall concept of the model training method for multi-modal learning provided in the embodiments of the present application will be described.

[0067] With the development of deep learning, multi-modal learning, as a technology capable of integrating different types of data, has received extensive attention. Multi-modal data usually includes various forms such as images, texts, and voices. These data contain different information in different modalities. How to effectively align these modalities (that is, fuse the information in different modalities into the same representation space) is an important issue in current multi-modal learning research.

[0068] The current methods mainly adopted in multi-modal learning are divided into two types. One is to directly model the modality alignment from multi-modal data, and the other is to use a cross-modal connector to align the source modality to the target modality. However, these methods each face different challenges.

[0069] Among them, the challenges of directly learning multimodal alignment from multimodal data mainly include: The training methods for directly modeling multimodal alignment from multimodal data tend to jointly model multimodal data from scratch without the support of prior knowledge or pre-trained models. For example, pre-trained multimodal models that connect text and images (Contrastive Language-Image Pre-training, CLIP) and general multimodal foundation models such as BEiT-3 are all used to model multimodal alignment directly from multimodal data. Among them, the CLIP model is trained using a large number of image and text pairs through contrastive learning to learn the alignment relationship between images and text. The CLIP model consists of two main parts: the Text Encoder and the Image Encoder, which convert text and images into low-dimensional vector representations respectively, and evaluate the correlation between images and text by calculating the similarity between these vectors.

[0070] Although directly learning multimodal alignment from multimodal data can indeed effectively achieve cross-modal alignment, this method not only requires high training costs, but also is difficult to utilize the rich intra-modal knowledge of existing pre-trained models. More critically, this method of directly modeling multimodal alignment on multimodal data will also lead to the loss of fine-grained information within a single modality.

[0071] In addition, the methods in related technologies that use cross-modal connectors to align one modality to the target modality usually rely on connectors for specific modalities and specific directions, which will result in the fragmentation of the knowledge representation of the final model and reduce the computational efficiency, limiting the model's ability to form a unified multimodal representation. For example, the multimodal large model (Large Language and Vision Assistant, LLaVA) consists of an LLM (Vicuna) and an image encoder (ViT-L / 14 of CLIP), and a linear layer W is added in the middle to convert the image features into features with the same dimension as the text Embedding; the vision-language large model Flamingo that interleaves text and images also includes a vision encoder (Vision Transformer) and a language encoder (MosaicGPT), as well as a Perceiver Resampler module for modality interaction; and the multimodal large model BLIP-2 that freezes the vision model and the language model also connects the frozen image encoder and the LLM through the Q-Former module.

[0072] Combined with the above, although the method of directly performing cross-modal alignment from multi-modal data can align modalities bidirectionally, on the one hand, such methods usually require training from scratch on multi-modal data, facing high training costs and large-scale data requirements. On the other hand, directly modeling multi-modal alignment easily causes the model to lose some local details (for example, text-guided visual representation learning methods limit the model's ability to obtain fine-grained visual information). In response to this, the embodiments of this application propose cross-modal knowledge alignment to solve these problems, that is, the embodiments of this application alleviate the expensive computational costs and large-scale data requirements by using pre-trained single-modal models (such as visual models and language models), while retaining more rich information within the modality. In addition, although the method of using a cross-modal connector to align the source modality to the target modality can use a cross-modal connection module to align two single-modal models. However, this method is limited to unidirectional connection, thus only enabling alignment from the source modality to the target modality (such as image-to-text alignment), and unable to capture the bidirectional semantic relationship and complex interaction between the two modalities. This may lead to asymmetry in information flow, restricting the model's ability to mine complementary information between modalities. Moreover, the unidirectional connection method easily ignores the reverse influence of the target modality on the source modality when processing multi-modal data, thereby affecting the alignment effect. In response to this, the embodiments of this application propose to introduce a mixture-of-experts connector to achieve bidirectional information flow and deep fusion based on a bidirectional alignment mechanism, thereby improving the alignment effect and expression ability of cross-modal tasks.

[0073] Based on this, the embodiments of this application provide a model training method, task processing method, device, and medium for multi-modal learning, aiming to improve the ability of an artificial intelligence model to perform multi-modal learning to form a unified multi-modal knowledge representation, thereby improving the alignment effect and expression ability of the model for cross-modal tasks.

[0074] The embodiments of this application obtain training data of the first modality and training data of the second modality, and then input the training data of the first modality into a pre-trained first modality model to obtain a first feature vector, and input the training data of the second modality into a pre-trained second modality model to obtain a second feature vector. Then, by inputting the first feature vector and the second feature vector into a to-be-trained mixture-of-experts connector to train the mixture-of-experts connector, a multi-modal learning model including the first modality model, the second modality model, and the trained mixture-of-experts connector is obtained. Among them, the trained mixture-of-experts connector is used to perform alignment from the first modality to the second modality, and to perform alignment from the second modality to the first modality.

[0075] Thus, compared with the training method of the multi-modal learning model in the related art that aligns the source modality with the target modality (one-way alignment), the embodiment of the present application performs two-way alignment processing between the first modality and the second modality by introducing a mixture-of-experts connector. Compared with one-way alignment, this can further capture the two-way semantic relationship and complex interaction between different modalities, thereby effectively improving the multi-modal alignment effect of the multi-modal learning model. Moreover, through the two-way alignment of the mixture-of-experts connector, the embodiment of the present application can also achieve symmetric information flow between the two modalities, which helps the multi-modal learning model to more comprehensively mine the complementary information between modalities. More importantly, the embodiment of the present application performs deep fusion between multi-modalities based on two-way alignment, improving the alignment effect and expression ability of the multi-modal learning model for cross-modal tasks. Especially when processing multi-modal data, the multi-modal learning model can better consider the reverse influence of the target modality on the source modality.

[0076] In addition, the embodiment of the present application also processes the training data through pre-trained single-modal models (the first-modal model and the second-modal model) to obtain corresponding feature vectors, which can not only reduce the high computing cost and large-scale data requirements for training the model from scratch, but also retain the rich information in the training data under each modality, avoiding the possible loss of local details when directly modeling multi-modal alignment.

[0077] Next, the model training method, device, computer device, and storage medium for multi-modal learning provided by the embodiment of the present application are specifically described through the following embodiments, and first, the model training method for multi-modal learning provided by the embodiment of the present application is described in detail.

[0078] It should be noted that the model training method for multimodal learning provided in the embodiments of the present application and the task processing method for multimodal learning that will also be described in detail later relate to the field of artificial intelligence technology and can be applied to terminals, server sides, or software running on terminals or server sides. In some embodiments, the terminal can be a computer device such as a smart phone, a tablet computer, a notebook computer, or a desktop computer. The server side can be the background server terminal device of a vehicle, which can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The software can be an application that implements the method provided in the embodiments of the present application, a computer program, and a storage medium carrying the computer program, etc. It should be understood that based on different design requirements of actual applications, in different feasible embodiments, the terminals, server sides, and software that apply the method provided in the embodiments of the present application can, of course, also be other forms not listed here, and the method provided in the embodiments of the present application does not specifically limit this.

[0079] In addition, the present application can also be used in many general or special computer system environments or configurations. For example: vehicles, personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0080] For the convenience of understanding and description, in the following text, the terminal device applying the model training method and task processing method of multi-modal learning provided in the embodiments of the present application is taken as an example to describe each specific embodiment of the present application in detail. The implementation of any of the above-mentioned forms of the subject applying the model training method of multi-modal learning provided in the embodiments of the present application can refer to the process of the model training method and task processing method of multi-modal learning applied by the terminal device described later.

[0081] It should be noted that in each specific implementation manner of the present application, when it comes to relevant processing that needs to be carried out according to data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0082] Please refer to Figure 1 , Figure 1 which is the schematic diagram of the step flow of the model training method of multi-modal learning provided in the embodiments of the present application in some embodiments. It should be understood that although Figure 1 and the subsequent other schematic diagrams of step flows show the execution order of some method steps, due to different design requirements in actual applications, the model training method of multi-modal learning provided in the embodiments of the present application can of course adopt an execution order different from that shown in the figures. That is, Figure 1 the order of the method steps shown does not constitute a limitation on the execution logic order of the model training method of multi-modal learning provided in the embodiments of the present application, and any reasonable changes based on Figure 1 the order of the method steps shown should be included in the protection scope of the model training method of multi-modal learning provided in the embodiments of the present application.

[0083] As Figure 1 shown, in some embodiments, when a vehicle applies the model training method of multi-modal learning provided in the embodiments of the present application, it may include step S101 to step S103.

[0084] Step S101: Obtain the training data of the first modality and the training data of the second modality.

[0085] During the process of the terminal device training the multi-modal learning model, it first obtains the training data of the first modality and the training data of the second modality.

[0086] It should be noted that the terminal device can specifically obtain multi-modal training data through the COCO dataset (the COCO dataset is a large and rich object detection, segmentation, and captioning dataset) and the flickr30k image annotation dataset. The multi-modal training data includes, but is not limited to, data of types such as images, text, and audio. The training data of the first modality is one of the types of data such as images, text, and audio, and the training data of the second modality is another type of data among images, text, and audio.

[0087] Step S102: Input the training data of the first modality into a pre-trained first modality model to obtain a first feature vector; and input the training data of the second modality into a pre-trained second modality model to obtain a second feature vector.

[0088] The terminal device inputs the obtained training data of the first modality into the pre-trained first modality model, so as to learn the training data of the first modality based on the prior knowledge of the first modality model by the first modality model itself and output a first feature vector. Moreover, the terminal device also inputs the obtained training data of the second modality into the pre-trained second modality model, so as to learn the training data of the second modality based on the prior knowledge of the second modality model by the second modality model itself and output a second feature vector.

[0089] In some embodiments, both the first modality model and the second modality model are advanced single-modal models in the related art. For example, the first modality model is a large visual model, the second modality model is a large language model, and so on. The terminal device uses the advanced single-modal models as the representation modules, so as to directly utilize the knowledge representation capabilities of these single-modal models in their respective fields to learn the training data and obtain the corresponding feature vectors. This can further improve the generalization ability of the finally trained multi-modal learning model and reduce the resource consumption of training from scratch (directly performing cross-modal alignment from multi-modal data).

[0090] Step S103: Input the first feature vector and the second feature vector into a to-be-trained hybrid expert connector to train the hybrid expert connector and obtain a multi-modal learning model; wherein, the multi-modal learning model includes the first modality model, the second modality model, and the trained hybrid expert connector. The trained hybrid expert connector is respectively connected to the first modality model and the second modality model, and the trained hybrid expert connector is used to perform alignment from the first modality to the second modality and perform alignment from the second modality to the first modality.

[0091] After obtaining the first feature vector and the second feature vector, the terminal device further inputs the first feature vector and the second feature vector into the hybrid expert connector to be trained, so as to train the process of multi-modal alignment of the hybrid expert connector, and thus obtain a multi-modal learning model including a first modal model, a second modal model, and a trained hybrid expert connector. Among them, when the terminal device trains the hybrid expert connector, through a bidirectional alignment mechanism, the hybrid expert connector processes the first feature vector and the second feature vector respectively, so as to perform alignment from the first modality to the second modality, and perform alignment from the second modality to the first modality. In this way, the trained hybrid expert connector in the multi-modal learning model can be used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0092] It should be noted that the hybrid expert connector can be a cross-modal alignment module based on the Mixture of Experts (MoE). Please refer to Figure 2 , the Mixture of Experts (MoE) is a dynamic network structure designed to improve the computational efficiency and hierarchical representation ability of deep learning models by partitioning model parameters and tasks. The core idea of MoE is to organize multiple experts together and selectively activate some experts through a gating network (Router) module by weights, so as to achieve on-demand computing. In multi-modal tasks (such as vision-language alignment, cross-modal generation, audio-visual fusion, etc.), the data distributions and feature expression forms of each modality often have significant differences. Traditional single-model architectures may have difficulty taking into account the specific features of each modality when dealing with multi-modal tasks. Based on the dynamic network structure of the mixture of experts, it is possible to design exclusive expert modules (i.e., modality-specific expert modules) for different modalities to achieve targeted modeling, and at the same time, through a cross-modal mechanism, efficient fusion and alignment are achieved by shared expert modules. For example, modality-specific experts process single-modal features to improve the understanding of that modality, while shared experts are responsible for handling cross-modal tasks to promote joint understanding between different modalities.

[0093] By introducing a hybrid expert connector to process the first feature vector and the second feature vector, the terminal device can retain rich information within different modalities and perform two-way alignment of the two modalities (alignment from the first modality to the second modality and alignment from the second modality to the first modality, specifically such as alignment from natural language to vision and alignment from vision to natural language).

[0094] In some embodiments, the training data of the first modality obtained by the terminal device may include first image data in the image modality, while the training data of the second modality obtained by the terminal device may include first text data in the text modality. Moreover, the first image data and the first text data are correlated with each other. For example, the first text data may be text data for characterizing the features of the first image data, or the first text data may also be text data indicating the processing of the first image data. In addition, the first modality model may include a visual model, and the second modality model may include a language model. In this case, the first feature vector may include a first image feature vector obtained by the terminal device through learning the first image data by the visual model, while the second feature vector may include a first text feature vector obtained by the terminal device through learning the first text data by the language model.

[0095] Based on this, the above step S103: inputting the first feature vector and the second feature vector into the hybrid expert connector to be trained to train the hybrid expert connector may include:

[0096] Inputting the first image feature vector and the first text feature vector into the hybrid expert connector to be trained;

[0097] Based on the hybrid expert connector, respectively process the first image feature vector and the first text feature vector to perform alignment from the image modality to the text modality and alignment from the text modality to the image modality.

[0098] When the terminal device trains the hybrid expert connector for bidirectional alignment between the image modality and the text modality, it first inputs both the first image feature vector and the first text feature vector learned by using the pre-trained single-modal model into the hybrid expert connector, and then processes the first image feature vector and the first text feature vector respectively through the hybrid expert connector to perform cross-modal interaction between the image modality and the text modality, so as to perform alignment from the image modality to the text modality and alignment from the text modality to the image modality. Among them, the terminal device may perform alignment from the image modality to the text modality during the process of processing the first image feature vector, and perform alignment from the text modality to the image modality during the process of processing the first text feature vector.

[0099] Exemplarily, for a cross-modal alignment problem, the terminal device uses pre-trained single-modal models such as a large language model and a large visual model (i.e., the above first modality model and second modality model) as modality encoders and connects them with a bidirectional connector (hybrid expert connector). Taking image-text alignment as an example, during the training process of the multi-modal learning model, the terminal device decomposes the image-text alignment into image-to-text alignment and text-to-image alignment that are alternately performed in time. AsFigure 3 As shown, the model architecture trained by the terminal device is mainly divided into three parts: 1. A pre-trained visual model (the first modality model), 2. A pre-trained language model (the second modality model), and 3. A bidirectional hybrid expert connection module (Bi-MoE) (hybrid expert connector). For the input image-text pair, the terminal device encodes the corresponding feature vectors using the visual model and the language model respectively, and then the terminal device inputs the feature vectors corresponding to the image and the text into the Bi-MoE. Through the Bi-MoE, cross-modal interaction between the image modality and the text modality is performed based on the feature vectors, and finally alignment is achieved in the latent space, that is, alignment from the image modality to the text modality and alignment from the text modality to the image modality). Among them, in the image-text pair input to the model architecture, the image data is the training data of the first modality, and the text data is the training data of the second modality.

[0100] It should be noted that the above visual model can be: a pure visual model trained only in the visual modality, such as DinoV2. The terminal device uses the visual model to process the training data of the image modality and can retain more information within the image modality. In addition, the above language model can be: an advanced language model trained only in the language modality, such as qwen2.5llama, etc. By selecting these language models, the terminal device can effectively capture rich semantic information within the language modality, thereby making full use of the context information and semantic characteristics of the language modality to provide high-quality language feature representations for the subsequent multi-modal alignment of the hybrid expert connector. In addition, the above bidirectional hybrid expert connection module Bi-MoE can be designed using the above classical mixture of experts model (MoE).

[0101] In the embodiments of the present application, the terminal device first obtains the training data of the first modality and the training data of the second modality, and then inputs the obtained training data of the first modality into the pre-trained first modality model. Thus, based on the first modality model, the prior knowledge of the model itself is used to learn the training data of the first modality, and a first feature vector is output. Moreover, the terminal device also inputs the obtained training data of the second modality into the pre-trained second modality model. Thus, based on the second modality model, the prior knowledge of the model itself is used to learn the training data of the second modality, and a second feature vector is output. Finally, the terminal device further inputs the first feature vector and the second feature vector into the hybrid expert connector to be trained, so as to train the multi-modal alignment process of the hybrid expert connector, and thus obtain a multi-modal learning model including the first modality model, the second modality model, and the trained hybrid expert connector. Among them, when the terminal device trains the hybrid expert connector, through a bidirectional alignment mechanism, the hybrid expert connector processes the first feature vector and the second feature vector respectively, so as to perform the alignment from the first modality to the second modality and the alignment from the second modality to the first modality. In this way, the trained hybrid expert connector in the multi-modal learning model can be used to perform the alignment from the first modality to the second modality and the alignment from the second modality to the first modality.

[0102] In this way, the embodiments of the present application introduce a hybrid expert connector to perform bidirectional alignment processing between the first modality and the second modality. Compared with unidirectional alignment, this can further capture the bidirectional semantic relationships and complex interactions between different modalities, thereby effectively improving the multi-modal alignment effect of the multi-modal learning model. Moreover, through the bidirectional alignment of the hybrid expert connector in the embodiments of the present application, the symmetric flow of information between the two modalities can be realized, which helps the multi-modal learning model to more comprehensively mine the complementary information between modalities. More importantly, the embodiments of the present application perform deep fusion between multi-modalities based on bidirectional alignment, improving the alignment effect and expression ability of the multi-modal learning model for cross-modal tasks. Especially when processing multi-modal data, the multi-modal learning model can better consider the reverse influence of the target modality on the source modality.

[0103] In addition, in the embodiments of the present application, the pre-trained single-modal models such as the first modality model and the second modality model are used to process the training data to obtain the corresponding feature vectors, which can not only reduce the high computing cost and large-scale data requirements for training the model from scratch, but also retain the rich information in the training data under each modality, avoiding the possible loss of local details when directly modeling multi-modal alignment.

[0104] During the process of the terminal device training the mixture-of-experts connector, the mixture-of-experts connector can perform different tasks on the feature vectors being processed respectively when processing the first feature vector and the second feature vector, so as to perform joint understanding between the first modality and the second modality, and achieve alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0105] Please refer to Figure 4 , Figure 4 for Figure 1 the schematic diagram of the refined step flow of step S103 in

[0106] In some embodiments, as Figure 4 shown, in the above step S103, "inputting the first feature vector and the second feature vector into the mixture-of-experts connector to be trained to train the mixture-of-experts connector" may include step S401 and step S402 shown as follows.

[0107] Step S401: Input the target feature vector into the mixture-of-experts connector to be trained; wherein, the target feature vector is any one of the first feature vector and the second feature vector.

[0108] During the process of the terminal device training the mixture-of-experts connector, the first feature vector in the first modality and the second feature vector in the second modality learned by using the pre-trained single-modal model will be used as the target feature vectors to be processed by the mixture-of-experts connector in sequence, and the target feature vectors will be input into the mixture-of-experts connector to be trained, so as to perform multi-modal alignment based on the target feature vectors to train the mixture-of-experts connector in the subsequent process.

[0109] Step S402: Based on the mixture-of-experts connector, perform a prediction task and a contrast task on the target feature vector to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0110] After the terminal device inputs the target feature vector into the mixture-of-experts connector, by controlling the mixture-of-experts connector to perform a prediction task and a contrast task on the target feature vector simultaneously, the mixture-of-experts connector is trained to perform two-way alignment between the first modality and the second modality, that is, to train the mixture-of-experts connector to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0111] It should be noted that when the target feature vector input to the hybrid expert connector is the first feature vector, the prediction task that the terminal device controls the hybrid expert connector to perform on the first feature vector is: generating a predicted feature vector in the second modality based on the first feature vector, while the comparison task that the terminal device controls the hybrid expert connector to perform on the first feature vector is: comparing the generated predicted feature vector in the second modality with the actual second feature vector. In addition, when the target feature vector input to the hybrid expert connector is the second feature vector, the prediction task that the terminal device controls the hybrid expert connector to perform on the second feature vector is: generating a predicted feature vector in the first modality based on the second feature vector, and the comparison task that the terminal device controls the hybrid expert connector to perform on the second feature vector is: comparing the generated predicted feature vector in the first modality with the actual first feature vector.

[0112] In some embodiments, when the hybrid expert connector is designed based on the mixture-of-experts model MOE, in order to enable MoE to perceive inputs of different modalities (the first feature vector and the second feature vector) and perform different tasks, namely the prediction task and the comparison task, the terminal device can adopt the Cross Embedding method to guide the gating network Router to select different expert models in the hybrid expert network according to the input modality and the task to be performed. Among them, Cross embedding consists of task embedding and modality embedding. Modality embedding is used to inform MoE of the modality type of the currently input feature, and task embedding is used to inform MoE of the task type to be performed on the currently input feature.

[0113] Exemplarily, assuming that the first feature vector is an image feature and the second feature vector is a text feature, then as described above Figure 2 As shown, the image feature also controls the execution of the prediction task of the hybrid expert connector by adding image embedding and L2 embedding, that is, predicting the text feature (generating a predicted text feature based on the image feature). At the same time, the image feature also controls the execution of the contrast learning task of the hybrid expert connector designed based on MOE by adding image modality embedding and contrast embedding, that is, comparing the text feature (comparing the predicted text feature with the actual text feature). Based on the same principle of controlling the hybrid expert connector to process the image feature, for the text feature, different tasks are also performed by adding text embedding, L2, and contrast Embedding.

[0114] It should be noted that the features (image features and text features) after adding the embedding will generate a set of weights through the gating network router. The terminal device can sort these weights and only select the outputs of the k experts with the highest weight values for aggregation as the final predicted features, that is, the predicted image features or the predicted text features mentioned above.

[0115] During the process of the terminal device training the hybrid expert connector, it can, at the same time step, control the hybrid expert connector to perform prediction tasks and contrast tasks on the same modality in parallel. In this way, the hybrid expert connector can be trained to achieve two-way alignment of different modalities in the previous and subsequent time steps. For example, at time step t, alignment from the first modality to the second modality is performed, and at time step t + 1, alignment from the second modality to the first modality is performed.

[0116] Please refer to Figure 5 , Figure 5 For Figure 4 the detailed step flow diagram of step S402 in

[0117] In some embodiments, as Figure 5 shown, in the above step S402, "performing a prediction task and a contrast task on the target feature vector based on the hybrid expert connector" may include step S501 and step S502 shown below.

[0118] Step S501: At the t-th time step, based on the hybrid expert connector, compare the predicted feature vector in the second modality with the second feature vector; wherein, the hybrid expert connector generates the predicted feature vector in the second modality based on the first feature vector.

[0119] In order for the terminal device to train the hybrid expert connector to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality at different time steps (t, t + 1,...) respectively, first at the t-th time step, control the hybrid expert connector to perform a prediction task and a contrast task on the first feature vector simultaneously, that is, based on cross-embedding, control the hybrid expert connector to generate the predicted feature vector in the second modality based on the first feature vector at the t-th time step, and compare the generated predicted feature vector in the second modality with the actual second feature vector, so as to perform alignment from the first modality to the second modality first at the t-th time step.

[0120] Step S502: At the t + 1-th time step, based on the hybrid expert connector, compare the predicted feature vector in the first modality with the first feature vector; wherein, the hybrid expert connector generates the predicted feature vector in the first modality based on the second feature vector.

[0121] At the (t + 1)-th time step, the terminal device controls the mixture-of-experts connector to perform a prediction task and a contrast task on the second feature vector simultaneously. That is, based on cross-embedding, the terminal device controls the mixture-of-experts connector to generate a predicted feature vector in the first modality based on the second feature vector at the (t + 1)-th time step, and compares the generated predicted feature vector in the first modality with the actual first feature vector, so as to perform alignment from the second modality to the first modality at the (t + 1)-th time step.

[0122] In some embodiments, in order to train the mixture-of-experts connector to achieve two-way alignment of different modalities, the terminal device can design a training strategy based on alternating gradient descent, and combine the dynamic activation ability of the bidirectional mixture-of-experts connection module Bi-MoE to alternately optimize in two directions from the first modality to the second modality and from the second modality to the first modality (specifically, such as from image to text and from text to image), so as to gradually achieve efficient alignment of multiple modalities.

[0123] Exemplarily, when the terminal device trains the mixture-of-experts connector for two-way alignment of images and texts, the image data and text data are respectively encoded by a visual model and a language model to obtain an image feature vector (the first feature vector) and a text feature vector (the second feature vector). Then, the terminal device controls the mixture-of-experts connector to perform alignment from image to text and from text to image at different time steps. Among them, as Figure 6 shown in the left image-to-text alignment process, at time step t, the terminal device processes the image feature vector through Bi-MoE to obtain a predicted text feature vector, and then compares the predicted text feature vector with the actual text feature vector and calculates the contrast loss and the L2 loss. After that, as Figure 6 shown in the right text-to-image alignment process, at time step t + 1, the terminal device performs the reverse process. First, the text feature vector is passed through Bi-MoE to obtain a predicted image feature vector, and then the predicted image feature vector is compared with the actual image feature vector and the contrast loss and the L2 loss between the actual image feature vector and the predicted image feature vector are calculated.

[0124] In this embodiment, during the process of training the hybrid expert connector by the terminal device, the hybrid expert connector is controlled to perform different tasks on the feature vectors being processed respectively when processing the first feature vector and the second feature vector, so as to perform joint understanding between the first modality and the second modality, and achieve alignment from the first modality to the second modality and alignment from the second modality to the first modality. Moreover, during the process of training the hybrid expert connector by the terminal device, at the same time step, the hybrid expert connector is controlled to perform a prediction task and a contrast task on the same modality in parallel, so that the hybrid expert connector can be trained to achieve two-way alignment of different modalities at two consecutive time steps. For example, alignment from the first modality to the second modality is performed at time step t, and alignment from the second modality to the first modality is performed at time step t+1.

[0125] Compared with the traditional method of training a multimodal learning model by aligning the source modality to the target modality based on unidirectional alignment, this embodiment improves the unidirectional cross-modal connector. By adopting the cross-modal alignment module Bi-MoE designed based on the mixture of experts (MoE) as the hybrid expert connector and training the hybrid expert connector to enable two-way alignment of different modules, the understanding of the interaction between modalities by the finally trained multimodal learning model can be effectively enhanced.

[0126] When the terminal device controls the hybrid expert connector to perform a prediction task and a contrast task on the feature vector, contrast learning improves the discriminability of the feature by maximizing the mutual information of positive samples and minimizing the shared information of negative samples, while the L2 loss ensures the semantic consistency between modalities by minimizing the error between the predicted feature and the target feature. Contrast learning and the L2 loss complement each other, optimizing both the discriminability of information and enhancing the precise cross-modal alignment.

[0127] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of the steps of the model training method for multimodal learning provided in an embodiment of this application in some other embodiments.

[0128] In some embodiments, as Figure 7 shown, the model training method for multimodal learning provided in an embodiment of this application may further include step S701 and step S702 as follows.

[0129] Step S701: Calculate the multimodal alignment loss between the predicted feature vector in the second modality and the second feature vector; wherein, the multimodal alignment loss includes a contrast loss and an L2 squared difference loss.

[0130] Step S702: Optimize the alignment of the hybrid expert connector from the first modality to the second modality based on the contrast loss and the L2 squared difference loss.

[0131] When the terminal device trains the hybrid expert connector for bidirectional alignment of different modalities, during the process of controlling the hybrid expert connector to perform prediction tasks and contrast tasks on the feature vector being processed, it can further calculate the contrast loss and L2 loss between the predicted feature vector and the true feature vector simultaneously, so as to optimize the process of multimodal alignment of the hybrid expert connector by combining the two losses.

[0132] Taking the terminal device training the hybrid expert connector for alignment from the first modality to the second modality as an example (the alignment from the second modality to the first modality is similar and can directly refer to this description), when the terminal device controls the hybrid expert connector to perform prediction tasks and contrast tasks on the first feature vector, after the hybrid expert connector generates the predicted feature vector in the second modality based on the first feature vector, it can simultaneously calculate the multimodal alignment loss between the generated predicted feature vector in the second modality and the actual second feature vector, that is, the contrast loss and the L2 squared difference loss between the predicted feature vector in the second modality and the actual second feature vector. Then, the terminal device optimizes the process of the hybrid expert connector performing prediction tasks and contrast tasks on the first feature vector through the contrast loss and the L2 squared difference loss together (such as adjusting the parameters used when generating the predicted feature vector, and adjusting the parameters used when comparing the predicted feature vector with the true feature vector, etc.), so as to optimize the alignment of the hybrid expert connector from the first modality to the second modality.

[0133] Exemplarily, when the terminal device trains a multimodal alignment model for pairing image and text features, the overall process includes three main stages: feature encoding, alternating alignment based on Bi-MoE, and joint optimization combining contrast loss and L2 loss.

[0134] As Figure 8 shown, when the terminal device performs feature encoding on the obtained image data and text data, for the input paired image data and text data (the image is the training data of the first modality, and the text is the training data of the second modality), the terminal device first encodes the image data I into an image feature vector f I (the first feature vector) through an image model (the first modality model, specifically a convolutional neural network for example), and encodes the text data T into a text feature vector f T (the second feature vector) through a language model (the second modality model). Then, these encoded features are used as the input for the alignment process.

[0135] After that, when the terminal device performs alternating alignment of the first modality and the second modality in the bidirectional expert model module Bi-MoE (the first modality is the image modality and the second modality is the text modality), it trains Bi-MoE to perform image-to-text alignment at odd time steps t. At this time, the terminal device takes the image feature vector f I and generates a predicted text feature vector through Bi-MoE and calculates the contrast loss and L2 loss between the predicted text feature vector and the real text feature vector f T to optimize the alignment effect. Also, it trains Bi-MoE to perform text-to-image alignment at even time steps t + 1. Similarly, the terminal device takes the text feature vector f T at this time and generates a predicted image feature vector through the Bi-MoE module and then calculates the contrast loss and L2 loss between the predicted image feature vector and the real image feature vector f I to optimize the alignment effect.

[0136] When the terminal device jointly optimizes Bi-MoE for image-text bidirectional alignment by combining the contrast loss and L2 loss, at each time step, the contrast loss is used to ensure that the features between the image modality and the text modality are well-distinguishable from each other, thereby enhancing the discriminative ability of the finally trained multi-modal learning model for different modal features. At the same time, the L2 loss is used to minimize the reconstruction error between the predicted feature vector and the real feature vector, thereby ensuring the semantic consistency between the image modality and the text modality. That is to say, by jointly optimizing the contrast loss and L2 loss when training Bi-MoE for image-text bidirectional alignment, the terminal device can enable the finally trained multi-modal learning model to achieve robust bidirectional alignment of image and text features.

[0137] In some embodiments, in order to verify the effectiveness of the multi-modal learning model finally trained by the terminal device, the terminal device performs image-text retrieval experiments on the COCO dataset and the flickr30k dataset. As Figure 9 shown, the terminal device trains on the COCO dataset and the Flickr30k dataset and tests on their standard test datasets. The finally obtained experimental results show that the multi-modal learning model (ALt-MoE) obtained by the model training method proposed in the embodiments of the present application outperforms the current state-of-the-art models.

[0138] In some other embodiments, such as Figure 10As shown, the terminal device also conducts experiments on audio-text retrieval in the Clotho dataset and the Audiocaps dataset. The final experimental results also show that the multi-modal learning model ALt-MoE obtained by the model training method proposed in the embodiments of the present application outperforms the current state-of-the-art models.

[0139] Next, various embodiments of the task processing method for multi-modal learning provided in the embodiments of the present application are presented.

[0140] Please refer to Figure 11 , Figure 11 which is a schematic flowchart of the steps in some embodiments of the task processing method for multi-modal learning provided in the embodiments of the present application.

[0141] In some embodiments, as Figure 11 shown, the task processing method for multi-modal learning provided in the embodiments of the present application may include step S1101 and step S1102 as follows.

[0142] Step S1101: Obtain the data to be processed for the multi-modal task; wherein, the data to be processed includes the processed data of the first modality and the processed data of the second modality.

[0143] It should be noted that the multi-modal tasks include but are not limited to vision-language alignment, cross-modal generation, audio-video fusion, etc. Among them, the data to be processed for each multi-modal task includes data of different modalities, and there are significant differences in the data distribution and feature expression forms of each modality.

[0144] The terminal device can perform any one or more of the multi-modal tasks such as vision-language alignment, cross-modal generation, audio-video fusion, etc. based on the multi-modal learning model finally trained in the various embodiments of the above model training method. And, the terminal device first obtains the data to be processed for the multi-modal task to be executed based on the multi-modal learning modality. Among them, the data to be processed obtained by the terminal device includes the processed data of the first modality and the processed data of the second modality.

[0145] Exemplarily, when the terminal device executes the multi-modal task of vision-language alignment, it first obtains the image data of the image modality (the first modality) and the text data of the text modality (the second modality), and uses the obtained image data and text data as the data to be processed for the vision-language alignment task to be executed currently.

[0146] Step S1102: Input the processed data of the first modality and the processed data of the second modality into a preset multi-modal learning model to obtain the processing result of the multi-modal task; wherein, the first modality model in the multi-modal learning model learns the processed data of the first modality to obtain a first feature vector; the second modality model in the multi-modal learning model learns the processed data of the second modality to obtain a second feature vector; the mixture-of-experts connector in the multi-modal learning model is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0147] The terminal device inputs all the acquired data to be processed into the multi-modal learning model. After processing the data to be processed through the multi-modal learning model, the processing result of the multi-modal task is output. Among them, the multi-modal learning modality first learns the processed data of the first modality based on the first modality model in its own model architecture to obtain a first feature vector, and learns the processed data of the second modality based on the second modality model in its own model architecture to obtain a second feature vector. Then, the multi-modal learning modality further uses the mixture-of-experts connector connecting the first modality model and the second modality model in its own model architecture to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality based on the first feature vector and the second feature vector.

[0148] In some embodiments, among the data to be processed for the multi-modal task acquired by the terminal device, the processed data of the first modality may include second image data in the image modality, and the processed data of the second modality may include second text data in the text modality, and the second image data is also associated with the second text data. Moreover, in the multi-modal learning model adopted by the terminal device, the first modality model may include a pre-trained vision model, and the second modality model may include a pre-trained language model. In this case, the above-mentioned first feature vector includes the second image feature vector obtained by the vision model learning the second image data, and the above-mentioned second feature vector includes the second text feature vector obtained by the language model learning the second text data.

[0149] It should be noted that when the terminal device executes any multi-modal task such as the above-mentioned vision-language alignment, cross-modal generation, audio-video fusion, etc., the acquired data to be processed may include second image data in the image modality and second text data in the text modality. Thus, when the terminal device uses the above-mentioned multi-modal learning model to process any of the above multi-modal tasks, it may involve performing alignment from the image modality to the text modality and alignment from the text modality to the image modality.

[0150] Please refer to Figure 12 , Figure 12Schematic diagram of the step flow of the task processing method for multimodal learning provided by the embodiments of the present application in some other embodiments.

[0151] In some embodiments, as Figure 12 shown, the task processing method for multimodal learning provided by the embodiments of the present application may further include steps S1201 to S1203 as shown below.

[0152] Step S1201: Input the second image feature vector and the second text feature vector into the mixture of experts connector in the multimodal learning model.

[0153] When the terminal device processes the data to be processed for the multimodal task based on the multimodal learning model, after the multimodal learning model uses the pre-trained visual model to learn the second image data to obtain the second image feature vector, and uses the pre-trained language model to train the second text data to obtain the second text feature vector, the second image feature vector and the second text feature vector are further input together into the mixture of experts connector in the multimodal learning model.

[0154] It should be noted that in the model architecture of the multimodal learning model, the mixture of experts connector is respectively connected to the first modality model (pre-trained visual model) and the second modality model (pre-trained text model). Therefore, the second image feature vector obtained by the visual model learning the second image data and the second text feature vector obtained by the text model training the second text data can both be automatically input into the mixture of experts connector.

[0155] Step S1202: Based on the mixture of experts connector, generate a predicted feature vector of the second text feature vector based on the second image feature vector, and compare the predicted feature vector of the second text feature vector with the second text feature vector based on the mixture of experts connector to perform alignment from the image modality to the text modality.

[0156] When the terminal device processes the data to be processed for the multimodal task based on the multimodal learning model, after inputting the second image feature vector into the mixture of experts connector, the mixture of experts connector simultaneously performs prediction and comparison tasks on the second image feature vector to perform alignment from the image modality to the text modality. For example, the mixture of experts connector first generates a predicted feature vector of the second text feature vector based on the second image feature vector based on cross-embedding, and further compares the generated predicted feature vector of the second text feature vector with the real second text feature vector, thereby performing alignment from the image modality to the text modality.

[0157] Step S1203: Based on the hybrid expert connector, generate a predicted feature vector of the second image feature vector based on the second text feature vector, and compare the predicted feature vector of the second image feature vector with the second image feature vector to perform alignment from the text modality to the image modality.

[0158] After the terminal device inputs the second text feature vector into the hybrid expert connector based on the multi-modal learning model, the hybrid expert connector simultaneously performs prediction and comparison tasks on the second text feature vector to perform alignment from the text modality to the image modality. For example, the hybrid expert connector first generates a predicted feature vector of the second image feature vector based on the second text feature vector through cross-embedding, and further compares the generated predicted feature vector of the second image feature vector with the true second image feature vector, thereby performing alignment from the text modality to the image modality.

[0159] In this embodiment, the terminal device executes any one or more multi-modal tasks such as visual-language alignment, cross-modal generation, and audio-video fusion based on the multi-modal learning model finally trained in each of the above model training methods. Among them, the terminal device first obtains the data to be processed for the multi-modal task to be executed currently, and the data to be processed includes the processed data of the first modality and the processed data of the second modality. Then, the terminal device inputs both the processed data of the first modality and the processed data of the second modality into the multi-modal learning model, so as to process the data to be processed through the multi-modal learning model and output the processing result of the multi-modal task. Herein, the multi-modal learning modality first learns the first feature vector based on the first modality model in its own model architecture for the processed data of the first modality, and learns the second feature vector based on the second modality model in its own model architecture for the processed data of the second modality; then, the multi-modal learning modality further performs alignment from the first modality to the second modality and alignment from the second modality to the first modality based on the first feature vector and the second feature vector through the hybrid expert connector connecting the first modality model and the second modality model in its own model architecture.

[0160] In this way, this embodiment performs a bidirectional alignment process between the first modality and the second modality through the hybrid expert connector of the multimodal learning model, which can further capture the bidirectional semantic relationship and complex interaction between different modalities compared to unidirectional alignment, thereby effectively improving the multimodal alignment effect of the multimodal learning model. Moreover, this embodiment can also realize the symmetrical flow of information between the two modalities through the bidirectional alignment of the hybrid expert connector in the multimodal learning model, which helps to more comprehensively mine the complementary information between different modalities. More importantly, this embodiment realizes the deep fusion between multimodalities by performing bidirectional alignment of different modalities based on the multimodal learning model, and also improves the alignment effect and expression ability of the multimodal learning model for cross-modal tasks, especially when processing the processed data of the first modality and the processed data of the second modality, the multimodal learning model can better consider the reverse influence of the target modality on the source modality.

[0161] In addition, this embodiment uses a pre-trained single-modal model to process the data to be processed through a multimodal learning model to obtain a corresponding feature vector, which can effectively utilize the existing prior knowledge of the pre-trained single-modal model and retain the rich information in each modality data, thereby being more conducive to the hybrid expert connector to perform bidirectional alignment of different modalities.

[0162] See also Figure 13 The embodiment of the present application also provides a multimodal learning model training device, which can implement the above multimodal learning model training method, and the device includes: a first acquisition module 1301, a single-modal learning module 1302 and a multimodal alignment module 1303. Among them,

[0163] A first acquisition module 1301 is used to acquire training data of a first modality and training data of a second modality;

[0164] The single-modality learning module 1302 is used to input the training data of the first modality into the pre-trained first modality model to obtain a first feature vector; and input the training data of the second modality into the pre-trained second modality model to obtain a second feature vector;

[0165] A multimodal alignment module 1303 is used to input the first feature vector and the second feature vector into a hybrid expert connector to be trained to train the hybrid expert connector and obtain a multimodal learning model;

[0166] Among them, the multimodal learning model includes the first modal model, the second modal model, and a trained mixture-of-experts connector. The trained mixture-of-experts connector is respectively connected to the first modal model and the second modal model, and the trained mixture-of-experts connector is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0167] In some embodiments, the multimodal alignment module 1303 is further configured to input a target feature vector into a mixture-of-experts connector to be trained; wherein the target feature vector is any one of the first feature vector and the second feature vector; and, based on the mixture-of-experts connector, perform a prediction task and a contrast task on the target feature vector to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality;

[0168] Among them, when the target feature vector is the first feature vector, the prediction task is to generate a predicted feature vector in the second modality based on the first feature vector, and the contrast task is to compare the predicted feature vector in the second modality with the second feature vector; when the target feature vector is the second feature vector, the prediction task is to generate a predicted feature vector in the first modality based on the second feature vector, and the contrast task is to compare the predicted feature vector in the first modality with the first feature vector.

[0169] In some embodiments, the multimodal alignment module 1303 is further configured to, at the t-th time step, compare the predicted feature vector in the second modality with the second feature vector based on the mixture-of-experts connector; wherein the mixture-of-experts connector generates the predicted feature vector in the second modality based on the first feature vector; and, at the (t + 1)-th time step, compare the predicted feature vector in the first modality with the first feature vector based on the mixture-of-experts connector; wherein the mixture-of-experts connector generates the predicted feature vector in the first modality based on the second feature vector.

[0170] In some embodiments, the model training apparatus for multimodal learning provided in the embodiments of the present application further includes a loss optimization module, configured to calculate a multimodal alignment loss between the predicted feature vector in the second modality and the second feature vector; wherein the multimodal alignment loss includes a contrast loss and an L2 squared difference loss; and, based on the contrast loss and the L2 squared difference loss, optimize the mixture-of-experts connector to perform alignment from the first modality to the second modality.

[0171] In some embodiments, the training data of the first modality includes first image data in the image modality, and the training data of the second modality includes first text data in the text modality, and the first image data is associated with the first text data; the first modality model includes a visual model, and the second modality model includes a language model; the first feature vector includes a first image feature vector obtained by the visual model learning the first image data, and the second feature vector includes a first text feature vector obtained by the language model learning the first text data;

[0172] The multi-modal alignment module 1303 is further configured to input the first image feature vector and the first text feature vector into a to-be-trained mixture-of-experts connector; and, based on the mixture-of-experts connector, process the first image feature vector and the first text feature vector respectively to perform alignment from the image modality to the text modality and alignment from the text modality to the image modality.

[0173] The specific implementation manner of the model training device for multi-modal learning provided by the embodiments of the present application is basically the same as the specific embodiments of the above-mentioned model training method for multi-modal learning, and will not be elaborated here.

[0174] Please refer to Figure 14 , the embodiments of the present application further provide a task processing device for multi-modal learning, which can implement the above-mentioned task processing method for multi-modal learning. The device includes: a second acquisition module 1401 and a multi-modal task processing module 1402. Wherein,

[0175] The second acquisition module 1401 is configured to acquire data to be processed for a multi-modal task; wherein, the data to be processed includes processing data of the first modality and processing data of the second modality;

[0176] The multi-modal task processing module 1402 is configured to input the processing data of the first modality and the processing data of the second modality into a preset multi-modal learning model to obtain a processing result of the multi-modal task;

[0177] Wherein, the first modality model in the multi-modal learning model learns the processing data of the first modality to obtain a first feature vector; the second modality model in the multi-modal learning model learns the processing data of the second modality to obtain a second feature vector; the mixture-of-experts connector in the multi-modal learning model is used to perform alignment from the first modality to the second modality and alignment from the second modality to the first modality.

[0178] In some embodiments, the processed data of the first modality includes second image data of the image modality, and the processed data of the second modality includes second text data of the text modality, and the second image data is associated with the second text data; the first modality model includes a pre-trained vision model, and the second modality model includes a pre-trained language model; the first feature vector includes a second image feature vector obtained by the vision model learning the second image data, and the second feature vector includes a second text feature vector obtained by the language model learning the second text data;

[0179] The multi-modal task processing module 1402 is further configured to input the second image feature vector and the second text feature vector into the mixture-of-experts connector in the multi-modal learning model; generate a predicted feature vector of the second text feature vector based on the second image feature vector by the mixture-of-experts connector, and compare the predicted feature vector of the second text feature vector with the second text feature vector based on the mixture-of-experts connector to perform alignment from the image modality to the text modality; and, generate a predicted feature vector of the second image feature vector based on the second text feature vector by the mixture-of-experts connector, and compare the predicted feature vector of the second image feature vector with the second image feature vector to perform alignment from the text modality to the image modality.

[0180] The specific implementation manner of the multi-modal learning task processing device provided in the embodiments of the present application is basically the same as the specific embodiments of the above multi-modal learning task processing method, and will not be elaborated herein.

[0181] The embodiments of the present application further provide a multi-modal learning device. The multi-modal learning device provided in the embodiments of the present application may be a computer device, and the computer device may specifically be an intelligent robot, a terminal device equipped with an embodied intelligent system, or may also be a computer device such as a smart phone, a tablet computer, a notebook computer, or a desktop computer. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above multi-modal learning model training method is implemented. In some embodiments, the multi-modal learning device may also be any intelligent terminal such as an in-vehicle computer, a portable PC, and a wearable device.

[0182] Please refer to Figure 15 , Figure 15 which schematically shows the hardware structure of a computer device in an embodiment. The computer device includes:

[0183] The processor 1501 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0184] The memory 1502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1502 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 1502 and are called by the processor 1501 to execute the model training method of multi-modal learning and / or the task processing method of multi-modal learning in the embodiments of the present application;

[0185] The input / output interface 1503 is used to implement information input and output;

[0186] The communication interface 1504 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);

[0187] The bus 1505 transmits information between the various components of the device (such as the processor 1501, the memory 1502, the input / output interface 1503, and the communication interface 1504);

[0188] Among them, the processor 1501, the memory 1502, the input / output interface 1503, and the communication interface 1504 achieve communication connections with each other inside the device through the bus 1505.

[0189] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned model training method of multi-modal learning and / or the task processing method of multi-modal learning.

[0190] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0191] The embodiments of the present application also provide a computer program product, including a computer program, and the steps implemented when the computer program is executed by a processor are substantially the same as the specific embodiments of the above-mentioned model training method for multi-modal learning and / or the task processing method for multi-modal learning, and will not be elaborated herein.

[0192] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0193] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0196] In the description of the present application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0197] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0198] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0199] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0202] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A model training method for multimodal learning, characterized in that: The method comprises: Obtaining training data of a first modality and training data of a second modality; Inputting the training data of the first modality into a pre-trained first modality model to obtain a first eigenvector; and inputting the training data of the second modality into a pre-trained second modality model to obtain a second eigenvector; Inputting the first feature vector and the second feature vector into a hybrid expert connector to be trained to train the hybrid expert connector to obtain a multimodal learning model; The multimodal learning model includes the first modal model, the second modal model and a trained hybrid expert connector, the trained hybrid expert connector is connected to the first modal model and the second modal model respectively, and the trained hybrid expert connector is used to align the first modality to the second modality and to align the second modality to the first modality.

2. The method according to claim 1, characterized in that The step of inputting the first feature vector and the second feature vector into a hybrid expert connector to be trained to train the hybrid expert connector includes: Inputting a target feature vector into a hybrid expert connector to be trained; wherein the target feature vector is any one of the first feature vector and the second feature vector; Performing a prediction task and a comparison task on the target feature vector based on the hybrid expert connector to align the first modality to the second modality and align the second modality to the first modality; Among them, when the target feature vector is the first feature vector, the prediction task is to generate a predicted feature vector under the second modality based on the first feature vector, and the comparison task is to compare the predicted feature vector under the second modality with the second feature vector; when the target feature vector is the second feature vector, the prediction task is to generate a predicted feature vector under the first modality based on the second feature vector, and the comparison task is to compare the predicted feature vector under the first modality with the first feature vector.

3. The method according to claim 2, characterized in that The performing a prediction task and a comparison task on the target feature vector based on the hybrid expert connector includes: At the tth time step, the predicted feature vector under the second modality is compared with the second feature vector based on the hybrid expert connector; wherein the hybrid expert connector generates the predicted feature vector under the second modality based on the first feature vector; At the t+1th time step, the predicted feature vector under the first modality is compared with the first feature vector based on the hybrid expert connector; wherein the hybrid expert connector generates the predicted feature vector under the first modality based on the second feature vector.

4. The method according to claim 2, characterized in that: The method further comprises: Calculating a multimodal alignment loss between the predicted feature vector under the second modality and the second feature vector; wherein the multimodal alignment loss includes a contrast loss and an L2 square difference loss; The hybrid expert connector is optimized based on the contrast loss and the L2 squared error loss to align the first modality to the second modality.

5. The method according to any one of claims 1 to 4, characterized in that The training data of the first modality includes first image data of an image modality, and the training data of the second modality includes first text data of a text modality, wherein the first image data is associated with the first text data; the first modality model includes a visual model, and the second modality model includes a language model; the first feature vector includes a first image feature vector obtained by learning the first image data with the visual model, and the second feature vector includes a first text feature vector obtained by learning the first text data with the language model; The step of inputting the first feature vector and the second feature vector into a hybrid expert connector to be trained to train the hybrid expert connector includes: Inputting the first image feature vector and the first text feature vector into a hybrid expert connector to be trained; The first image feature vector and the first text feature vector are processed respectively based on the hybrid expert connector to perform alignment from the image modality to the text modality and to perform alignment from the text modality to the image modality.

6. A multimodal learning task processing method, characterized in that: The method comprises: Acquire data to be processed for a multimodal task; wherein the data to be processed includes processed data of a first modality and processed data of a second modality; Inputting the processed data of the first modality and the processed data of the second modality into a preset multimodal learning model to obtain a processing result of the multimodal task; Among them, the first modal model in the multimodal learning model learns the processed data of the first modality to obtain a first feature vector; the second modal model in the multimodal learning model learns the processed data of the second modality to obtain a second feature vector; the hybrid expert connector in the multimodal learning model is used to align the first modality to the second modality and to align the second modality to the first modality.

7. The method according to claim 6, characterized in that The processed data of the first modality includes second image data of an image modality, the processed data of the second modality includes second text data of a text modality, and the second image data is associated with the second text data; the first modality model includes a pre-trained visual model, and the second modality model includes a pre-trained language model; the first feature vector includes a second image feature vector obtained by learning the second image data with the visual model, and the second feature vector includes a second text feature vector obtained by learning the second text data with the language model; The method further comprises: Inputting the second image feature vector and the second text feature vector into a hybrid expert connector in the multimodal learning model; generating a predicted feature vector of the second text feature vector based on the second image feature vector based on the hybrid expert connector, and comparing the predicted feature vector of the second text feature vector with the second text feature vector based on the hybrid expert connector to perform alignment from the image modality to the text modality; A predicted feature vector of the second image feature vector is generated based on the second text feature vector based on the hybrid expert connector, and the predicted feature vector of the second image feature vector is compared with the second image feature vector to align the text modality to the image modality.

8. A multimodal learning device, characterized in that: The multimodal learning device includes a memory and a processor, the memory stores a computer program, and the processor implements the model training method for multimodal learning described in any one of claims 1 to 5 when executing the computer program, and implements the task processing method for multimodal learning described in claim 6 or 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the model training method for multimodal learning described in any one of claims 1 to 5, and implements the task processing method for multimodal learning described in claim 6 or 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the steps of the model training method for multimodal learning described in any one of claims 1 to 5, and implements the task processing method for multimodal learning described in claim 6 or 7.