Multi-modal model training method, visual question and answer task processing method and equipment

Through cross-modal global and local alignment training, combined with dynamic expert mechanism, the problem that existing multimodal models are difficult to capture the correspondence between the problem and the image area in visual question-and-answer tasks is solved, and the accuracy of the answer is improved.

CN120164059APending Publication Date: 2025-06-17ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510235192.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When handling visual question-and-answer tasks, existing multimodal models are difficult to effectively capture the correspondence between natural language problems and specific areas in the image, resulting in poor accuracy of answers.

Method used

A multimodal model training method is proposed. Through cross-modal global and local alignment training, combined with dynamic expert mechanism, the coarse and fine granular alignment of images and problems is achieved, and the model's ability to capture the correspondence between text semantics and image areas is enhanced.

Benefits of technology

The answer accuracy of visual question-and-answer tasks is improved, so that the model can not only understand the image and the problem globally, but also carefully capture the correspondence between the semantics of the problem text and the image area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164059A_ABST
    Figure CN120164059A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal model training method and a visual question and answer task processing method and device, and belongs to the technical field of artificial intelligence. The training method comprises the steps that image training data and text training data are acquired; performing cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data to obtain a first hybrid expert connector, and performing cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data to obtain a second hybrid expert connector; obtaining a multi-modal model comprising a second hybrid expert connector; and the multi-modal model is used for performing global alignment and local alignment of the image modal information and the text modal information based on the second hybrid expert connector to obtain an answer of the visual question and answer task. According to the method, cross-modal alignment of coarse and fine granularities can be performed in combination with images and questions, so that the accuracy of answers of the visual question and answer tasks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method for a multimodal model, a processing method for a visual question answering task, and a device. Background Art

[0002] Visual Question Answering (VQA) is a multimodal task that combines computer vision, natural language processing, and cross-modal understanding, and requires a multimodal model to understand the input image and the relevant natural language question and generate a reasonable answer. In related technologies, in order to effectively align the information of the visual modality (images and videos) and the language modality (text questions) so as to accurately generate an answer related to the semantics of the question, some multimodal models perform cross-modal alignment and fusion of multimodal information based on feature fusion, and some multimodal models further introduce an additional object detection model on the basis of feature fusion to perform more comprehensive cross-modal information alignment. However, only based on feature fusion, the corresponding relationship between the natural language question and a specific region in the image cannot be captured, and introducing an additional object detection model will inevitably introduce additional noise, which results in poor accuracy of the answers generated by the multimodal model for processing the visual question answering task. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a training method for a multimodal model, a processing method for a visual question answering task, and a device, aiming to combine coarse-grained and fine-grained cross-modal alignment of images and questions, so that the multimodal model can globally understand images and questions and can also carefully capture the corresponding relationship between the semantic meaning of the question text and the image region, thereby improving the accuracy of the answers to the visual question answering task.

[0004] To achieve the above object, the first aspect of the embodiments of this application proposes a training method for a multimodal model, and the method includes:

[0005] Obtain image training data and text training data, where the text training data is the natural language question of the image training data;

[0006] Based on the image training data and the text training data, perform cross-modal global alignment training on a mixture of experts connector to obtain a first mixture of experts connector, where the first mixture of experts connector is used to globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data;

[0007] Based on the image training data and the text training data, perform cross-modal local alignment training on the first mixture of experts connector to obtain a multimodal model including a second mixture of experts connector;

[0008] Among them, the second hybrid expert connector is obtained by performing cross-modal local alignment training on the first hybrid expert connector, and the second hybrid expert connector is used to perform local alignment on the fine-grained image features of the image training data and the fine-grained text features of the text training data; the multi-modal model is used to perform global alignment and local alignment of the image modality information and the text modality information based on the second hybrid expert connector to obtain the answer to the visual question answering task.

[0009] In some embodiments, the cross-modal global alignment training of the hybrid expert connector based on the image training data and the text training data includes:

[0010] Obtaining the coarse-grained image features of the image training data; and obtaining the coarse-grained text features of the text training data;

[0011] Performing cross-modal global alignment training on the hybrid expert connector to be trained based on the coarse-grained image features and the coarse-grained text features.

[0012] In some embodiments, the cross-modal global alignment training of the hybrid expert connector to be trained based on the coarse-grained image features and the coarse-grained text features includes:

[0013] Inputting the target coarse-grained feature into the hybrid expert connector to be trained; wherein, the target coarse-grained feature is any one of the coarse-grained image features and the coarse-grained text features;

[0014] Performing a prediction task and a contrast task on the target coarse-grained feature based on the hybrid expert connector to perform cross-modal global alignment of the coarse-grained image features and the coarse-grained text features;

[0015] Wherein, when the target coarse-grained feature is the coarse-grained image feature, the prediction task is to generate a coarse-grained predicted text feature based on the coarse-grained image feature, and the contrast task is to compare the coarse-grained predicted text feature with the target coarse-grained text feature in the non-sequence dimension; the target coarse-grained text feature is obtained based on aggregating the coarse-grained text features; when the target coarse-grained feature is the coarse-grained text feature, the prediction task is to generate a coarse-grained predicted image feature based on the coarse-grained text feature, and the contrast task is to compare the coarse-grained predicted image feature with the target coarse-grained image feature in the non-sequence dimension; the target coarse-grained image feature is obtained based on aggregating the coarse-grained image features.

[0016] In some embodiments, the method further includes:

[0017] Calculate the cross-modal global alignment loss between the predicted text features and the target coarse-grained text features; wherein, the cross-modal global alignment loss at least includes a contrastive loss;

[0018] Optimize the hybrid expert connector based on the contrastive loss for cross-modal global alignment of the coarse-grained image features and the coarse-grained text features.

[0019] In some embodiments, the multimodal model further includes a visual model and a language model, and the second hybrid expert connector is respectively connected to the visual model and the language model;

[0020] The cross-modal local alignment training of the first hybrid expert connector based on the image training data and the text training data includes:

[0021] Input the target fine-grained features into the first hybrid expert connector; wherein, the target fine-grained features are either fine-grained image features or fine-grained text features; the fine-grained image features are obtained by the visual model learning the image training data, and the fine-grained text features are obtained by the language model learning the text training data;

[0022] Perform a similarity comparison process on the target fine-grained features based on the first hybrid expert connector, and perform a distribution constraint process on the global features corresponding to the target fine-grained features based on the first hybrid expert connector, so as to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

[0023] In some embodiments, the performing a similarity comparison process on the target fine-grained features based on the first hybrid expert connector, and performing a distribution constraint process on the global features corresponding to the target fine-grained features based on the first hybrid expert connector includes:

[0024] Based on the first hybrid expert connector, perform a similarity comparison between the image feature sequence corresponding to the fine-grained image features and the fine-grained text features to obtain sample similarity data; and use an L2 regularization loss function to constrain the distributions of the global features corresponding to the fine-grained image features and the fine-grained text features respectively; wherein, the sequence length of the image feature sequence is the same as the sequence length of the fine-grained text features; the sample similarity data includes positive sample similarity data and negative sample similarity data; the global features include the image global features corresponding to the fine-grained image features and the text global features corresponding to the fine-grained text features;

[0025] Calculate a local contrast loss based on the positive sample similarity data and the negative sample similarity data; and calculate a global alignment loss between the global image feature and the global text feature;

[0026] Optimize the trained hybrid expert connector based on the local contrast loss and the global alignment loss to perform cross-modal local alignment between the fine-grained image feature and the fine-grained text feature.

[0027] In some embodiments, the multi-modal model further includes a visual model and a language model, and the second hybrid expert connector is respectively connected to the visual model and the language model;

[0028] The cross-modal local alignment training of the first hybrid expert connector based on the image training data and the text training data includes:

[0029] Input a target fine-grained feature into the first hybrid expert connector; wherein, the target fine-grained feature is any one of a fine-grained image feature and a fine-grained text feature; the fine-grained image feature is obtained by the visual model learning the image training data, and the fine-grained text feature is obtained by the language model learning the text training data;

[0030] Perform a similarity comparison process on the target fine-grained feature based on the first hybrid expert connector;

[0031] Input the target fine-grained feature into the first hybrid expert connector again;

[0032] Perform a distribution constraint process on the global feature corresponding to the target fine-grained feature based on the first hybrid expert connector to perform cross-modal local alignment between the fine-grained image feature and the fine-grained text feature.

[0033] To achieve the above object, a second aspect of the embodiments of the present application proposes a method for processing a visual question answering task, the method including:

[0034] Obtain data to be processed for a visual question answering task, the data to be processed including image data and text data, and the text data being a natural language question about the image data;

[0035] Input the image data and the text data into a preset multi-modal model to obtain an answer to the visual question answering task;

[0036] Wherein, the hybrid expert connector in the multi-modal model performs global alignment and local alignment on the image modality information corresponding to the image data and the text modality information corresponding to the text data.

[0037] To achieve the above object, a third aspect of the embodiments of the present application provides a training device for a multimodal model, the device including:

[0038] A first acquisition module, configured to acquire image training data and text training data, where the text training data is a natural language question of the image training data;

[0039] A first-stage training module, configured to perform cross-modal global alignment training on a mixture-of-experts connector based on the image training data and the text training data to obtain a first mixture-of-experts connector, where the first mixture-of-experts connector is used to perform global alignment on the coarse-grained image features of the image training data and the coarse-grained text features of the text training data;

[0040] A second-stage training module, configured to perform cross-modal local alignment training on the first mixture-of-experts connector based on the image training data and the text training data to obtain a multimodal model including a second mixture-of-experts connector;

[0041] Wherein, the second mixture-of-experts connector is obtained by performing local alignment training on the first mixture-of-experts connector, and the second mixture-of-experts connector is used to perform local alignment on the fine-grained image features of the image training data and the fine-grained text features of the text training data; the multimodal model is used to perform global alignment and local alignment on the image modality information and the text modality information based on the second mixture-of-experts connector to obtain an answer to a visual question answering task.

[0042] To achieve the above object, a fourth aspect of the embodiments of the present application provides a processing device for a visual question answering task, the device including:

[0043] A second acquisition module, configured to acquire data to be processed for a visual question answering task, where the data to be processed includes image data and text data, and the text data is a natural language question of the image data;

[0044] A task processing module, configured to input the image data and the text data into a preset multimodal model to obtain an answer to the visual question answering task;

[0045] Wherein, the mixture-of-experts connector in the multimodal model performs global alignment and local alignment on the image modality information corresponding to the image data and the text modality information corresponding to the text data.

[0046] To achieve the above object, a fifth aspect of the embodiments of the present application provides a processing device for visual question answering tasks. The processing device for visual question answering tasks includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method described in the first aspect above, and / or implements the method described in the second aspect above.

[0047] To achieve the above object, a sixth aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect above, and / or implements the method described in the second aspect above.

[0048] To achieve the above object, a seventh aspect of the embodiments of the present application provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method provided in the first aspect above, and / or implements the method provided in the second aspect above.

[0049] The training method of the multi-modal model, the processing method of the visual question answering task, the device, the equipment, the computer-readable storage medium and the computer program product provided by the embodiments of the present application obtain image training data and text training data (natural language questions of the image training data), and use the image training data and the text training data as an image-text pair to input into the model to be trained, so as to perform cross-modal global alignment training on the mixture-of-experts connector based on the image training data and the text training data, and obtain a first mixture-of-experts connector. Among them, the first mixture-of-experts connector is used to globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data. Then, further perform cross-modal local alignment training on the first mixture-of-experts connector based on the image training data and the text training data, so as to obtain a multi-modal model including a second mixture-of-experts connector. In the multi-modal model, the second mixture-of-experts connector is obtained by performing cross-modal local alignment training on the first mixture-of-experts connector, and the second mixture-of-experts connector is used to locally align the fine-grained image features of the image training data and the fine-grained text features of the text training data; and, the multi-modal model is used to perform global alignment and local alignment of the image modality information and the text modality information based on the second mixture-of-experts connector to obtain the answer to the visual question answering task.

[0050] Compared with the method of aligning visual modal information and text model information based on simple feature fusion and introducing an additional object detection model, the embodiment of the present application performs global (coarse-grained) and local (fine-grained) cross-modal alignment on the training data of the image modality and the training of the text modality through a mixture-of-experts connector. As a result, when the trained multi-modal model processes visual question answering tasks, it can not only globally understand images and text, but also carefully capture the correspondence between text semantics and image regions, thereby improving the accuracy of the answers to visual question answering tasks.

[0051] In addition, the multi-modal model trained by introducing a mixture-of-experts connector in the embodiment of the present application can also, when processing complex scenarios such as visual question answering tasks and image caption generation tasks, select specific experts according to specific questions through a dynamic expert mechanism, and implement dynamic adjustment of the expert module for cross-modal hierarchical semantic alignment (global semantic alignment and local semantic alignment), so as to be able to process multi-modal tasks in complex scenarios more efficiently, accurately and stably. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic flowchart of the steps in some embodiments of the training method of the multi-modal model provided by the embodiment of the present application;

[0053] Figure 2 It is a schematic structural diagram of the dynamic mixture-of-experts model MoE involved in some embodiments of the training method of the multi-modal model provided by the embodiment of the present application;

[0054] Figure 3 It is a schematic flowchart of the process of generating answers to visual question answering tasks involved in some embodiments of the training method of the multi-modal model provided by the embodiment of the present application;

[0055] Figure 4 is Figure 1 a schematic flowchart of the refined steps of step S102 in;

[0056] Figure 5 It is a schematic logical flowchart of cross-modal global alignment training involved in some embodiments of the training method of the multi-modal model provided by the embodiment of the present application;

[0057] Figure 6 is Figure 1 a schematic flowchart of a kind of refined steps of step S103 in;

[0058] Figure 7 It is a schematic logical flowchart of cross-modal local alignment training involved in some embodiments of the training method of the multi-modal model provided by the embodiment of the present application;

[0059] Figure 8 is Figure 1Another refined step flow diagram of step S103;

[0060] Figure 9 It is a schematic logical flow diagram of cross-modal alignment training based on task embedding involved in the training method of the multi-modal model provided by the embodiments of the present application in some embodiments;

[0061] Figure 10 It is a schematic step flow diagram of the alternating alignment algorithm involved in the training method of the multi-modal model provided by the embodiments of the present application in some embodiments;

[0062] Figure 11 It is a schematic step flow diagram of the global and local alignment algorithms involved in the training method of the multi-modal model provided by the embodiments of the present application in some embodiments;

[0063] Figure 12 It is a schematic diagram of the experimental results of the multi-modal model involved in the training method of the multi-modal model provided by the embodiments of the present application in some embodiments;

[0064] Figure 13 It is a schematic step flow diagram of the processing method of the visual question answering task provided by the embodiments of the present application in some embodiments;

[0065] Figure 14 It is a schematic step flow diagram of the local attention algorithm involved in the training method of the multi-modal model provided by the embodiments of the present application in some embodiments;

[0066] Figure 15 It is a schematic structural diagram of the training device of the multi-modal model provided by the embodiments of the present application;

[0067] Figure 16 It is a schematic structural diagram of the processing device of the visual question answering task provided by the embodiments of the present application

[0068] Figure 17 It is a schematic hardware structure diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners

[0069] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0070] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0072] First, the overall concept of the embodiments of this application will be described.

[0073] In the field of artificial intelligence technology, visual question answering provides innovative solutions for intelligent transportation and vehicle applications. Combining multi-modal perception and reasoning capabilities, visual question answering has broad application prospects in intelligent driving, in-vehicle interaction, and vehicle-mounted systems. As a multi-modal task that combines computer vision, natural language processing, and cross-modal understanding, visual question answering has open-ended question answering types and multiple-choice question answering types. Among them, open-ended question answering is also called generation-based visual question answering. When processing this type of visual question answering task, the answer can be any natural language phrase; while multiple-choice question answering is also called retrieval-based visual question answering, and its answer is obtained by selecting from a predefined candidate set.

[0074] The main difficulties in processing visual question answering tasks lie in visual information processing, language understanding, and multi-modal alignment. Among them, visual information processing refers to extracting meaningful features from images. These features usually require global features to contain high-level semantics of the images because for visual-language cross-modal understanding, global alignment can measure the general consistency between an image and a text description. In addition, language understanding refers to parsing the semantic information contained in the question, and these semantic information include word-level semantics and sentence-level understanding. Multi-modal alignment refers to aligning language features and visual features to achieve information fusion and reasoning, which usually involves understanding from global to local, and thus requires achieving an alignment from coarse-grained to fine-grained, such as from sentence-image alignment to words-patches alignment.

[0075] In related technologies, in order to effectively align the information of the visual modality and the language modality, and the alignment is from coarse-grained to fine-grained, there are basically some problems in existing methods. This is because traditional multimodal alignment usually aligns semantics only globally, while visual question answering requires not only global alignment but also local alignment. Exemplarily, some early methods based on feature fusion (such as MMTM: Multimodal Transfer Module for CNN Fusion) can extract image features by using a convolutional neural network CNN and process text features by using a recurrent neural network RNN (specifically, a long short-term memory-based recurrent neural network LSTM), and then achieve multimodal fusion through simple feature concatenation or linear projection. Using this method, the correspondence between the question and a specific image region cannot be captured, only global correspondence can be modeled, and the method lacks the ability to model fine-grained semantics in the question. In addition, there are also some methods that introduce region features (such as LXMERT: Learning Cross-Modality Encoder Representations from Transformers), which extract region features in the image by introducing an object detection model such as Faster-RCNN and perform dynamic alignment in combination with the question semantics. However, this method requires introducing an additional model, and the bounding box is usually a rectangular box, so the predicted bounding box may contain noise. Furthermore, there are also some methods based on multimodal pre-training for processing visual question answering tasks and have shown significant advantages, such as BEIT3: Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks. This method enables the model to learn rich cross-modal representations and establish deeper semantic associations between images and texts by pre-training on a large-scale image-text pair dataset. However, this method usually relies on a large-scale, high-quality image-text dataset, and this type of method usually requires training a cross-modal alignment model from scratch and is difficult to directly integrate the knowledge in existing large-scale unimodal pre-trained models (such as large-scale visual models and language models), resulting in insufficient knowledge reuse. Therefore, the model training cost of this type of method is high, and a large amount of computing resources and time are required for large-scale multimodal pre-training.

[0076] Based on this, embodiments of the present application provide a training method for a multimodal model, a processing method for a visual question answering task, a device, a device, a computer-readable storage medium, and a computer program product, aiming to combine cross-modal alignment of coarse and fine grains for images and questions, so that the multimodal model can not only globally understand images and questions but also carefully capture the correspondence between the question text semantics and the image region, thereby improving the accuracy of the answers to visual question answering tasks.

[0077] Embodiments of this application utilize the coarse-grained to fine-grained information in large vision models (LVMs) and large language models (LLMs) to facilitate visual understanding tasks. By performing an overall alignment from coarse-grained to fine-grained across modalities, the semantic gap between coarse and fine granularities is significantly addressed, achieving more complete semantic alignment.

[0078] It should be noted that the coarse-to-fine alignment ensures semantic consistency and complementarity of features in different modalities during the fusion process, and it is a top-down, coarse-to-fine alignment. Among them, based on the coarse-grained alignment, it is ensured that the model can understand the overall semantics of the image and the question. For example, for the question "What is the theme of this picture?", the global information of the image is captured through coarse-grained alignment: this is a beach scene. In addition, based on the fine-grained alignment, it helps the model identify specific regions or objects and align them with the details in the question. For example, for the question "What color are the shoes worn by the person in the image?", the fine-grained alignment helps the model focus on the target area (the person's shoes) and extract accurate color information. That is to say, the embodiments of this application can support the finally trained model to handle complex tasks by adopting this coarse-to-fine alignment. For example, it can capture the overall context through coarse-grained alignment to handle factual questions: "What season is shown in this picture?" It can also identify specific objects based on fine-grained alignment to handle detail questions: "What species is the bird in the upper left corner of the picture?", and it can also help the model gradually build context relationships through the combination of coarse and fine granularity alignments to handle multi-step reasoning questions: "What brand is the car on the far right in the picture?" The model first performs local localization (fine-grained alignment) and then combines the global context (coarse-grained alignment) to answer the question.

[0079] Embodiments of this application effectively guide and model the interoperability and independence between coarse-grained information flow and fine-grained information flow by introducing a dynamic expert mechanism and jointly optimizing global and local contrast learning and L2 loss, achieving hierarchical alignment from coarse-grained to fine-grained, and being able to dynamically adjust the feature expressions and alignment strategies of coarse-grained information flow and fine-grained information flow, so that the model can maintain the independence of both when processing global semantics and local details, and can also achieve efficient information interoperability through a shared mechanism.

[0080] Next, the training method of the multi-modal model, the processing method of the visual question answering task, the device, the equipment, the computer-readable storage medium, and the computer program product provided by the embodiments of this application are specifically described through the following embodiments, and first, the training method of the multi-modal model and the processing method of the visual question answering task provided by the embodiments of this application are described in detail.

[0081] It should be noted that the training method of the multimodal model and the processing method of the visual question answering task provided in the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be an in-vehicle terminal on a vehicle, or a computer device such as a smart phone, a tablet computer, a laptop computer, or a desktop computer. The server can be the background server terminal device of the aforementioned terminal, which can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The software can be an application, a computer program, and a storage medium carrying the computer program that implement the training method of the multimodal model and / or the processing method of the visual question answering task. It should be understood that, based on different design requirements of actual applications, in different feasible embodiments, the terminal, server, and software that apply the training method of the multimodal model and the processing method of the visual question answering task provided in the embodiments of the present application can, of course, also be other forms not listed here. The training method of the multimodal model and the processing method of the visual question answering task provided in the embodiments of the present application do not specifically limit this.

[0082] In addition, the present application can also be used in many general or special computer system environments or configurations. For example: vehicles, personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0083] For ease of understanding and description, in the following text, the training method of the multi-modal model and the processing method of the visual question answering task provided by the embodiments of the present application are taken as examples for the terminal device to illustrate each specific embodiment of the present application in detail. For the implementation of any of the above-mentioned forms of the subject applying the embodiments of the present application, the operation process of the terminal device in each specific embodiment described later can be referred to.

[0084] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing that needs to be carried out according to data related to the user's identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0085] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the step flow of the training method of the multi-modal model provided by the embodiments of the present application in some embodiments. It should be understood that although Figure 1 and the subsequent other schematic diagrams of the step flow show the execution order of some method steps, due to different design requirements in actual applications, the training method of the multi-modal model provided by the embodiments of the present application can of course adopt an execution order different from that shown in the figure. That is, Figure 1 the order of the shown method steps does not constitute a limitation on the execution logic order of the training method of the multi-modal model provided by the embodiments of the present application, and any reasonable changes based on Figure 1 the order of the shown method steps should be included in the protection scope of the training method of the multi-modal model provided by the embodiments of the present application.

[0086] As Figure 1 shown, in some embodiments, the terminal device applying the training method of the multi-modal model provided by the embodiments of the present application may include the following steps S101 to S103.

[0087] Step S101: Obtain image training data and text training data, where the text training data is the natural language question of the image training data.

[0088] The terminal device constructs an image-text pair by obtaining image training data and obtaining the natural language question of the image training data as text training data, so as to perform subsequent operations of training the multi-modal model with the image-text pair as the input.

[0089] It should be noted that during the model training process of the terminal device, image-to-text and text-to-image alignment training are alternately performed based on the mixture-of-experts mechanism. That is to say, the terminal device realizes semantic alignment between modalities of image-to-text (I2T) and text-to-image (T2I) through an alternating training mechanism.

[0090] In some embodiments, the terminal device uses a dynamic mixture-of-experts module (MoE) as a connection hub for cross-modal alignment, mapping image features to the text space or mapping text features to the image space respectively to achieve the intercommunication of fine-grained and coarse-grained semantics. Among them, the MoE module can optimize some expert parameters according to different tasks (I2T and T2I), so as to dynamically adapt to task requirements, and can improve the computational efficiency and feature expression ability of the model at the same time.

[0091] It should be noted that as Figure 2 shown, the MoE module consists of multiple expert networks (Experts) and a gating network (Router). MoE is a dynamic network structure designed to improve the computational efficiency and hierarchical representation ability of deep learning models by partitioning model parameters and tasks. The core idea of MoE is to organize multiple Experts modules together, and a Router module selectively activates some Experts through weights to achieve computing on demand. Based on the dynamic network structure of MoE, the terminal device can design exclusive expert modules for different modalities (such as image modality and text modality) to achieve targeted modeling.

[0092] In some embodiments, the process of the terminal device performing alternating training includes an I2T stage and a T2I stage. Among them, the I2T stage refers to extracting coarse-grained global features from the image modality and converting them to the text semantic space through the MoE module for alignment with text features; the T2I stage is the opposite, which refers to extracting coarse-grained global features from the text modality and converting them to the image semantic space through the MoE module for alignment with image features. The terminal device jointly optimizes the dynamic routing and expert parameters of the MoE module by alternately updating the alignment targets of the I2T stage and the T2I stage to perform image-to-text and text-to-image alignment training.

[0093] Step S102: Perform cross-modal global alignment training on the mixture-of-experts connector based on the image training data and the text training data to obtain a first mixture-of-experts connector, where the first mixture-of-experts connector is used to globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data.

[0094] It should be noted that the hybrid expert connector can be designed based on the above-mentioned MoE module. Therefore, the structure of the hybrid expert connector can be basically the same as the network structure of the above-mentioned MoE module.

[0095] When the terminal device uses the obtained image training data and text training data as an image-text pair input for model training, it first performs cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data, so as to obtain the trained first hybrid expert connector. Among them, based on the cross-modal global alignment training performed by the terminal device, the first hybrid expert connector can globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data.

[0096] In some embodiments, the terminal device can use the first hybrid expert connector as the initial model after cross-modal global alignment training. In other embodiments, the terminal device can also use the first hybrid expert connector as a component to form an initial model that has undergone cross-modal global alignment training together with other components such as an encoder component, a denoiser component, and a decoder component.

[0097] Step S103: Perform cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data to obtain a multi-modal model including a second hybrid expert connector; wherein, the second hybrid expert connector is obtained by performing cross-modal local alignment training on the first hybrid expert connector, and the second hybrid expert connector is used to locally align the fine-grained image features of the image training data and the fine-grained text features of the text training data; the multi-modal model is used to globally and locally align the image modal information and the text modal information based on the second hybrid expert connector to obtain the answer to the visual question answering task.

[0098] After the terminal device performs cross-modal global alignment training on the hybrid expert connector based on image training data and text training data to obtain the first hybrid expert connector, it further performs cross-modal local alignment training on the first hybrid expert connector based on image training data and text training data to obtain the second hybrid expert connector. In this way, the terminal device can directly use the second hybrid expert connector as the finally trained multi-modal model; alternatively, the terminal device can also use the first hybrid expert connector together with other components such as encoder components, denoiser components, and decoder components to form a multi-modal model that has undergone both cross-modal global alignment training and cross-modal local alignment training. Among them, based on the cross-modal local alignment training performed by the terminal device, the second hybrid expert connector can perform local alignment on the fine-grained image features of the image training data and the fine-grained text features of the text training data; in addition, the terminal device can use the finally trained multi-modal model to process visual question answering tasks, and when the multi-modal model processes visual question answering tasks, it can perform global alignment and local alignment of image modality information and text modality information based on the second hybrid expert connector to obtain the answer to the visual question answering task.

[0099] In some embodiments, after the terminal device performs two-stage alignment training (cross-modal global alignment training in the first stage and cross-modal local alignment training in the second stage) on the hybrid expert connector designed based on the MoE module, it can use the finally trained multi-modal model including the second hybrid expert connector to perform the visual question answering VQA task.

[0100] Exemplarily, as Figure 3 shown, when the terminal device uses the multi-modal module to perform the VQA task, in the first stage, the second hybrid expert connector selects some experts to train and align the global parts of the visual model and the language model, and then in the second stage, it selects some experts to align the local information. Finally, the visual-linguistic representation obtained after global alignment and local alignment by the second hybrid expert connector is input into the classification layer to predict the answer to the VQA task. Among them, the input of the VQA task includes an image and a question about this image, and the answer to the VQA task generated by the multi-modal model is an accurate answer to this question. In an actual scenario, for the visual question answering VQA task, the terminal device will receive a natural image and a related question, and the task is to generate or select the correct answer based on the multi-modal model.

[0101] In some embodiments, after the terminal device performs two - stage alignment training on the hybrid expert connector designed based on the MoE module, it can first train and evaluate the multimodal model on the VQA 2.0 dataset, and then use the multimodal model to perform a specific visual question - answering (VQA) task. Among them, according to common practice, when the terminal device trains and evaluates the multimodal model on the VQA 2.0 dataset, it can convert VQA 2.0 into a classification task and select answers from a shared set containing 3129 answers.

[0102] In some embodiments, during the process that the terminal device trains the hybrid expert connector in two stages based on image training data and text training data to obtain the second hybrid expert connector, the terminal device can control the coarse - grained and fine - grained information flows in the image training data and text training data through a dynamic expert mechanism. For example, the terminal device designs a dynamic routing mechanism through the MoE module. In the first - stage training, for the coarse - grained information flow, only the adapted expert modules in the hybrid expert connector are activated to process the global features, while in the second - stage training, for the fine - grained information flow, other expert modules in the hybrid expert connector are activated to process the local features. Moreover, the terminal device also realizes the inter - communication and cooperation between the coarse and fine information flows through shared experts. Among them, the terminal device can introduce shared expert modules through a hierarchical information - sharing mechanism to perform dynamic transfer of complementary information between the coarse - grained and fine - grained information flows, so as to not only maintain the independence of the two types of information flows of coarse and fine, but also fully explore the cooperation potential between the two types of information flows, thereby improving the alignment effect.

[0103] In some embodiments, the terminal device can perform refined control on the coarse - grained and fine - grained information flows by combining the regularization effect of the L2 loss function, enhancing the stability of the coarse - grained and fine - grained alignment while avoiding feature redundancy or overfitting, thereby ensuring the effectiveness of the dynamic expert mechanism.

[0104] In some embodiments, the terminal device may adopt a hierarchical alignment method from coarse to fine to train the mixture-of-experts connector in two stages. For example, the terminal device divides the input image training data and text training data into coarse-grained information flows (such as scenes, topics, etc.) and fine-grained information flows (such as target objects, attributes, etc.) respectively. Then, when performing cross-modal global alignment training on the mixture-of-experts connector, the alignment process of the coarse-grained information flow is optimized through global contrast learning, and when performing cross-modal local alignment training on the first mixture-of-experts connector, the alignment process of the fine-grained information flow is optimized through local contrast learning. In this way, the consistency and accuracy of semantic information from the whole to the local can be ensured. Moreover, by using the training method of alignment from coarse to fine and combining the control framework of the MoE module, the terminal device can enable the finally trained multi-modal model including the second mixture-of-experts controller to have a more refined semantic alignment ability, so as to accurately handle multi-granularity problems when processing multi-modal tasks such as visual question answering and image caption generation, thereby greatly improving the generalization ability and task performance of the model.

[0105] In the embodiments of the present application, the terminal device obtains image training data and text training data (natural language questions of the image training data), and first performs cross-modal global alignment training on the mixture-of-experts connector based on the image training data and the text training data, so as to obtain the trained first mixture-of-experts connector, enabling the first mixture-of-experts connector to globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data based on the cross-modal global alignment training performed by the terminal device. Then, the terminal device further performs cross-modal local alignment training on the first mixture-of-experts connector based on the image training data and the text training data, so as to obtain a multi-modal module including the second mixture-of-experts connector. Among them, the second mixture-of-experts connector can locally align the fine-grained image features of the image training data and the fine-grained text features of the text training data based on the cross-modal local alignment training performed by the terminal device. In addition, the terminal device can use the finally trained multi-modal model to process visual question answering tasks, and when the multi-modal model processes visual question answering tasks, it can globally and locally align the image modality information and the text modality information based on the second mixture-of-experts connector, so as to obtain the answer to the visual question answering task.

[0106] In this way, the embodiments of the present application perform global (coarse-grained) and local (fine-grained) cross-modal alignment on the training data of the image modality and the training of the text modality through the mixture-of-experts connector, so that the trained multi-modal model can not only globally understand the image and the text, but also carefully capture the correspondence between the text semantics and the image regions when processing visual question answering tasks, thereby improving the accuracy of the answers to the visual question answering tasks.

[0107] In addition, the multi-modal model obtained by training with the hybrid expert connector in the embodiments of the present application can also, when processing complex scenarios such as visual question answering tasks and image caption generation tasks, select specific experts according to specific questions through a dynamic expert mechanism, and achieve dynamic adjustment of the expert module for cross-modal hierarchical semantic alignment (global semantic alignment and local semantic alignment), so as to be able to process multi-modal tasks in complex scenarios more efficiently, accurately and stably.

[0108] When the terminal device performs the first-stage global alignment stage based on the image training data and the text training data, the goal is to model and align the coarse-grained information flows in the image training data and the text training data, so as to ensure the global semantic consistency between the input modalities (image modality and text modality), and further lay a foundation for the subsequent second-stage fine-grained alignment.

[0109] Please refer to Figure 4 , Figure 4 For Figure 1 the detailed step flow diagram of step S102 in

[0110] In some embodiments, as Figure 4 shown, the step of "performing cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data" in the above step S102 may include step S401 and step S402 as shown below.

[0111] Step S401: Obtain the coarse-grained image features of the image training data; and, obtain the coarse-grained text features of the text training data.

[0112] When the terminal device performs training in the global alignment stage based on the image training data and the text training data, it first performs extraction processing on the coarse-grained information flow, so as to obtain the coarse-grained image features of the image training data and the coarse-grained text features of the text training data.

[0113] In some embodiments, when the terminal device performs extraction processing on the coarse-grained information flow of the image training data and the text training data, it performs coarse-grained feature extraction on the input image-text pair (image training data and text training data) to obtain the global semantic information I g . Among them, the global semantic information I g includes the coarse-grained image features of the image modality and the coarse-grained text features of the text modality.

[0114] Exemplarily, to obtain the coarse-grained image features of the image modality, the terminal device may use the pre-trained large vision model Dino-v2 large to extract the global features of the image training data, so as to obtain the coarse-grained image features by capturing the overall semantics such as the scene and the theme. In addition, to obtain the coarse-grained text features of the text modality, the terminal device may use the pre-trained language model Qwen2 to extract the global features of the text training data, so as to obtain the coarse-grained text features by capturing the sentence-level or paragraph-level semantics.

[0115] Step S402: Perform cross-modal global alignment training on the hybrid expert connector to be trained based on the coarse-grained image features and the coarse-grained text features.

[0116] After the terminal device obtains the coarse-grained image features and the coarse-grained text features, it further performs cross-modal global alignment training on the hybrid expert connector designed based on the MoE module and not yet trained based on the coarse-grained image features and the coarse-grained text features, so as to obtain the trained first hybrid expert connector.

[0117] In some embodiments, based on the role of the MoE module in learning the global correlation relationship between modalities and further optimizing the coarse-grained features, the terminal device may, based on the dynamic expert mechanism of the MoE module, have the hybrid expert connector dynamically activate the adapted expert module for feature processing according to the distribution characteristics of the coarse-grained information flow (coarse-grained image features and coarse-grained text features) (such as scene complexity, text length). In addition, the terminal device also establishes a shared coarse-grained feature channel between the image modality and the text modality through the hybrid expert connector, and fuses cross-modal information through the attention mechanism or feature splicing operation to extract the global shared semantics.

[0118] In some embodiments, the above step S402 may include the following steps:

[0119] Input the target coarse-grained feature into the hybrid expert connector to be trained; wherein, the target coarse-grained feature is any one of the coarse-grained image features and the coarse-grained text features;

[0120] Based on the hybrid expert connector, perform a prediction task and a contrast task on the target coarse-grained feature to perform cross-modal global alignment of the coarse-grained image features and the coarse-grained text features.

[0121] It should be noted that since the terminal device alternately performs alignment training for image-to-text and text-to-image based on the mixture-of-experts mechanism, when the terminal device performs alignment training for image-to-text, the target coarse-grained feature is the coarse-grained image feature. In this case, the prediction task is to generate a coarse-grained predicted text feature based on the coarse-grained image feature, and the contrast task is to compare the coarse-grained predicted text feature with the target coarse-grained text feature in the non-sequence dimension. Among them, the target coarse-grained text feature is obtained based on the aggregated coarse-grained text feature. In addition, when the terminal device performs alignment training for text-to-image, the target coarse-grained feature is the coarse-grained text feature. In this case, the prediction task is to generate a coarse-grained predicted image feature based on the coarse-grained text feature, and the contrast task is to compare the coarse-grained predicted image feature with the target coarse-grained image feature in the non-sequence dimension. Among them, the target coarse-grained image feature is obtained based on the aggregated coarse-grained image feature.

[0122] Exemplarily, as Figure 5 shown, both the coarse-grained image feature and the coarse-grained text feature extracted by the terminal device have a sequence dimension (i.e., the dimension is: sequence, embedding dimension). In order to obtain the global feature I of the image g and obtain the global feature T of the text g , the terminal device aggregates the coarse-grained image feature and the coarse-grained text feature along the sequence respectively to obtain the target coarse-grained image feature I without the sequence dimension g and the target coarse-grained text feature T g , while retaining the original coarse-grained image feature I and coarse-grained text feature T.

[0123] After that, in order for the terminal device to map the extracted global feature I g and the global feature T g to the same semantic space, cross-modal global alignment training for image-to-text and text-to-image is alternately performed on the mixture-of-experts connector designed based on the dynamic mixture-of-experts MoE module. First, in the image-to-text stage, the terminal device inputs the coarse-grained image feature I with the sequence dimension into the mixture-of-experts connector to select a part of the experts for aligning the global features in the image-to-text stage. Then, in the subsequent text-to-image stage, the terminal device inputs the coarse-grained text feature T with the sequence dimension into the mixture-of-experts connector, and the mixture-of-experts connector selects a part of the experts for aligning the global features in the text-to-image stage. Among them, in the image-to-text stage, after the coarse-grained image feature I is processed by selecting a part of the experts through the mixture-of-experts connector, the global feature I with the sequence dimension is output i2t (coarse-grained predicted text feature), and then the contrast loss between the predicted text feature I i2t and the target coarse-grained text feature T g is calculated through global contrast learning.

[0124] L i2t-cl = CL(I i2t , T g ).

[0125] The text-to-image stage is the opposite of the image-to-text stage. After the coarse-grained text feature T is processed by selecting a part of the experts through the mixture-of-experts connector, the global feature T with a sequence dimension is output t2i (coarse-grained predicted image feature), and then the predicted text feature T is calculated through global contrast learning t2i and the target coarse-grained image feature I g to calculate the contrast loss therebetween.

[0126] In some embodiments, the training method of the multi-modal model provided by the embodiments of the present application may further include the following steps:

[0127] Calculate the cross-modal global alignment loss between the predicted text feature and the target coarse-grained text feature; wherein, the cross-modal global alignment loss at least includes a contrast loss;

[0128] Optimize the mixture-of-experts connector based on the contrast loss for cross-modal global alignment of the coarse-grained image feature and the coarse-grained text feature.

[0129] During the process of the terminal device obtaining the first mixture-of-experts connector through cross-modal global alignment training of the mixture-of-experts connector, during the image-to-text alignment training stage, the contrast loss between the predicted text feature and the target coarse-grained text feature can also be calculated as the cross-modal global alignment loss, and then the mixture-of-experts connector is optimized based on the contrast loss for cross-modal global alignment of the coarse-grained image feature and the coarse-grained text feature. And, during the text-to-image alignment training stage, the contrast loss between the predicted image feature and the target coarse-grained image feature is calculated, so as to optimize the mixture-of-experts connector based on the contrast loss for cross-modal global alignment of the coarse-grained text feature and the coarse-grained image feature.

[0130] In some embodiments, the terminal device can perform cross-modal global alignment training on the mixture-of-experts connector by controlling the mixture-of-experts connector to perform global contrastive learning on the coarse-grained image features and coarse-grained text features. For example, the terminal device inputs the coarse-grained features (coarse-grained image features and coarse-grained text features) of a pair of corresponding image training data and text training data as positive samples into the mixture-of-experts connector to be trained, and inputs the coarse-grained features of some unrelated image training data and text training data as negative samples into the mixture-of-experts connector. Moreover, when the terminal device performs cross-modal global alignment training on the mixture-of-experts connector based on the coarse-grained image features and coarse-grained text features of the positive samples and the coarse-grained image features and coarse-grained text features of the negative samples, it controls the mixture-of-experts connector to use the contrastive learning objective InfoNCE loss to optimize the coarse-grained feature alignment (image-to-text alignment and text-to-image alignment). Among them, using the InfoNCE loss as the contrastive learning objective is to maximize the semantic similarity of the positive sample pairs and minimize the similarity of the negative sample pairs. The formula of the InfoNCE loss is as follows:

[0131]

[0132] where f i and g i are the coarse-grained features of the image modality and the text modality respectively (f i is the coarse-grained image feature, and g i is the coarse-grained text feature), sim(*) represents the similarity function, and τ is the temperature parameter.

[0133] In this embodiment, the terminal device extracts the coarse-grained information flow of the image training data and the text training data, thereby obtaining the coarse-grained image features of the image training data and the coarse-grained text features of the text training data, and alternately performs cross-modal global alignment training on the mixture-of-experts connector designed based on the MoE module based on the coarse-grained image features and coarse-grained text features, so as to train the mixture-of-experts connector to effectively complete the semantic modeling of the coarse-grained information and the inter-modal consistency alignment in the global alignment stage.

[0134] Please refer to Figure 6 , Figure 6 which is Figure 1 a schematic diagram of a refined step flow of step S103 in

[0135] In some embodiments, the multi-modal model finally trained by the terminal device further includes a pre-trained visual model and a language model in addition to the second mixture-of-experts connector, and the second mixture-of-experts connector is respectively connected to the visual model and the language model. In this case, as Figure 6As shown, in step S103 above, "performing cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data" may include step S601 and step S602 shown below.

[0136] Step S601: Input the target fine-grained feature into the first hybrid expert connector; wherein, the target fine-grained feature is either a fine-grained image feature or a fine-grained text feature; the fine-grained image feature is obtained by the visual model learning from the image training data, and the fine-grained text feature is obtained by the language model learning from the text training data.

[0137] After the terminal device performs cross-modal global alignment training on the hybrid expert connector to obtain the first hybrid expert connector, it further inputs the fine-grained image feature obtained by the visual model learning and encoding the image training data, and the fine-grained text feature obtained by the language model learning and encoding the text training data into the first hybrid expert connector to perform subsequent cross-modal local alignment training on the first hybrid expert connector.

[0138] In some embodiments, when the terminal device performs cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data, it inputs the image training data into the visual model DINO to extract the image feature I (fine-grained image feature) based on the visual model as the image encoder; and inputs the text training data into the language model to extract the text feature T (fine-grained text feature) based on the language model as the text encoder. Among them, the feature dimension of the initial feature I is (L I ×D I ), L I is the sequence length of the image blocks, D I is the feature dimension of each block. In addition, the feature dimension of the text feature T is (L T ×D T ), L T is the text sequence length, and D T is the feature dimension of each word token.

[0139] In some embodiments, since the hybrid expert connector is respectively connected to the visual model and the language model, the fine-grained image feature extracted by the terminal device based on the visual model and the fine-grained text feature extracted based on the language model will be automatically input into the first hybrid expert connector for the first hybrid expert connector to perform subsequent cross-modal local alignment training.

[0140] In some embodiments, in the image-to-text stage, the terminal device may input image training data into a vision model, and based on the vision model, extract fine-grained image features and input them into a first mixture-of-experts connector. Then, in the text-to-image stage, the terminal device inputs text training data into a language model, and based on the language model, extracts fine-grained text features and inputs them into the first mixture-of-experts connector.

[0141] Step S602: Perform a similarity comparison process on the target fine-grained features based on the first mixture-of-experts connector, and perform a distribution constraint process on the global features corresponding to the target fine-grained features based on the first mixture-of-experts connector, so as to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

[0142] After the terminal device inputs the target fine-grained features (fine-grained image features or fine-grained text features) into the first mixture-of-experts connector, it controls the first mixture-of-experts connector to perform a similarity comparison process on the input target fine-grained features, and perform a distribution constraint process on the global features corresponding to the target fine-grained features, so as to perform cross-modal local alignment training of image-to-text and text-to-image alternately, and perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

[0143] In some embodiments, the above step S602 may include the following steps:

[0144] Based on the first mixture-of-experts connector, perform a similarity comparison between the image feature sequence corresponding to the fine-grained image features and the fine-grained text features to obtain sample similarity data; and use the L2 regularization loss function to constrain the distributions of the global features corresponding to the fine-grained image features and the fine-grained text features respectively; wherein, the sequence length of the image feature sequence is the same as the sequence length of the fine-grained text features; the sample similarity data includes positive sample similarity data and negative sample similarity data; the global features include the image global features corresponding to the fine-grained image features and the text global features corresponding to the fine-grained text features;

[0145] Calculate a local contrast loss based on the positive sample similarity data and the negative sample similarity data; and calculate a global alignment loss between the image global features and the text global features;

[0146] Optimize the trained mixture-of-experts connector based on the local contrast loss and the global alignment loss to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

[0147] Such as Figure 7As shown, when the terminal device alternately performs cross-modal local alignment training of image-to-text and text-to-image on the first hybrid expert connector, in the image-to-text stage (the operation process in the text-to-image stage is the same and will not be elaborated here), after the terminal device inputs the image feature I (fine-grained image feature) into the first hybrid expert connector, based on the local attention mechanism of the first hybrid expert connector, it first measures the importance of all image patches relative to each text token to generate a weight matrix

[0148] (A ij = softmax j (sim(I j , T i ))),

[0149] where (sim(·)) represents similarity calculation (cosine similarity).

[0150] After that, based on the weight matrix (A) through the first hybrid expert connector, the image feature I is aggregated into the text space to generate an image feature sequence (I agg ) with the same sequence length as the fine-grained text feature:

[0151] (I agg = A · I),

[0152] Finally, through the first hybrid expert connector, similarity comparison is performed to compare the similarity of each token between the aligned image feature sequence (I agg ) and the fine-grained text feature (T) to generate the total positive sample similarity (S pos ):

[0153]

[0154] Moreover, the terminal device generates the negative sample similarity (S pos ) based on the same operation of generating the positive sample similarity (S neg ). Among them, the terminal device generates negative samples by shuffling the pairing of image training data and text training data, and then calculates the negative sample similarity (S neg ).

[0155] During the process of the terminal device performing similarity comparison through the first hybrid expert connector, it also calculates the local contrast loss through the first hybrid expert connector, so as to optimize the local alignment using the contrast learning loss:

[0156]

[0157] When the terminal device performs similarity comparison through the first hybrid expert connector, it can also perform global feature distribution constraint processing through the first hybrid expert connector later or simultaneously. That is, in the image-to-text stage, the terminal device aggregates the fine-grained image feature (I) and the fine-grained text feature (T) through the first hybrid expert connector to generate the image global feature (I global ). In the subsequent text-to-image stage, the fine-grained text feature (T) and the fine-grained image feature (I) are aggregated to generate the text global feature (T global ). Then, the terminal device calculates the global alignment loss through the first hybrid expert connector, that is, constrains the distribution consistency of the image global feature (I global ) and the text global feature (T global ) through L2 regularization:

[0158]

[0159] In addition, during the process of the terminal device alternately performing cross-modal local alignment training on the first hybrid expert connector, the first hybrid expert connector combines the local contrast loss and the global alignment loss for joint optimization:

[0160]

[0161] Among them, λ is a balance coefficient used to adjust the importance of local and global alignment.

[0162] In this embodiment, after obtaining the first hybrid expert connector through the terminal device performing cross-modal global alignment training on the hybrid expert connector, the fine-grained image feature obtained by the visual model learning and encoding the image training data and the fine-grained text feature obtained by the language model learning and encoding the text training data are further input into the first hybrid expert connector, and the first hybrid expert connector is controlled to perform similarity comparison processing on the input target fine-grained features and perform distribution constraint processing on the global features corresponding to the target fine-grained features, so as to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features by alternately performing cross-modal local alignment training of image-to-text and text-to-image. In this way, through a hierarchical alignment framework from coarse-grained to fine-grained for multi-modal model training, and combining global-local contrast and L2 for hierarchical semantic alignment from coarse to fine, cross-modal semantic hierarchical modeling is achieved, thus meeting the alignment requirements for processing multi-granularity information.

[0163] Please refer to Figure 8 , Figure 8 for Figure 1 another schematic diagram of the refined step process of step S103 in

[0164] In some embodiments, the multi-modal model finally trained by the terminal device includes, in addition to the second mixture-of-experts connector, a pre-trained visual model and a language model, and the second mixture-of-experts connector is respectively connected to the visual model and the language model. In this case, as Figure 8 shown, the step of "performing cross-modal local alignment training on the first mixture-of-experts connector based on the image training data and the text training data" in the above step S103 may further include steps S801 to S804 as shown below.

[0165] Step S801: Input the target fine-grained feature into the first mixture-of-experts connector; wherein, the target fine-grained feature is any one of a fine-grained image feature and a fine-grained text feature; the fine-grained image feature is obtained by the visual model learning the image training data, and the fine-grained text feature is obtained by the language model learning the text training data.

[0166] When the terminal device further performs cross-modal local alignment training on the first mixture-of-experts connector alternately, in the image-to-text stage or the text-to-image stage, the target fine-grained feature is input into the first mixture-of-experts connector twice. That is, the terminal device first inputs the target fine-grained feature into the first mixture-of-experts connector for the first time, so that the first mixture-of-experts connector first performs a similarity comparison process on the target fine-grained feature.

[0167] Step S802: Perform a similarity comparison process on the target fine-grained feature based on the first mixture-of-experts connector.

[0168] After the terminal device inputs the target fine-grained feature into the first mixture-of-experts connector, the first mixture-of-experts connector performs a similarity comparison process on the target fine-grained feature. Among them, the process of the first mixture-of-experts connector performing a similarity comparison process on the target fine-grained feature is the same as the step and its refinement step flow of "performing a similarity comparison process on the target fine-grained feature based on the first mixture-of-experts connector" in the above step S602, and the same content will not be repeated here.

[0169] Step S803: Input the target fine-grained feature into the first mixture-of-experts connector again.

[0170] After the terminal device performs a similarity comparison process on the target fine-grained feature through the first mixture-of-experts connector, it inputs the target fine-grained feature into the first mixture-of-experts connector for the second time, so that the first mixture-of-experts connector performs a distribution constraint process on the global feature corresponding to the target fine-grained feature.

[0171] Step S804: Perform distribution constraint processing on the global feature corresponding to the target fine-grained feature based on the first hybrid expert connector, so as to perform cross-modal local alignment between the fine-grained image feature and the fine-grained text feature.

[0172] After the terminal device inputs the target fine-grained feature into the first hybrid expert connector for the second time, the first hybrid expert connector performs distribution constraint processing on the global feature corresponding to the target fine-grained feature, so as to perform cross-modal local alignment between the fine-grained image feature and the fine-grained text feature by alternately performing cross-modal local alignment training from image to text and from text to image. Among them, the process of the first hybrid expert connector performing distribution constraint processing on the global feature is the same as the steps and their refinement steps of "performing distribution constraint processing on the global feature corresponding to the target fine-grained feature based on the first hybrid expert connector" in the above step S602, and the same content will not be elaborated here.

[0173] In this embodiment, the terminal device inputs the fine-grained feature into the hybrid expert connector twice successively, so as to train the hybrid expert connector to gradually realize the multi-granularity alignment from local to global of the fine-grained image feature and the fine-grained text feature. In this way, it can provide an accurate semantic basis for the final multi-modal model to process cross-modal tasks.

[0174] In some embodiments, during the process of the terminal device performing cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data, and performing cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data, in order to enable the hybrid expert connector to perceive the inputs of different modalities and perform different tasks, the terminal device adopts the method of task embedding Cross Embedding, so that the hybrid expert connector can guide the gating network Router to select different experts according to the input modality and the task to be performed. Among them, Crossembedding includes modality embedding and task embedding. Modality embedding includes text modality and image modality, and task embedding includes global contrast learning E CL-g and local contrast learning and L2-E l2Three different tasks. When the Mixture of Experts Connector performs cross-modal alignment training for image-to-text and text-to-image, each time it combines the modal embedding and task embedding according to the input to prompt the gating network Router to select different experts to perform the corresponding tasks. For example, the terminal device embeds the image embedding plus the global contrast embedding in the target coarse-grained features input to the Mixture of Experts Connector to inform the gating network Router that the input is an image and the task to be performed is global contrast learning.

[0175] In some embodiments, when the terminal device performs cross-modal global alignment training on the Mixture of Experts Connector and cross-modal local alignment training on the first Mixture of Experts Connector, it can control the Mixture of Experts Connector to perform multiple task combinations at different time steps respectively, so as to alternately perform the alignment in the image-to-text stage and the text-to-image stage. For example, as Figure 9 shown, at the same time step, the terminal device adds the image modal embedding and the local contrast embedding to the image features input to the Mixture of Experts Connector to control the Mixture of Experts Connector to perform the contrast learning task, and at the same time adds the image embedding and the L2 embedding to control the Mixture of Experts Connector to perform the prediction task. At another time step, the terminal device adds the text embedding, L2, and contrast Embedding to the text features input to the Mixture of Experts Connector to control the Mixture of Experts Connector to perform the corresponding different tasks.

[0176] In some embodiments, the image features / text features after adding the task embedding generate a set of weights when passing through the gating network router. The Mixture of Experts Connector sorts these weights and only selects the outputs of the top k experts with the highest weight values to aggregate as the final prediction features.

[0177] Exemplarily, when the terminal device alternately performs cross-modal global alignment training or cross-modal local alignment training on the Mixture of Experts Connector, the overall process includes three main stages: feature encoding, alternating alignment based on Bi-MoE, and joint optimization combining the contrast loss and the L2 loss.

[0178] As Figure 10 shown, for the input image-text pair, the terminal device first encodes the image training data I into a feature vector f I (coarse-grained image features and fine-grained image features) through a visual model (such as a convolutional neural network), and encodes the text training data T into a feature vector f T(Coarse-grained text features and fine-grained text features). These encoded features will be used as the input for the terminal device to alternately perform cross-modal alignment training (cross-modal global alignment training and cross-modal local alignment training) on the hybrid expert connector.

[0179] At odd time steps t, the terminal device trains the hybrid expert connector to perform image-to-text alignment. The image feature vector f I Passes through the hybrid expert connector to generate a predicted text feature vector And by calculating And the true text feature vector f T Between the global and local contrast losses and the L2 loss, the alignment effect is optimized. At even time steps t, the terminal device trains the hybrid expert connector to perform text-to-image alignment. Similarly, the text feature vector f T Passes through the hybrid expert connector to generate a predicted image feature vector And by calculating And the true image feature vector f I Between the global and local-contrast losses and the L2 loss to optimize the alignment effect.

[0180] In each time step, the terminal device ensures good discriminability of features between different modalities by calculating the global and local contrast losses (CL-loss), thereby enhancing the discriminative ability of the finally trained multi-modal model for different modality features. The terminal device jointly optimizes these two losses to achieve robust bidirectional alignment of image features and text features. At the same time, it also minimizes the reconstruction error between the predicted features and the true features through the L2 loss, thereby ensuring semantic consistency between different modalities. Among them, the global contrast loss is used when the terminal device trains the hybrid expert connector to perform global representation comparison of images, emphasizing capturing the overall semantic information of samples, while the local contrast loss is used when the terminal device trains the hybrid expert connector to perform local contrast learning, paying more attention to extracting fine-grained features from local regions of the data and emphasizing the relationships and differences between local regions.

[0181] It should be noted that since the visual question answering VQA task requires the model to combine the semantics of the question and the local information of the image to generate an answer. The terminal device optimizes the local contrast loss when performing cross-modal local alignment training on the hybrid expert connector, thereby enhancing the attention ability of the final model to local regions of the image by introducing supervision of local features.

[0182] In addition, as Figure 11As shown, assume that the text training data is a sentence of length 4 and the number of patches in the image training data is 6. When the terminal device calculates the global and local contrast losses through the mixture of experts connector, first, it calculates the similarity between the first word token and the embeddings of all image patches to obtain a similarity sequence of length 6. Then, it normalizes this similarity sequence and aggregates all image patches using it as weights, thereby obtaining the weighted image embedding corresponding to the first word token. Next, it sequentially takes the second, third, and all subsequent word tokens and repeats the above steps. Finally, a sequence of weighted image embeddings consistent with the number of word tokens is obtained.

[0183] The similarity for local contrast learning is calculated based on the weighted image embedding of each word token and its corresponding token embedding, and the contrast loss is calculated after aggregating them in order. For global contrast learning, the weighted representations of all image patches and the representations of text tokens are respectively aggregated into global features, and then the global contrast loss is calculated.

[0184] In some embodiments, to verify the effectiveness of the multi-modal model finally trained by the terminal device, the terminal device fine-tunes and trains the multi-modal model Alt-MoE finally trained based on the training method provided in the embodiments of the present application on the public dataset VQAv2, and obtains the Figure 12 experimental results as shown, which show that Alt-MoE achieves advanced results compared with various baselines. Moreover, the terminal device also conducts experiments on NVLR-2, and the experimental results also prove that the multi-modal model finally trained based on the training method provided in the embodiments of the present application achieves advanced results when compared with multiple baselines.

[0185] Next, various embodiments of the processing method for the visual question answering task provided in the embodiments of the present application are presented.

[0186] Please refer to Figure 13 , Figure 13 which is a schematic flowchart of the steps in some embodiments of the processing method for the visual question answering task provided in the embodiments of the present application.

[0187] In some embodiments, as Figure 13 shown, the task processing method for multi-modal learning provided in the embodiments of the present application may include step S1301 and step S1302 as shown below.

[0188] Step S1301: Obtain the data to be processed for the visual question answering task, where the data to be processed includes image data and text data, and the text data is the natural language question for the image data.

[0189] The terminal device can perform a visual question answering task based on the multi-modal model finally trained in each of the above embodiments of the training method. Moreover, the terminal device first obtains the data to be processed for the visual question answering task to be performed currently. Among them, the data to be processed obtained by the terminal device includes image data and text data, and the text data is a natural language question for the image data.

[0190] Exemplarily, when performing a visual question answering task, the terminal device first obtains the image data in the image modality and the text data in the text modality (a natural language question for the image data), and uses the obtained image data and text data as the data to be processed for the visual question answering task to be performed currently.

[0191] Step S1302: Input the image data and the text data into a preset multi-modal model to obtain the answer to the visual question answering task; wherein, the mixture of experts connector in the multi-modal model performs global alignment and local alignment on the image modality information corresponding to the image data and the text modality information corresponding to the text data.

[0192] The terminal device inputs the obtained data to be processed into the multi-modal model. Thus, after processing the data to be processed through the multi-modal model, the answer to the visual question answering task is output. Among them, based on the mixture of experts connector in its own model architecture, the multi-modal model first performs cross-modal global alignment on the coarse-grained image features of the image data and the coarse-grained text features of the text data, and then performs cross-modal local alignment on the fine-grained image features of the image data and the fine-grained text features of the text data, so as to globally understand the image data and the text data, and also carefully capture the corresponding relationship between the text semantics and the image regions, and further accurately generate the answer to the visual question answering task.

[0193] Exemplarily, when the terminal device uses the above multi-modal model to perform a retrieval-based visual question answering (VQA) task, it models an image and its corresponding question by adopting the above training method to find the correct answer. For the input image and question, the terminal device encodes them respectively using a visual model and a language model and then aligns them through the pre-trained MoE. And in order to model the fine-grained semantic relationship between the question and the corresponding image, the terminal device further uses local attention to model the fine-grained semantic relationship between the image and the question. Among them, as Figure 14As shown, local attention is a local attention algorithm for aligning text tokens and image patches. For a given sequence of text tokens and a sequence of image patches, the local attention algorithm traverses each text token one by one, calculates its cosine similarity with all image patches, normalizes the similarity using softmax, and then uses the normalized weights to perform weighted aggregation on the image patches to generate the image representation corresponding to the text token. Finally, an image weighted representation sequence with the same number of text tokens is output. Finally, after the terminal device uses local attention for local alignment, the aligned features are concatenated together as the visual language alignment representation, and then passed through a classifier to compress the dimension of the aligned features to be the same as the category label, so as to find the correct answer.

[0194] In the embodiments of the present application, a hybrid expert connector designed based on the dynamic mixture of experts (MoE) module is introduced through the terminal device to achieve multi-granularity dynamic control of image features and text features, so as to complete hierarchical alignment from global to local. In this way, when processing cross-modal tasks such as visual question answering (VQA), for the diverse requirements for determining the final task result (since the alignment requirements between the question and the image vary depending on the specific scenario: some questions need to capture the coarse-grained overall semantics of the image, such as "What is the main color of this picture?", some questions focus on the fine-grained detailed information of the image, such as "How many apples are there on the table?", and some questions may require both global and local information at the same time, such as "What is the object closest to the red chair in the picture?"), in the embodiments of the present application, through the hybrid expert connector in the multi-modal model, according to the dynamic requirements of the task, the feature information flow is flexibly allocated between coarse-grained and fine-grained alignment. By using the hybrid expert connector as the connection hub, the coarse-grained alignment stage mainly focuses on the mapping and alignment of global semantic features, while the fine-grained alignment stage focuses on the correlation modeling of local features. In this way, not only can the mutual interference between coarse-grained and fine-grained alignment be effectively isolated, but also the most suitable alignment granularity can be dynamically selected according to the specific questions of the visual question answering (VQA) task.

[0195] Please refer to Figure 15 , the embodiments of the present application also provide a training device for a multi-modal model, which can implement the above-mentioned training method of the multi-modal model. The device includes a first acquisition module 1501, a first-stage training module 1502, and a second-stage training module 1503. Among them,

[0196] The first acquisition module 1501 is used to acquire image training data and text training data, and the text training data is the natural language question of the image training data;

[0197] The first-stage training module 1502 is configured to perform cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data to obtain a first hybrid expert connector, where the first hybrid expert connector is used to globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data;

[0198] The second-stage training module 1503 is configured to perform cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data to obtain a multi-modal model including a second hybrid expert connector;

[0199] Wherein, the second hybrid expert connector is obtained by performing local alignment training on the first hybrid expert connector, and the second hybrid expert connector is used to locally align the fine-grained image features of the image training data and the fine-grained text features of the text training data; the multi-modal model is configured to perform global alignment and local alignment of the image modality information and the text modality information based on the second hybrid expert connector to obtain the answer to the visual question answering task.

[0200] In some embodiments, the first-stage training module 1502 is further configured to obtain the coarse-grained image features of the image training data; and obtain the coarse-grained text features of the text training data; and perform cross-modal global alignment training on the hybrid expert connector to be trained based on the coarse-grained image features and the coarse-grained text features.

[0201] In some embodiments, the first-stage training module 1502 is further configured to input the target coarse-grained feature into the hybrid expert connector to be trained; wherein, the target coarse-grained feature is any one of the coarse-grained image features and the coarse-grained text features; and perform a prediction task and a contrast task on the target coarse-grained feature based on the hybrid expert connector to perform cross-modal global alignment of the coarse-grained image features and the coarse-grained text features;

[0202] Wherein, when the target coarse-grained feature is the coarse-grained image feature, the prediction task is to generate coarse-grained predicted text features based on the coarse-grained image feature, and the contrast task is to compare the coarse-grained predicted text features with the target coarse-grained text features in the non-sequence dimension; the target coarse-grained text features are obtained based on aggregating the coarse-grained text features; when the target coarse-grained feature is the coarse-grained text feature, the prediction task is to generate coarse-grained predicted image features based on the coarse-grained text feature, and the contrast task is to compare the coarse-grained predicted image features with the target coarse-grained image features in the non-sequence dimension; the target coarse-grained image features are obtained based on aggregating the coarse-grained image features.

[0203] In some embodiments, the first-stage training module 1502 is further configured to calculate a cross-modal global alignment loss between the predicted text feature and the target coarse-grained text feature; wherein, the cross-modal global alignment loss at least includes a contrast loss; and, optimize the hybrid expert connector based on the contrast loss to perform cross-modal global alignment of the coarse-grained image feature and the coarse-grained text feature.

[0204] In some embodiments, the multi-modal model further includes a visual model and a language model, and the second hybrid expert connector is respectively connected to the visual model and the language model; the second-stage training module 1503 is further configured to input the target fine-grained feature into the first hybrid expert connector; wherein, the target fine-grained feature is any one of a fine-grained image feature and a fine-grained text feature; the fine-grained image feature is obtained by the visual model learning the image training data, and the fine-grained text feature is obtained by the language model learning the text training data; and, perform a similarity comparison process on the target fine-grained feature based on the first hybrid expert connector, and perform a distribution constraint process on the global feature corresponding to the target fine-grained feature based on the first hybrid expert connector to perform cross-modal local alignment of the fine-grained image feature and the fine-grained text feature.

[0205] In some embodiments, the second-stage training module 1503 is further configured to, based on the first hybrid expert connector, compare the image feature sequence corresponding to the fine-grained image feature with the fine-grained text feature to obtain sample similarity data; and use an L2 regularization loss function to constrain the distributions of the global features corresponding to the fine-grained image feature and the fine-grained text feature respectively; wherein, the sequence length of the image feature sequence is the same as the sequence length of the fine-grained text feature; the sample similarity data includes positive sample similarity data and negative sample similarity data; the global features include the image global feature corresponding to the fine-grained image feature and the text global feature corresponding to the fine-grained text feature; calculate a local contrast loss based on the positive sample similarity data and the negative sample similarity data; and calculate a global alignment loss between the image global feature and the text global feature; and optimize the trained hybrid expert connector based on the local contrast loss and the global alignment loss to perform cross-modal local alignment of the fine-grained image feature and the fine-grained text feature.

[0206] In some embodiments, the multimodal model further includes a visual model and a language model, and the second hybrid expert connector is respectively connected to the visual model and the language model; the second-stage training module 1503 is further configured to input the target fine-grained feature into the first hybrid expert connector; wherein, the target fine-grained feature is any one of a fine-grained image feature and a fine-grained text feature; the fine-grained image feature is obtained by the visual model learning the image training data, and the fine-grained text feature is obtained by the language model learning the text training data; performing a similarity comparison process on the target fine-grained feature based on the first hybrid expert connector; inputting the target fine-grained feature into the first hybrid expert connector again; and performing a distribution constraint process on the global feature corresponding to the target fine-grained feature based on the first hybrid expert connector to perform cross-modal local alignment between the fine-grained image feature and the fine-grained text feature.

[0207] The specific implementation manner of the training device for the multimodal model provided in the embodiments of the present application is basically the same as the specific embodiments of the above-mentioned multimodal model training method, and will not be elaborated here.

[0208] Please refer to Figure 16 , the embodiments of the present application further provide a processing device for a visual question answering task, which can implement the above-mentioned processing method for the visual question answering task. The device includes a second acquisition module 1601 and a task processing module 1602. Among them,

[0209] The second acquisition module 1601 is configured to acquire the data to be processed for the visual question answering task. The data to be processed includes image data and text data, and the text data is a natural language question about the image data;

[0210] The task processing module 1602 is configured to input the image data and the text data into a preset multimodal model to obtain an answer to the visual question answering task;

[0211] Among them, the hybrid expert connector in the multimodal model performs global alignment and local alignment on the image modality information corresponding to the image data and the text modality information corresponding to the text data.

[0212] The specific implementation manner of the processing device for the visual question answering task provided in the embodiments of the present application is basically the same as the specific embodiments of the above-mentioned processing method for the visual question answering task, and will not be elaborated here.

[0213] The embodiments of the present application also provide a processing device for visual question answering tasks. The processing device for visual question answering tasks provided by the embodiments of the present application may be a computer device, which may specifically be an in-vehicle terminal device, an intelligent robot, a terminal device equipped with an embodied intelligent system, or may also be a computer device such as a smart phone, a tablet computer, a laptop computer, or a desktop computer. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned model training method for multi-modal learning. In some embodiments, the computer device may also be any intelligent terminal such as an in-vehicle computer, a portable PC, and a wearable device.

[0214] Please refer to Figure 17 , Figure 17 which schematically shows the hardware structure of a computer device in an embodiment. The computer device includes:

[0215] A processor 1701, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0216] A memory 1702, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1702 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1702, and the processor 1701 is used to call and execute the training method of the multi-modal model and / or the processing method of the visual question answering task of the embodiments of the present application;

[0217] An input / output interface 1703, which is used to implement information input and output;

[0218] A communication interface 1704, which is used to implement communication and interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);

[0219] A bus 1705, which transmits information between various components of the device (such as the processor 1701, the memory 1702, the input / output interface 1703, and the communication interface 1704);

[0220] Among them, the processor 1701, the memory 1702, the input / output interface 1703, and the communication interface 1704 are communicatively connected to each other inside the device through the bus 1705.

[0221] An embodiment of the present application further provides a vehicle, on which a computer device is configured. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the training method of the above multi-modal model and / or the processing method of the visual question answering task are implemented.

[0222] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the training method of the above multi-modal model and / or the processing method of the visual question answering task are implemented.

[0223] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0224] An embodiment of the present application further provides a computer program product, including a computer program. The steps implemented when the computer program is executed by a processor are substantially the same as the specific embodiments of the training method of the above multi-modal model and / or the task processing method of multi-modal learning, and will not be described in detail here.

[0225] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0226] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine some steps, or different steps.

[0227] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0228] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0229] It should be understood that in this application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0230] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0231] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0232] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0233] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0234] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0235] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings, but this does not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A method for training a multimodal model, characterized in that: The method comprises: Acquire image training data and text training data, wherein the text training data is a natural language question of the image training data; Performing cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data to obtain a first hybrid expert connector, wherein the first hybrid expert connector is used to globally align the coarse-grained image features of the image training data and the coarse-grained text features of the text training data; Performing cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data to obtain a multimodal model including a second hybrid expert connector; Among them, the second hybrid expert connector is obtained by performing cross-modal local alignment training on the first hybrid expert connector, and the second hybrid expert connector is used to perform local alignment on the fine-grained image features of the image training data and the fine-grained text features of the text training data; the multimodal model is used to perform global and local alignment of image modal information and text modal information based on the second hybrid expert connector to obtain the answer to the visual question answering task.

2. The method according to claim 1, characterized in that The performing cross-modal global alignment training on the hybrid expert connector based on the image training data and the text training data includes: Acquire coarse-grained image features of the image training data; and acquire coarse-grained text features of the text training data; Based on the coarse-grained image features and the coarse-grained text features, cross-modal global alignment training is performed on the hybrid expert connector to be trained.

3. The method according to claim 2, characterized in that The performing cross-modal global alignment training on the hybrid expert connector to be trained based on the coarse-grained image features and the coarse-grained text features includes: Inputting the target coarse-grained feature into the hybrid expert connector to be trained; wherein the target coarse-grained feature is any one of the coarse-grained image feature and the coarse-grained text feature; Based on the hybrid expert connector, a prediction task and a comparison task are performed on the target coarse-grained features to perform cross-modal global alignment of the coarse-grained image features and the coarse-grained text features; Among them, when the target coarse-grained feature is the coarse-grained image feature, the prediction task is to generate a coarse-grained predicted text feature based on the coarse-grained image feature, and the comparison task is to compare the coarse-grained predicted text feature with the target coarse-grained text feature of the non-sequential dimension; the target coarse-grained text feature is obtained based on aggregating the coarse-grained text feature; when the target coarse-grained feature is the coarse-grained text feature, the prediction task is to generate a coarse-grained predicted image feature based on the coarse-grained text feature, and the comparison task is to compare the coarse-grained predicted image feature with the target coarse-grained image feature of the non-sequential dimension; the target coarse-grained image feature is obtained based on aggregating the coarse-grained image feature.

4. The method according to claim 3, characterized in that The method further comprises: Calculating a cross-modal global alignment loss between the predicted text feature and the target coarse-grained text feature; wherein the cross-modal global alignment loss at least includes a contrast loss; The hybrid expert connector is optimized based on the contrast loss to perform cross-modal global alignment of the coarse-grained image features and the coarse-grained text features.

5. The method according to claim 1, characterized in that The multimodal model further includes a visual model and a language model, and the second hybrid expert connector is connected to the visual model and the language model respectively; The performing cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data includes: Inputting a target fine-grained feature into the first hybrid expert connector; wherein the target fine-grained feature is any one of a fine-grained image feature and a fine-grained text feature; the fine-grained image feature is obtained by learning the image training data based on the visual model, and the fine-grained text feature is obtained by learning the text training data based on the language model; Based on the first hybrid expert connector, similarity comparison processing is performed on the target fine-grained features, and based on the first hybrid expert connector, distribution constraint processing is performed on the global features corresponding to the target fine-grained features to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

6. The method according to claim 5, characterized in that The performing similarity comparison processing on the target fine-grained features based on the first hybrid expert connector, and the performing distribution constraint processing on the global features corresponding to the target fine-grained features based on the first hybrid expert connector, include: Based on the first hybrid expert connector, the image feature sequence corresponding to the fine-grained image feature is compared with the fine-grained text feature for similarity to obtain sample similarity data; and the L2 regularization loss function is used to constrain the distribution of the global features corresponding to the fine-grained image feature and the fine-grained text feature; wherein the sequence length of the image feature sequence is consistent with the sequence length of the fine-grained text feature; the sample similarity data includes positive sample similarity data and negative sample similarity data; the global features include image global features corresponding to the fine-grained image feature and text global features corresponding to the fine-grained text feature; Calculating a local contrast loss based on the positive sample similarity data and the negative sample similarity data; and calculating a global alignment loss between the image global feature and the text global feature; The trained hybrid expert connector is optimized based on the local contrast loss and the global alignment loss to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

7. The method according to claim 1, characterized in that The multimodal model further includes a visual model and a language model, and the second hybrid expert connector is connected to the visual model and the language model respectively; The performing cross-modal local alignment training on the first hybrid expert connector based on the image training data and the text training data includes: Inputting a target fine-grained feature into the first hybrid expert connector; wherein the target fine-grained feature is any one of a fine-grained image feature and a fine-grained text feature; the fine-grained image feature is obtained by learning the image training data based on the visual model, and the fine-grained text feature is obtained by learning the text training data based on the language model; performing similarity comparison processing on the target fine-grained features based on the first hybrid expert connector; inputting the target fine-grained features into the first hybrid expert connector again; Based on the first hybrid expert connector, distribution constraint processing is performed on the global features corresponding to the target fine-grained features to perform cross-modal local alignment of the fine-grained image features and the fine-grained text features.

8. A method for processing a visual question answering task, characterized in that: The method comprises: Obtaining data to be processed for a visual question answering task, wherein the data to be processed includes image data and text data, and the text data is a natural language question about the image data; Inputting the image data and the text data into a preset multimodal model to obtain an answer to the visual question answering task; Among them, the hybrid expert connector in the multimodal model globally aligns and locally aligns the image modality information corresponding to the image data and the text modality information corresponding to the text data.

9. A processing device for a visual question answering task, characterized in that: The processing device for the visual question answering task includes a memory and a processor, the memory stores a computer program, and the processor implements the training method of the multimodal model described in any one of claims 1 to 7 and / or implements the processing method of the visual question answering task described in claim 8 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the training method of the multimodal model described in any one of claims 1 to 7, and / or implements the processing method of the visual question answering task described in claim 8.

Citation Information

Cited By

  • Human-centered vision model training method and system based on dynamic mode alignment

    CN121582747A

  • Cell image classification method based on morphological semantic guidance

    CN121937808A