Multi-vision model training method, multi-modal task processing method and equipment
By training multi-vision models, using hybrid expert connectors to align image feature vectors output from different visual models, the problem that a single visual model cannot fully capture and analyze image features is solved, and more efficient multi-modal task processing performance is achieved.
Patent Information
- Application Number
- CN202510235188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The existing single vision model cannot fully capture and analyze image features, resulting in poor performance in its model processing multimodal tasks based on visual understanding.
A multi-vision model training method is proposed. By obtaining image training data and inputting it into two visual models trained in different ways, the respective image feature vectors are obtained, and then these feature vectors are input into the hybrid expert connector for alignment, and a multi-modal task processing model based on visual understanding is trained.
Through the hybrid expert connector, aligning the image feature vectors output from different visual models can achieve comprehensive capture and analysis of image features, improving the performance of model processing multimodal tasks.
Smart Images

Figure CN120164058A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method for multi-vision models, a multi-modal task processing method, and a device. Background Art
[0002] With the development of multi-modal image classification technology, using deep learning technology to effectively fuse information from different modalities to improve the accuracy of image classification has become an important research direction. In related technologies, visual models are mostly used to process various multi-modal tasks that require visual understanding, such as image classification. However, visual models trained in different ways focus on different directions of learning image features. For example, some visual models focus on capturing and analyzing the features of local details of images, while some visual models are more focused on capturing and analyzing the high-level semantic features of images. That is to say, existing single visual models cannot comprehensively capture and analyze image features, resulting in poor model performance in processing multi-modal tasks based on visual understanding. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a training method for multi-vision models, a multi-modal task processing method, and a device, aiming to comprehensively capture and analyze image features based on visual models, thereby improving the performance of models in processing multi-modal tasks based on visual understanding.
[0004] To achieve the above object, in the first aspect of the embodiments of this application, a training method for multi-vision models is proposed, and the method includes:
[0005] Obtain model training data; wherein, the model training data at least includes image training data;
[0006] Input the image training data into a first vision model and a second vision model to obtain a first image feature vector output by the first vision model and a second image feature vector output by the second vision model;
[0007] Input the first image feature vector and the second image feature vector into a to-be-trained hybrid expert connector to train the hybrid expert connector to obtain a multi-modal task processing model based on visual understanding;
[0008] Wherein, the multi-modal task processing model includes the first vision model, the second vision model, and the trained hybrid expert connector; the trained hybrid expert connector is used to align the first image feature and the second image feature; the multi-modal task processing model is used to process multi-modal data including at least image data to obtain a processing result of the multi-modal task.
[0009] In some embodiments, training the hybrid expert connector with the first image feature vector and the second image feature vector includes:
[0010] Inputting a target image feature vector into the hybrid expert connector to be trained; wherein, the target image feature vector is any one of the first image feature vector and the second image feature vector;
[0011] Performing an alignment task on the target image feature vector based on the hybrid expert connector to align the first image feature vector and the second image feature vector.
[0012] Inputting the target image feature vector into the hybrid expert connector to be trained includes:
[0013] Adding alignment task parameters to the target image feature vector; wherein, the alignment task parameters include contrast learning alignment task parameters and L2 squared difference alignment task parameters;
[0014] Inputting the target image feature vector added with the contrast learning alignment task parameters into the hybrid expert connector to be trained, and inputting the target image feature vector added with the L2 squared difference alignment task parameters into the hybrid expert connector to be trained;
[0015] Performing an alignment task on the target image feature vector based on the hybrid expert connector includes:
[0016] Performing a contrast learning alignment task on the target image feature vector added with the contrast learning alignment task parameters based on the hybrid expert connector, and performing an L2 squared difference alignment task on the target image feature vector added with the L2 squared difference alignment task parameters based on the hybrid expert connector.
[0017] In some embodiments, the model training data further includes text training data; the method further includes:
[0018] Inputting the text training data into a language model to obtain text feature vectors;
[0019] Matching the alignment feature vector output by the hybrid expert connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the alignment feature vector and the text feature vector; wherein, the alignment loss includes a contrast loss and an L2 squared difference loss;
[0020] Optimizing the hybrid expert connector for aligning the first image feature vector and the second image feature vector based on the contrast loss and the L2 squared difference loss.
[0021] In some embodiments, the alignment task includes a contrastive learning alignment task and an L2 squared difference alignment task;
[0022] Matching the alignment feature vector output by the hybrid expert connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the alignment feature vector and the text feature vector, includes:
[0023] Matching the first alignment feature vector output by the hybrid expert connector performing the contrastive learning alignment task on the first image feature vector with the text feature vector, and calculating the contrastive loss between the first alignment feature vector and the text feature vector;
[0024] Matching the second alignment feature vector output by the hybrid expert connector performing the L2 squared difference alignment task on the second image feature vector with the text feature vector, and calculating the L2 squared difference loss between the second alignment feature vector and the text feature vector.
[0025] In some embodiments, the alignment task includes a contrastive learning alignment task and an L2 squared difference alignment task;
[0026] Matching the alignment feature vector output by the hybrid expert connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the alignment feature vector and the text feature vector, includes:
[0027] Aggregating the first alignment feature vector output by the hybrid expert connector performing the contrastive learning alignment task on the first image feature vector and the third alignment feature vector output by the hybrid expert connector performing the contrastive learning alignment task on the second image feature vector to obtain a first aggregated feature vector;
[0028] Matching the first aggregated feature vector with the text feature vector, and calculating the contrastive loss between the first aggregated feature vector and the text feature vector;
[0029] Aggregating the fourth alignment feature vector output by the hybrid expert connector performing the L2 squared difference alignment task on the first image feature vector and the second alignment image feature vector output by the hybrid expert connector performing the L2 squared difference alignment task on the second image feature vector to obtain a second aggregated feature vector;
[0030] Matching the second aggregated feature vector with the text feature vector, and calculating the L2 squared difference loss between the second aggregated feature vector and the text feature vector.
[0031] To achieve the above object, a second aspect of the embodiments of the present application proposes a multi-modal task processing method, and the method includes:
[0032] Obtain multi-modal data to be processed for a multi-modal task; wherein, the multi-modal data at least includes image data;
[0033] Input the image data into a preset multi-modal task processing model based on visual understanding to obtain a processing result of the multi-modal task;
[0034] Wherein, a first visual model and a second visual model in the multi-modal task processing model respectively learn the image data to obtain corresponding first image feature vectors and second image feature vectors; a hybrid expert connector in the multi-modal task processing model aligns the first image feature vector and the second image feature vector.
[0035] To achieve the above object, a third aspect of the embodiments of the present application proposes a training device for a multi-visual model, and the device includes:
[0036] A first acquisition module, configured to acquire model training data; wherein, the model training data at least includes image training data;
[0037] A single visual model processing module, configured to input the image training data into a first visual model and a second visual model to obtain a first image feature vector output by the first visual model and a second image feature vector output by the second visual model;
[0038] A multi-level image feature alignment module, configured to input the first image feature vector and the second image feature vector into a hybrid expert connector to be trained, and train the hybrid expert connector to obtain a multi-modal task processing model based on visual understanding;
[0039] Wherein, the multi-modal task processing model includes the first visual model, the second visual model and the trained hybrid expert connector; the trained hybrid expert connector is used to align the first image feature and the second image feature; the multi-modal task processing model is used to process multi-modal data including at least image data to obtain a processing result of the multi-modal task.
[0040] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a multi-modal task processing device, and the device includes:
[0041] A second acquisition module, configured to acquire multi-modal data to be processed for a multi-modal task; wherein, the multi-modal data at least includes image data;
[0042] A multi-modal task processing module, configured to input the image data into a preset multi-modal task processing model based on visual understanding to obtain a processing result of the multi-modal task;
[0043] Wherein, a first visual model and a second visual model in the multi-modal task processing model respectively learn the image data to obtain corresponding first and second image feature vectors; a mixture-of-experts connector in the multi-modal task processing model aligns the first image feature vector and the second image feature vector.
[0044] To achieve the above object, a fifth aspect of the embodiments of the present application provides a multi-modal task processing device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method described in the first aspect above, and / or implements the method described in the second aspect above.
[0045] To achieve the above object, a sixth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect above, and / or implements the method described in the second aspect above.
[0046] To achieve the above object, a seventh aspect of the embodiments of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method provided in the first aspect above, and / or implements the method provided in the second aspect above.
[0047] In the embodiments of the present application, image training data is obtained and input into a first visual model. The first visual model learns the image training data and outputs a first image feature vector, and the image training data is input into a second visual model. The second visual model learns the image training data and outputs a second image feature vector. Then, the first image feature vector and the second image feature vector are input into a to-be-trained mixture-of-experts connector to train the mixture-of-experts connector to align the first image feature vector and the second image feature vector, thereby training a multi-modal task processing model based on visual understanding. Among them, the multi-modal task processing model includes a first visual model, a second visual model, and a mixture-of-experts connector, and the multi-modal task processing model is used to process multi-modal data including at least image data to obtain a processing result of the multi-modal task.
[0048] Thus, the multi-modal task processing model finally trained in the embodiments of the present application can adapt to the feature distribution differences of at least two visual models through the trained hybrid expert connector, and achieve precise alignment of the image feature vectors output by different visual models. Compared with using a single visual model for image feature capture and analysis, in the embodiments of the present application, local detail features of an image are captured and analyzed based on one visual model, and global semantic features of the image are captured and analyzed based on another visual model, and then the image feature vectors output by the two visual models are aligned through the hybrid expert connector. That is to say, the multi-modal task processing model finally trained in the embodiments of the present application can efficiently align multi-source visual features through the hybrid expert connector. In this way, it is possible to comprehensively capture and analyze the features of the entire image, so that when processing multi-modal tasks based on visual understanding using the multi-modal task processing model, multi-modal data can be processed more accurately and flexibly, thereby improving the performance of the model in processing multi-modal tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 FIG. is a schematic flowchart of steps in some embodiments of the training method for multi-visual models provided by the embodiments of the present application;
[0050] Figure 2 is Figure 1 a schematic flowchart of the refined steps of step S103 in FIG.
[0051] Figure 3 FIG. is a schematic flowchart of steps in other embodiments of the training method for multi-visual models provided by the embodiments of the present application;
[0052] Figure 4 is Figure 3 a schematic flowchart of a refined step of step S302 in FIG.
[0053] Figure 5 is Figure 3 a schematic flowchart of another refined step of step S302 in FIG.
[0054] Figure 6 FIG. is a schematic diagram of the alignment process of multiple visual models involved in the training method for multi-visual models provided by the embodiments of the present application in some embodiments;
[0055] Figure 7 FIG. is a schematic diagram of the results of a classification comparison experiment based on a linear probe involved in the training method for multi-visual models provided by the embodiments of the present application in some embodiments;
[0056] Figure 8 FIG. is a schematic flowchart of steps in some embodiments of the multi-modal task processing method provided by the embodiments of the present application;
[0057] Figure 9 Structural schematic diagram of a training device for a multi-vision model provided by an embodiment of the present application;
[0058] Figure 10 Structural schematic diagram of a multi-modal task processing device provided by an embodiment of the present application;
[0059] Figure 11 It is a hardware structural schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0060] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0061] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application, and are not intended to limit this application.
[0063] First, the overall concept of the embodiments of the present application will be described.
[0064] In the field of multi-modal image classification, with the development of deep learning technology, how to effectively fuse information from different modalities to improve the accuracy and robustness of image classification has become an important research direction. Traditional image classification methods often rely on information of a single modality, but this method ignores the complementary information that may exist between different modality data. In recent years, with the development of multi-modal image classification technology, it has become possible to capture and analyze image features from multiple perspectives, thereby improving the accuracy of image classification.
[0065] In related technologies, visual models are mostly used to process various multimodal tasks that require visual understanding, such as image classification. For example, image features are captured and analyzed through pure visual models, or text-guided visual models are used for image feature capture and analysis. Among them, pure visual models (visual models trained only on images, such as Dinov2) can capture detailed features well by directly learning from the pixels of images, such as low-level and fine-grained information of images like texture, color, and edges. However, pure visual models usually rely on the saliency of pixel features and it is difficult to capture the global semantic information of images or the complex semantic relationships between objects. Moreover, when faced with unseen scenarios or open categories, the generalization ability of pure visual models is not as good as that of text-guided visual models. In addition, text-guided visual models (visual models that learn visual representations guided by the descriptive text corresponding to images, such as CLIP) can understand complex semantic relationships through the joint learning of language descriptions and images, such as attributes ("such as the color of the car") and context ("the car is parked next to the house"). And text-guided visual models can utilize the rich knowledge in language, thus having strong generalization ability on unseen categories. However, due to possible inaccuracies or incompleteness in language descriptions leading to biases, text-guided visual models are less sensitive to the low-level visual information of images and may even ignore pixel-level features.
[0066] Generally speaking, when using visual models to process various multimodal tasks that require visual understanding, such as image classification, visual models trained in different ways focus on different directions of learning image features. For example, some visual models focus on capturing and analyzing the features of local details of images, while some visual models are more focused on capturing and analyzing the high-level semantic features of images. That is to say, existing single visual models cannot comprehensively capture and analyze image features, resulting in poor model performance in processing multimodal tasks based on visual understanding.
[0067] Based on this, in the embodiments of the present application, by introducing an additional mixture-of-experts module, the visual models trained in different ways are aligned, so that the advantages of different visual models complement each other. For example, when the pure visual model focuses on pixel-level details of images, while the text-guided visual model is good at understanding the high-level semantics of images, aligning the pure visual model and the text-guided visual model through the mixture-of-experts module can achieve the complementarity of image details and semantics. In this way, by combining the pure visual model and the text-guided visual model, the local details and global semantic information of the image can be captured simultaneously, thereby improving the expression ability of the model (including the pure visual model, the text-guided visual model, and the model of the mixture-of-experts module) for complex scenarios such as cross-modal alignment and domain transfer. Among them, aligning the pure visual model and the text-guided visual model through the mixture-of-experts module means: aligning the fine-grained features of the image output by the pure visual model with the high-level semantic information of the image output by the text-guided model. In the embodiments of the present application, by aligning the fine-grained features with the high-level semantic information, multi-modal feature fusion is enhanced, so that the trained model can obtain a more consistent and comprehensive image representation, thereby improving the model performance in multi-modal tasks such as classification, retrieval, and generation.
[0068] In summary, in the embodiments of the present application, by combining the pure visual model and the text-guided visual model and aligning their feature representations through a unified training framework, the respective advantages can be fully utilized to solve the limitations of a single model. This combination can not only improve the performance of the finally trained model for multi-modal tasks, but also promote the comprehensive development of the model in aspects such as detail capture, semantic understanding, and open-world adaptability, thereby providing a more powerful solution for visual understanding tasks.
[0069] Next, the training method of the multi-visual model, the multi-modal task processing method, the training device of the multi-visual model, the multi-modal task processing device, the multi-modal task processing equipment, the computer-readable storage medium, and the computer program product provided by the embodiments of the present application will be specifically described through the following embodiments.
[0070] It should be noted that the training method for multi-vision models and the multi-modal task processing method provided in the embodiments of this application relate to the field of artificial intelligence technology. The training method for multi-vision models and the multi-modal task processing method provided in the embodiments of this application can be applied to terminals, can also be applied to server sides, or can be software running on terminals or server sides. In some embodiments, the terminal can be a terminal device such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart wearable device, etc. The server side can be the background server device of the foregoing various terminal devices, which can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The software can be an application, a computer program, and a storage medium carrying the computer program that implement the training method for multi-vision models and / or the multi-modal task processing method, etc. It should be understood that based on different design requirements of actual applications, in different feasible embodiments, the terminals, server sides, and software, etc. that apply the training method for multi-vision models and / or the multi-modal task processing method provided in the embodiments of this application can, of course, also be other forms not listed here. The training method for multi-vision models and the multi-modal task processing method provided in the embodiments of this application do not specifically limit this.
[0071] In addition, this application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: vehicles, personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0072] For ease of understanding and description, in the following text, the terminal device applying the training method of the multi-vision model and the multi-modal task processing method provided by the embodiments of the present application is taken as an example to elaborate on each specific embodiment of the present application in detail. The implementation of any of the above-mentioned forms of the subject applying the training method of the multi-vision model and the multi-modal task processing method provided by the embodiments of the present application can refer to the process of the terminal device applying the training method of the multi-vision model and the multi-modal task processing method described later.
[0073] It should be noted that in each specific implementation manner of the present application, when it comes to relevant processing that needs to be carried out based on data related to the user's identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0074] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the steps of the training method of the multi-vision model provided by the embodiments of the present application in some embodiments. It should be understood that although Figure 1 and the subsequent other schematic flowcharts show the execution order of some method steps, due to different design requirements in actual applications, the training method of the multi-vision model provided by the embodiments of the present application can of course adopt an execution order different from that shown in the figures. That is, Figure 1 the order of the shown method steps does not constitute a limitation on the execution logic order of the training method of the multi-vision model provided by the embodiments of the present application, and any reasonable changes based on Figure 1 the order of the shown method steps should be included in the protection scope of the training method of the multi-vision model provided by the embodiments of the present application.
[0075] As Figure 1 shown, in some embodiments, when a vehicle applies the training method of the multi-vision model provided by the embodiments of the present application, it may include steps S101 to S103.
[0076] Step S101: Obtain model training data; wherein, the model training data at least includes image training data.
[0077] When the terminal device trains a multi-modal task processing model based on visual understanding, it first obtains model training data that at least includes image training data.
[0078] In some embodiments, the terminal device can obtain multi-modal model training data through the COCO dataset (the COCO dataset is a large and rich object detection, segmentation, and caption dataset) and the flickr30k image annotation dataset. Among them, the model training data obtained by the terminal device includes, in addition to image training data, data of modal types such as text, audio, or video.
[0079] Step S102: Input the image training data into the first visual model and the second visual model to obtain a first image feature vector output by the first visual model and a second image feature vector output by the second visual model.
[0080] After the terminal device obtains the model training data, it inputs the image training data in the model training data into the first visual model and the second visual model at the same time. Then, after the first visual model learns and encodes the image training data to obtain the first image feature vector and outputs the first image feature vector, and after the second visual model learns and encodes the image training data to obtain the second image feature vector and outputs the second image feature vector.
[0081] It should be noted that the first visual model and the second visual model are visual models trained in different ways, and the first visual model and the second visual model focus on different contents when capturing and analyzing image features. For example, the first visual model is a pure visual model that focuses on fine-grained features in the image, while the second visual model is a text-guided visual model that focuses on high-level semantic information in the image.
[0082] Step S103: Input the first image feature vector and the second image feature vector into the hybrid expert connector to be trained, and train the hybrid expert connector to obtain a multi-modal task processing model based on visual understanding; wherein, the multi-modal task processing model includes the first visual model, the second visual model, and the trained hybrid expert connector; the trained hybrid expert connector is used to align the first image feature and the second image feature; the multi-modal task processing model is used to process multi-modal data including at least image data to obtain the processing result of the multi-modal task.
[0083] After obtaining the first image feature vector output by the first vision model and the second image feature vector output by the second vision model, the terminal device further inputs the first image feature vector and the second image feature vector into the hybrid expert connector to be trained, so as to train the hybrid expert connector to align the first image feature with the second image feature vector, and finally obtain a multi-modal task processing model based on visual understanding. Among them, the multi-modal task processing model finally trained by the terminal device includes the first vision model, the second vision model and the trained hybrid expert connector; moreover, the trained hybrid expert connector in the multi-modal task processing model is used to align the first image feature with the second image feature, and the multi-modal task processing model as a whole can be used to process multi-modal data including at least image data to obtain the processing result of the multi-modal task.
[0084] In some embodiments, in the architecture of the multi-modal task processing model trained by the terminal device, the hybrid expert connector can be connected to both the first vision model and the second vision model at the same time. Based on this, the image feature vectors (the first image feature vector and the second image feature vector) respectively output by the first vision model and the second vision model can be automatically input into the hybrid expert connector.
[0085] In some embodiments, the hybrid expert connector can be designed by the terminal device based on the Mixture of Experts (MoE). The terminal device uses the multi-expert network and dynamic routing mechanism of MoE to design the hybrid expert connector as a bridge to align the first vision model and the second vision model. Among them, the multi-expert network is multiple sub-networks (experts) included in the MoE module, and each sub-network is good at processing different features and patterns; the dynamic routing mechanism means that the MoE module dynamically assigns weights according to the input features, so that some experts play a major role in specific features of the input, while other experts are partially or completely blocked. By using the multi-expert network and dynamic routing mechanism of MoE, the terminal device can flexibly adapt to the feature differences between the two vision models during the process of aligning the two vision models, and optimize the combination of global and local information.
[0086] For example, for the first vision model and the second vision model, when they have different feature distributions, especially when there are differences in the expression of low-level details and high-level semantics, the terminal device can use MoE to capture these differences through multiple experts, that is: some experts can focus on aligning low-level features (such as texture, edges), and other experts focus on aligning high-level semantics (such as object category, scene understanding). In addition, the terminal device also uses MoE to automatically assign weights according to the input features through dynamic routing, so that each input is processed by the most suitable expert, thereby achieving efficient alignment.
[0087] For another example, for diverse data, when the image data (including the image training data and the image data in the multi-modal task to be processed) input by the terminal device to two vision models is diverse (for example, different resolutions, color spaces, or noise patterns), since it is difficult to adapt to all samples simultaneously using a traditional fixed and non-dynamic mechanism, the terminal device can utilize the expert mechanism of MoE to dynamically adjust the alignment strategy according to the input, thereby improving the generalization ability for unseen data.
[0088] In the embodiments of the present application, when training a multi-modal task processing model based on visual understanding through a terminal device, first, model training data including at least image training data is obtained. Then, the image training data in the model training data is simultaneously input into a first vision model and a second vision model. Thus, the first vision model learns the image training data and outputs a first image feature vector, and the second vision model learns the image training data and outputs a second image feature vector. Then, the first image feature vector and the second image feature vector are input into a hybrid expert connector to be trained, so as to train the hybrid expert connector to align the first image feature with the second image feature vector, and finally obtain a multi-modal task processing model based on visual understanding. Among them, the multi-modal task processing model finally trained by the terminal device includes a first vision model, a second vision model, and a trained hybrid expert connector. Moreover, the trained hybrid expert connector in the multi-modal task processing model is used to align the first image feature with the second image feature, and the entire multi-modal task processing model can be used to process multi-modal data including at least image data to obtain the processing result of the multi-modal task. In this way, the multi-modal task processing model finally trained in the embodiments of the present application can adapt to the feature distribution differences of at least two vision models through the trained hybrid expert connector, and achieve precise alignment of the image feature vectors output by different vision models.
[0089] Compared with using a single vision model for image feature capture and analysis, the embodiments of the present application can efficiently align multi-source visual features (image feature vectors output by different vision models) through a hybrid expert connector. In this way, the multi-modal task processing model finally trained in the embodiments of the present application can comprehensively capture and analyze the features of the entire image, thereby being able to process multi-modal data more accurately and flexibly, and improving the performance of the model in processing multi-modal tasks.
[0090] In addition, compared with directly aligning the image feature vectors of two actual model outputs through a unified non-dynamic model (which may require a significant increase in the parameter scale to cover all possible input patterns), in the embodiments of the present application, the terminal device uses a hybrid expert connector designed based on the MoE module. Through the division of labor among experts, each expert processes specific tasks, thereby reducing redundant parameters and improving the overall computational efficiency of the model.
[0091] In some embodiments, the terminal device aligns the semantic spaces of two visually models trained in different ways by introducing a hybrid expert module MoE during model training, so as to finally obtain a multi-modal task processing model based on visual understanding.
[0092] Please refer to Figure 2 , Figure 2 For Figure 1 the detailed step flow diagram of step S103 in
[0093] In some embodiments, as Figure 2 shown, in the above step S103, "inputting the first image feature vector and the second image feature vector into the hybrid expert connector to be trained and training the hybrid expert connector" may include step S201 and step S202 shown below.
[0094] Step S201: Input the target image feature vector into the hybrid expert connector to be trained; wherein, the target image feature vector is any one of the first image feature vector and the second image feature vector.
[0095] After the first image feature vector and the second image feature vector, the terminal device inputs any one of them as the target image feature vector into the hybrid expert connector to be trained.
[0096] In some embodiments, the terminal device may input the first image feature vector as the target image feature vector into the hybrid expert connector to be trained, and at the same time, also input the second image feature vector as the target image feature vector into the hybrid expert connector.
[0097] Step S202: Perform an alignment task on the target image feature vector based on the hybrid expert connector to align the first image feature vector and the second image feature vector.
[0098] It should be noted that the alignment task is a control logic designed by the terminal device to guide the hybrid expert connector to align the first image feature vector and the second image feature vector. For example, the alignment task may be a control instruction for controlling the hybrid expert connector to align the first image feature vector and the second image feature vector based on a multi-expert network and a dynamic routing mechanism.
[0099] After the terminal device inputs the first image feature vector and / or the second image feature vector as the target image feature vector into the mixture-of-experts connector, in order to train the mixture-of-experts connector to align the first image feature vector and the second image feature vector, the terminal device controls the mixture-of-experts connector to perform an alignment task on the target image feature vector.
[0100] In some embodiments, the terminal device can guide the mixture-of-experts connector designed based on the MoE module to align the first image feature vector and the second image feature vector by adding a task embedding to the image feature vector.
[0101] Based on this, the step S201 of "inputting the target image feature vector into the mixture-of-experts connector to be trained" can include:
[0102] Adding alignment task parameters to the target image feature vector; wherein the alignment task parameters include contrastive learning alignment task parameters and L2 squared difference alignment task parameters;
[0103] Inputting the target image feature vector added with the contrastive learning alignment task parameters into the mixture-of-experts connector to be trained, and inputting the target image feature vector added with the L2 squared difference alignment task parameters into the mixture-of-experts connector to be trained.
[0104] When the terminal device inputs the target image feature vector into the mixture-of-experts connector, it can first add an alignment task parameter on the target image feature vector to guide the mixture-of-experts connector to perform an alignment task on the target image feature vector. Among them, the alignment task parameter added by the terminal device can be the contrastive learning alignment task parameter for guiding the mixture-of-experts connector to perform contrastive learning alignment on the target image feature vector, or the L2 squared difference alignment task parameter for guiding the mixture-of-experts connector to perform L2 alignment on the target image feature vector.
[0105] After the terminal device adds the contrastive learning alignment task parameters to the target image feature vector, it inputs the target image feature vector added with the contrastive learning alignment task parameters into the mixture-of-experts connector to be trained, and after adding the L2 squared difference alignment task parameters to the target image feature vector, it also inputs the target image feature vector added with the L2 squared difference alignment task parameters into the mixture-of-experts connector to be trained.
[0106] In this case, the step S202 of "performing an alignment task on the target image feature vector based on the mixture-of-experts connector" can include:
[0107] Perform a contrastive learning alignment task on the target image feature vector with the contrastive learning alignment task parameters added based on the hybrid expert connector, and perform an L2 squared difference alignment task on the target image feature vector with the L2 squared difference alignment task parameters added based on the hybrid expert connector.
[0108] After the terminal device inputs the target image feature vector with the contrastive learning alignment task parameters added into the hybrid expert connector, the hybrid expert connector processes the target image feature vector through the dynamic routing mechanism under the guidance of the contrastive learning alignment task parameters to perform the contrastive learning alignment task on the target image feature vector for contrastive learning alignment. In addition, after the terminal device inputs the target image feature vector with the L2 squared difference alignment task parameters added into the hybrid expert connector, the hybrid expert connector processes the target image feature vector through the dynamic routing mechanism under the guidance of the L2 squared difference alignment task parameters to perform the L2 squared difference alignment task on the target image feature vector for L2 alignment.
[0109] In some embodiments, when the terminal device inputs the target image feature vector into the hybrid expert connector, it can also add the contrastive learning alignment task parameters and the L2 squared difference alignment task parameters to the target image feature vector at the same time. Then, the terminal device inputs the target image feature vector with the contrastive learning alignment task parameters and the L2 squared difference alignment task parameters added into the hybrid expert connector to be trained. After that, the terminal device controls the hybrid expert connector to first process the target image feature vector through the dynamic routing mechanism under the guidance of the contrastive learning alignment task parameters for contrastive learning alignment (at this time, the hybrid expert connector ignores the existence of the L2 squared difference alignment task parameters). In addition, the terminal device also controls the hybrid expert connector to process the target image feature vector through the dynamic routing mechanism under the guidance of the L2 squared difference alignment task parameters for L2 alignment (at this time, the hybrid expert connector ignores the existence of the contrastive learning alignment task parameters).
[0110] It should be noted that for the terminal device to train the hybrid expert connector to align the first image feature vector and the second image feature vector, the alignment task parameters added by the terminal device to the target image feature vector (the first image feature vector or the second image feature vector) input into the hybrid expert connector can be a type of task embedding (TaskEmbedding). Among them, the contrastive learning alignment task parameters are contrastive embedding (Contrastive Embedding), which guides the hybrid expert connector designed based on the mixture of experts module MoE to perform contrastive learning alignment, and the L2 squared difference alignment task parameters are L2 embedding (L2 Embedding), which guides the hybrid expert connector to perform L2 alignment.
[0111] In this embodiment, the terminal device trains a mixture-of-experts connector to align the first image feature with the second image feature vector, so that the finally trained multi-modal task processing model can efficiently align the visual features of different visual models. That is, when different visual models can capture different features (for example, one visual model focuses on local details and another visual model focuses on global semantics), the terminal device uses the mixture-of-experts connector designed based on the MoE module through a dynamic routing mechanism to flexibly adapt to various feature distributions, enabling multiple experts to be selectively activated according to the input data, thereby aligning the differences in feature distributions. In this way, compared with using a fixed alignment mechanism to align different-source visual features, the terminal device realizes a more flexible and precise efficient alignment of multi-source visual features through MoE.
[0112] In addition, by using the dynamic routing mechanism of MoE in the terminal device to align the first image feature with the second image feature vector, it can also enable the finally trained multi-modal task processing model to activate the most suitable experts for different input data (such as image data with different resolutions, different color spaces, or different noise patterns), thereby improving the generalization ability of the model to unseen data based on this dynamic adaptability and showing stronger robustness in complex scenarios.
[0113] In some embodiments, the terminal device can introduce a mixture-of-experts module MoE to align two frozen visual models to the semantic representation space of the text, so as to train the final multi-modal task processing model by complementing the advantages of the two visual models. And the terminal device can set the alignment target of the mixture-of-experts connector designed based on MoE by designing a loss function, so as to integrate the high-level information of visual features and text semantics and at the same time retain the low-level information in the visual model.
[0114] Among them, when designing the loss function, the terminal device can use the L2 loss and the contrast loss to jointly optimize the process of aligning the image feature vectors by the mixture-of-experts connector, so as to achieve the alignment of the low-level information and high-level information of the model. Among them, in order to ensure that the trained mixture-of-experts connector performs low-level semantic alignment on the image feature vectors encoded and output by different visual models, the model can be constrained to retain details such as texture, edges, and local features. Based on the L2 loss directly measuring the per-element difference between two features in the feature map, the L2 loss can be used to constrain the model to capture low-level information. For example, the terminal device designs the L2 loss through per-element difference constraints to ensure that the mixture-of-experts connector retains the similarity of a specific model (such as color distribution) of local structures (such as texture, edges, etc.) of low-level features when aligning image feature vectors. In addition, since the core objective of the contrast loss is to align features in the high-level semantic space by pulling closer the feature representations of positive samples (semantically similar) and pulling away those of negative samples (semantically irrelevant), in order to ensure that the trained mixture-of-experts connector performs high-level semantic alignment on the image feature vectors encoded and output by different visual models, the contrast loss can be used to optimize the representation through the similarity and distance relationships between features, so as to ensure that the mixture-of-experts connector captures global or semantic-level concepts when aligning image feature vectors.
[0115] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the steps of the training method for multi-visual models provided by the embodiments of the present application in some other embodiments.
[0116] In some embodiments, the model training data obtained by the terminal device further includes text training data. In this case, as Figure 3 shown, the training method for multi-visual models provided by the embodiments of the present application may further include steps S301 to S303 as follows.
[0117] Step S301: Input the text training data into a language model to obtain text feature vectors.
[0118] When the terminal device aligns two visual models to the semantic representation space of text by introducing the mixture-of-experts module MoE to train a multi-modal task processing model, when obtaining model training data, in addition to obtaining image training data, it is also necessary to obtain text training data for feature description of the image training data, so as to form an image-text pair with the image training data and text training data to train the mixture-of-experts connector and obtain the final multi-modal task processing model.
[0119] It should be noted that the multi-modal task processing model is trained by the terminal device by aligning two visual models to the semantic representation space of the text. In addition to including at least two different visual models (the first visual model and the second visual model) and the mixture of experts connector, it also includes a language model. The language model, like any visual model, is also connected to the mixture of experts connector.
[0120] After the terminal device obtains the model training data, while inputting the image training data in the model training data into the first visual model and the second visual model, it can input the text training data obtained in the model training data into the language model, so that the language model learns and encodes the text training data to output a text feature vector.
[0121] Step S302: Match the alignment feature vector output by the mixture of experts connector when performing the alignment task on the target image feature vector with the text feature vector, and calculate the alignment loss between the alignment feature vector and the text feature vector; wherein, the alignment loss includes a contrast loss and an L2 squared difference loss.
[0122] When the terminal device trains the multi-modal task processing model by aligning two visual models to the semantic representation space of the text by introducing the mixture of experts module MoE, it matches the alignment feature vector output by the mixture of experts connector when performing the alignment task on the target image feature vector with the text feature vector output by the language model, and calculates the alignment loss between the alignment feature vector and the text feature vector. Among them, the alignment loss calculated by the terminal device includes a contrast loss and an L2 squared difference loss.
[0123] In some embodiments, when the target image feature vector is the first image feature vector, after the terminal device inputs the first image feature vector into the mixture of experts connector and controls the mixture of experts connector to perform the alignment task on the first image feature vector based on adding a task embedding (Task Embedding), the terminal device matches the alignment feature output by the mixture of experts connector when performing the alignment task with the text feature vector output by the language model, and calculates the contrast loss and the L2 squared difference loss. In addition, when the target image feature vector is the second image feature vector, after the terminal device inputs the second image feature vector into the mixture of experts connector and controls the mixture of experts connector to perform the alignment task on the second image feature vector based on adding a task embedding (Task Embedding), the terminal device matches the alignment feature output by the mixture of experts connector when performing the alignment task with the text feature vector output by the language model, and calculates the contrast loss and the L2 squared difference loss.
[0124] Step S303: Optimize the mixture-of-experts connector based on the contrastive loss and the L2 squared loss to align the first image feature vector and the second image feature vector.
[0125] After the terminal device calculates the contrastive loss and the L2 squared loss, it jointly optimizes the mixture-of-experts connector to perform the alignment task on the target image feature vector, so as to align the first image feature vector and the second image feature vector, and loops until the trained mixture-of-experts connector is finally obtained, and then a multi-modal task processing model is obtained.
[0126] In this embodiment, the terminal device combines the L2 loss (L2 squared loss) and the contrastive loss to train the mixture-of-experts connector, so that the mixture-of-experts connector simultaneously aligns the low-level and high-level features of the image feature vector (wherein, the L2 loss guides the mixture-of-experts connector to perform low-level feature alignment of the image feature vector, and the contrastive loss guides the mixture-of-experts connector to perform high-level semantic alignment of the image feature vector), achieving the consistency of multi-level features of the image feature vector under a unified framework. In addition, the terminal device uses the L2 loss and the contrastive loss to jointly optimize the two visual models, so that the finally trained multi-modal task processing model retains the respective characteristics of different visual models and eliminates conflicts through the alignment mechanism, thereby enhancing the cross-model collaboration ability of the multi-modal task processing model. That is to say, compared with optimizing a single visual model or performing simple image feature splicing, the multi-modal task processing module trained by the terminal device in combination with the MoE module in this embodiment can provide a better collaboration effect and is suitable for processing a wider range of multi-modal tasks.
[0127] Please refer to Figure 4 , Figure 4 as Figure 3 a schematic diagram of a refined step flow of step S302 in
[0128] In some embodiments, the alignment task performed by the terminal device controlling the mixture-of-experts connector on the target image feature vector includes a contrastive learning alignment task and an L2 squared alignment task. In this case, as Figure 4 shown, the above step S302: Matching the aligned feature vector output by the mixture-of-experts connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the aligned feature vector and the text feature vector, may include the following steps S401 and S402.
[0129] Step S401: Match the first aligned feature vector output by the hybrid expert connector when performing the contrast learning alignment task on the first image feature vector with the text feature vector, and calculate the contrast loss between the first aligned feature vector and the text feature vector.
[0130] When the terminal device aligns two visual models to the semantic representation space of the text by introducing the Mixture of Experts (MoE) module to train the multi-modal task processing model, in the case where the target image feature vector is the first image feature vector, the terminal device inputs the first image feature vector into the hybrid expert connector and controls the hybrid expert connector to perform the contrast learning alignment task on the first image feature vector based on the method of adding Task Embedding. Then, the terminal device matches the first aligned feature vector output by the hybrid expert connector when performing the contrast learning alignment task with the text feature vector output by the language model, and calculates the contrast loss between the first aligned feature vector and the text feature vector.
[0131] Step S402: Match the second aligned feature vector output by the hybrid expert connector when performing the L2 squared difference alignment task on the second image feature vector with the text feature vector, and calculate the L2 squared difference loss between the second aligned feature vector and the text feature vector.
[0132] When the terminal device aligns two visual models to the semantic representation space of the text by introducing the Mixture of Experts (MoE) module to train the multi-modal task processing model, in the case where the target image feature vector is the second image feature vector, the terminal device inputs the second image feature vector into the hybrid expert connector and controls the hybrid expert connector to perform the L2 squared difference alignment task on the second image feature vector based on the method of adding Task Embedding. Then, the terminal device matches the second aligned feature vector output by the hybrid expert connector when performing the L2 squared difference alignment task with the text feature vector output by the language model, and calculates the L2 squared difference loss between the second aligned feature vector and the text feature vector.
[0133] In this way, the terminal device can use the L2 squared difference loss and the contrast loss for joint optimization of the two visual models, so that the finally trained multi-modal task processing model for different visual models not only retains their respective characteristics but also eliminates conflicts through the alignment mechanism, thereby enhancing the cross-model collaboration ability of the multi-modal task processing model.
[0134] In some embodiments, when the terminal device aligns two visual models to the semantic representation space of text by introducing the Mixture of Experts (MoE) module to train the multi-modal task processing model, in the case where the target image feature vector is the first image feature vector, the terminal device can also input the first image feature vector into the Mixture of Experts connector and control the Mixture of Experts connector to perform the L2 squared difference alignment task on the first image feature vector based on adding the Task Embedding. After that, the terminal device matches the aligned feature vector output by the Mixture of Experts connector when performing the L2 squared difference alignment task with the text feature vector output by the language model and calculates the L2 squared difference loss. Moreover, in the case where the target image feature vector is the second image feature vector, the terminal device inputs the second image feature vector into the Mixture of Experts connector and controls the Mixture of Experts connector to perform the contrastive learning alignment task on the second image feature vector based on adding the Task Embedding. After that, the terminal device matches the aligned feature vector output by the Mixture of Experts connector when performing the contrastive learning alignment task with the text feature vector output by the language model and calculates the contrastive loss.
[0135] Please refer to Figure 5 , Figure 5 is Figure 3 another schematic diagram of the refined step process for step S302 in
[0136] In some embodiments, in the case where the alignment task includes the contrastive learning alignment task and the L2 squared difference alignment task, as Figure 5 shown, the above-mentioned step S302: matching the aligned feature vector output by the Mixture of Experts connector when performing the alignment task on the target image feature vector with the text feature vector and calculating the alignment loss between the aligned feature vector and the text feature vector may include steps S501 to S504 as shown below.
[0137] Step S501: Aggregate the first aligned feature vector output by the Mixture of Experts connector when performing the contrastive learning alignment task on the first image feature vector with the third aligned feature vector output by the Mixture of Experts connector when performing the contrastive learning alignment task on the second image feature vector to obtain the first aggregated feature vector.
[0138] When the terminal device aligns two visual models to the semantic representation space of text by introducing the Mixture of Experts (MoE) module to train a multi-modal task processing model, the terminal device controls the MoE connector to perform a contrastive learning alignment task on the first image feature vector based on adding a task embedding (Task Embedding), and obtains a first aligned feature vector output by the MoE connector. Moreover, the terminal device also controls the MoE connector to perform a contrastive learning alignment task on the second image feature vector based on adding a task embedding (Task Embedding), and obtains a third aligned feature vector output by the MoE connector. Then, the terminal device aggregates the first aligned feature vector and the third aligned feature vector to obtain a first aggregated feature vector.
[0139] Step S502: Match the first aggregated feature vector with the text feature vector, and calculate the contrastive loss between the first aggregated feature vector and the text feature vector.
[0140] After obtaining the first aggregated feature vector, the terminal device matches the first aggregated feature vector with the text feature vector output by the language model, and calculates the contrastive loss between the first aggregated feature vector and the text feature vector.
[0141] Step S503: Aggregate the fourth aligned feature vector output by the MoE connector when performing the L2 squared difference alignment task on the first image feature vector, and the second aligned image feature vector output by the MoE connector when performing the L2 squared difference alignment task on the second image feature vector, to obtain a second aggregated feature vector.
[0142] When the terminal device aligns two visual models to the semantic representation space of text by introducing the Mixture of Experts (MoE) module to train a multi-modal task processing model, the terminal device controls the MoE connector to perform an L2 squared difference alignment task on the first image feature vector based on adding a task embedding (Task Embedding), and obtains a fourth aligned feature vector output by the MoE connector. Moreover, the terminal device also controls the MoE connector to perform an L2 squared difference alignment task on the second image feature vector based on adding a task embedding (Task Embedding), and obtains a second aligned feature vector output by the MoE connector. Then, the terminal device aggregates the fourth aligned feature vector and the second aligned feature vector to obtain a second aggregated feature vector.
[0143] Step S504: Match the second aggregated feature vector with the text feature vector, and calculate the L2 squared difference loss between the second aggregated feature vector and the text feature vector.
[0144] After obtaining the second aggregated feature vector, the terminal device matches the second aggregated feature vector with the text feature vector output by the language model, and calculates the L2 squared difference loss between the second aggregated feature vector and the text feature vector.
[0145] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the alignment process of multiple visual models involved in the training method of the multi-visual model provided by the embodiments of the present application in some embodiments.
[0146] As Figure 6 shown, the terminal device aligns two visually trained models (Visual Model 1 and Visual Model 2) in different ways by introducing a Mixture of Experts network MoE (Mixture of Experts Connector) to achieve the purpose of complementary advantages of the two visual models. Specifically, Visual Model 1 is a pure visual model (the first visual model) trained only on images, and Visual Model 2 is a visual model (the second visual model) trained with text guidance. In addition, the Mixture of Experts network MoE consists of multiple experts (such as expert1, expert2, etc.). The terminal device takes image-text pairs (image training data and text training data) as input, where the images are encoded by a pure visual model (such as DINO-V2) and a text-guided visual model (such as CLIP) respectively to generate corresponding feature vectors (the first image feature vector and the second image feature vector).
[0147] To further align the features, the terminal device adds a task embedding (TaskEmbedding) to each encoded feature vector, where the contrastive embedding (Contrastive Embedding) guides the MoE module to perform contrastive learning alignment, and the L2 embedding (L2 Embedding) guides the MoE module to perform L2 alignment.
[0148] The feature vectors after adding the task embedding are respectively input into the MoE module, and after being processed by the MoE module through the dynamic routing mechanism, two contrastive vectors (the first aligned feature vector and the third aligned feature vector) and two L2 vectors (the second pair of feature vectors and the fourth aligned feature vector) are obtained. Subsequently, the terminal device aggregates the contrastive vectors and the L2 vectors respectively. The aggregated vectors (the first aggregated feature vector and the second aggregated feature vector) are matched with the text feature vectors encoded by the language model, and the contrastive loss and the L2 loss are calculated respectively, so as to realize the alignment of the low-level semantic image features (the first image feature vector) and the high-level image features (the second image feature vector) by training the MoE module.
[0149] In this embodiment, the MoE module is introduced through the terminal device and the alignment is optimized by combining the L2 loss and the contrastive loss (the MoE module plays a bridging role in the optimization process of the L2 loss and the contrastive loss by dynamically activating different experts). The L2 loss is used to align the underlying image features to ensure the consistency of the local image information (such as texture, edges, etc.). In addition, the high-level semantic alignment of the image is performed through the contrastive loss. Among them, some experts in the MoE module can focus on capturing the local patterns of these low-level features, such as processing low-frequency and high-frequency feature components. The MoE module can allocate tasks according to the input features based on the dynamic routing mechanism, so that the L2 loss only optimizes those experts responsible for low-level alignment without interfering with other experts. In addition, the contrastive loss focuses on the high-level semantic representation and completes the global semantic features through positive and negative sample pairs. Some other experts in the MoE module can focus on extracting high-level semantic features and align the visual features to a unified semantic space. The dynamic routing selects the appropriate experts according to the high-level feature distribution, making the optimization of the contrastive loss more focused. That is to say, through the terminal device, the MoE module is used to simultaneously activate the experts of low-level and high-level features in the same input. The low-level experts optimize the L2 loss to ensure the consistency of local features, while the high-level experts optimize the contrastive loss to improve the alignment effect of semantic representation. Finally, the output of the MoE module is the result of weighted fusion of multiple experts, taking into account both low-level and high-level information, so that the multi-modal task processing model including the trained MoE module can align the image features output by different visual models more comprehensively.
[0150] Next, a classification comparison experiment is proposed to evaluate the representation quality of the multi-modal task processing model involved in the above embodiments through linear probe classification.
[0151] It should be noted that linear probe is a method for evaluating the feature expression ability of pre-trained models and is widely used in representation learning and multi-modal tasks. By using the features of a certain layer of the pre-trained model as input, a simple linear classifier is trained to complete the downstream classification task. The core idea of linear probe is that if the features extracted by the model have sufficient expression ability, then even using only a linear classifier can achieve high performance in the task. Different from end-to-end fine-tuning, linear probe evaluates the quality of pre-trained features rather than the global optimization ability of the model. Therefore, linear probe provides a fast and efficient way to measure the applicability and generalization ability of features in a specific task, providing an important basis for analyzing and improving pre-trained models.
[0152] In the classification comparison experiment of the multi-modal task processing model based on linear probe, the test data set is: ImageNet1k (the ImageNet-1K data set is a classic benchmark data set in the field of computer vision).
[0153] For the visual model Dino-v2, the last two-layer representations of the frozen Dino-v2 were obtained for aggregation (based on testing various combinations, the aggregation effect of the two-layer representation has been experimentally shown to be the best), and input into a linear classification network for classification. To measure the quality of the representations, only the linear classification layer was trained.
[0154] For the visual model CLIP-ViT, the last two layers of the frozen CLIP-ViT were also obtained for aggregation. Input into a linear classification network for classification. To measure the quality of the representations, only the linear classification layer was trained.
[0155] For the multi-modal task processing model obtained by the training method of the multi-visual model provided in the embodiments of the present application, after completing the alignment training of the mixture-of-experts connector, the obtained contrast vector and L2 vector were connected and input into a linear classification layer for classification. Only the classification layer of the entire network was trained to measure the aligned vectors after aggregation.
[0156] Finally, the experimental results are as Figure 7 shown: By comparing the classification results of the three models on ImageNet1k under the same conditions, it shows that by aligning two visually models trained in different ways for complementary advantages, the quality of the model representations is further outlined, thus proving the feasibility of the training method of the multi-visual model proposed in the embodiments of the present application.
[0157] In the embodiments of the present application, by using the MoE module to align two visually models trained in different ways for complementary advantages to train a multi-modal task processing model, the difference in the feature distributions of the two visually models can be adapted, and precise alignment can be achieved through dynamic activation. Different experts in the MoE module are responsible for low-level feature alignment and high-level semantic alignment, thus realizing multi-task division of labor and avoiding conflicts between models. In addition, compared with the global model, the MoE module realizes diverse alignment functions with fewer parameters, achieving efficient utilization of parameters. And, the MoE module processes diverse inputs through dynamic routing, which can improve the adaptability of the final multi-modal task processing model to new tasks and unseen data, so that the model has stronger generalization ability. Also, by combining the L2 loss and the contrast loss to optimize the alignment objective of the MoE module, the consistency of global and local features is achieved under the weighted fusion of the MoE module, achieving the effect of optimized cooperation.
[0158] Next, various embodiments of the multi-modal task processing method provided in the embodiments of the present application are presented.
[0159] Please refer to Figure 8 , Figure 8 , which is the schematic diagram of the step flow of the multi-modal task processing method provided in the embodiments of the present application in some embodiments.
[0160] In some embodiments, asFigure 8 As shown in Figure 8 , the multi-modal task processing method provided by the embodiments of the present application may include steps S801 and S802 as shown below.
[0161] Step S801: Obtain multi-modal data to be processed for a multi-modal task; wherein, the multi-modal data includes at least image data.
[0162] It should be noted that multi-modal tasks include, but are not limited to, vision-language alignment, cross-modal generation, audio-visual fusion, etc. Among them, the data to be processed for each multi-modal task includes data of different modalities, and there are significant differences in the data distribution and feature expression forms of each modality.
[0163] The terminal device can execute any one or more of the multi-modal tasks such as vision-language alignment, cross-modal generation, audio-visual fusion, etc. based on the multi-modal task processing model finally trained in the various embodiments of the above training method. And, the terminal device first obtains the multi-modal data to be processed for the multi-modal task to be executed currently. Among them, the multi-modal data obtained by the terminal device includes at least image data.
[0164] Exemplarily, when the terminal device executes the multi-modal task of image classification, it first obtains the image data of the image modality, and uses the obtained image data as the multi-modal data to be processed for the multi-modal task to be executed currently.
[0165] Step S802: Input the image data into a preset multi-modal task processing model based on visual understanding to obtain the processing result of the multi-modal task; wherein, the first visual model and the second visual model in the multi-modal task processing model respectively learn the image data to obtain corresponding first image feature vectors and second image feature vectors; the mixture-of-experts connector in the multi-modal task processing model aligns the first image feature vector and the second image feature vector.
[0166] The terminal device takes the obtained image data as input, and inputs the image data into the multi-modal task processing model trained in advance by the above training method, so that the multi-modal task processing model processes the image data based on visual understanding and outputs the processing result of the multi-modal task. Among them, when the multi-modal task processing model processes the image data based on visual understanding, through the first visual model and the second visual model, it respectively learns and encodes the image data to output the corresponding first image feature vector and second image feature vector of the image data. Then, both the first image feature vector and the second image feature vector are input into the mixture-of-experts connector, so that the mixture-of-experts connector aligns the first image feature vector and the second image feature vector.
[0167] It should be noted that the alignment of the first image feature vector and the second image feature vector by the mixture-of-experts connector is consistent with the process of aligning the first image feature vector and the second image feature vector by the mixture-of-experts connector described in the foregoing embodiments of the training method. Therefore, the same content will not be elaborated herein.
[0168] In the embodiments of the present application, a terminal device first obtains multi-modal data to be processed for a multi-modal task to be executed. Among them, the multi-modal data obtained by the terminal device at least includes image data. Then, the multi-modal task processing model processes the image data based on visual understanding and outputs a processing result of the multi-modal task. Moreover, when the multi-modal task processing model processes the image data based on visual understanding, it respectively learns and encodes the image data through a first vision model and a second vision model to output a first image feature vector and a second image feature vector corresponding to the image data. After that, both the first image feature vector and the second image feature vector are input into the mixture-of-experts connector, so as to align the first image feature vector and the second image feature vector through the mixture-of-experts connector.
[0169] In this way, the embodiments of the present application adopt a multi-modal task processing model to process multi-modal tasks, and can adapt to the feature distribution differences of at least two vision models through the mixture-of-experts connector, so as to accurately align the image feature vectors output by different vision models. Compared with using a single vision model to capture and analyze image features, the multi-modal task processing model in the embodiments of the present application can efficiently align multi-source vision features through the mixture-of-experts connector, that is, capture and analyze local detail features of an image based on one vision model, capture and analyze global semantic features of the image based on another vision model, and then align the image feature vectors output by the two vision models through the mixture-of-experts connector. Therefore, when the embodiments of the present application process multi-modal tasks based on visual understanding using the multi-modal task processing model, they can comprehensively capture and analyze the features of the entire image, so as to process multi-modal data more accurately and flexibly, and further improve the performance of the model in processing multi-modal tasks.
[0170] Please refer to Figure 9 , the embodiments of the present application further provide a training device for multiple vision models, which can implement the above-mentioned training method for multiple vision models. The device includes a first acquisition module 901, a single vision model processing module 902, and a multi-level image feature alignment module 903. Among them,
[0171] The first acquisition module 901 is configured to acquire model training data; among them, the model training data at least includes image training data;
[0172] The single vision model processing module 902 is configured to input the image training data into a first vision model and a second vision model, and obtain a first image feature vector output by the first vision model and a second image feature vector output by the second vision model;
[0173] The multi-level image feature alignment module 903 is configured to input the first image feature vector and the second image feature vector into a to-be-trained mixture-of-experts connector, train the mixture-of-experts connector, and obtain a multi-modal task processing model based on visual understanding;
[0174] Wherein, the multi-modal task processing model includes the first vision model, the second vision model, and the trained mixture-of-experts connector; the trained mixture-of-experts connector is configured to perform alignment of the first image feature and the second image feature; the multi-modal task processing model is configured to process multi-modal data including at least image data to obtain a processing result of a multi-modal task.
[0175] In some embodiments, the multi-level image feature alignment module 903 is further configured to input a target image feature vector into the to-be-trained mixture-of-experts connector; wherein, the target image feature vector is any one of the first image feature vector and the second image feature vector; and, based on the mixture-of-experts connector, perform an alignment task on the target image feature vector to perform alignment between the first image feature vector and the second image feature vector.
[0176] In some embodiments, the multi-level image feature alignment module 903 is further configured to add alignment task parameters to the target image feature vector; wherein, the alignment task parameters include contrastive learning alignment task parameters and L2 squared difference alignment task parameters; input the target image feature vector added with the contrastive learning alignment task parameters into the to-be-trained mixture-of-experts connector, and input the target image feature vector added with the L2 squared difference alignment task parameters into the to-be-trained mixture-of-experts connector; and, based on the mixture-of-experts connector, perform a contrastive learning alignment task on the target image feature vector added with the contrastive learning alignment task parameters, and perform an L2 squared difference alignment task on the target image feature vector added with the L2 squared difference alignment task parameters.
[0177] In some embodiments, the model training data further includes text training data; the multi-level image feature alignment module 903 is further configured to input the text training data into a language model to obtain text feature vectors; match the alignment feature vectors output by the hybrid expert connector performing alignment tasks on the target image feature vectors with the text feature vectors, and calculate the alignment loss between the alignment feature vectors and the text feature vectors; wherein, the alignment loss includes a contrast loss and an L2 squared difference loss; and, optimize the hybrid expert connector for aligning the first image feature vector and the second image feature vector based on the contrast loss and the L2 squared difference loss.
[0178] In some embodiments, the alignment tasks include a contrast learning alignment task and an L2 squared difference alignment task; the multi-level image feature alignment module 903 is further configured to match the first alignment feature vectors output by the hybrid expert connector performing the contrast learning alignment task on the first image feature vectors with the text feature vectors, and calculate the contrast loss between the first alignment feature vectors and the text feature vectors; and, match the second alignment feature vectors output by the hybrid expert connector performing the L2 squared difference alignment task on the second image feature vectors with the text feature vectors, and calculate the L2 squared difference loss between the second alignment feature vectors and the text feature vectors.
[0179] In some embodiments, the alignment tasks include a contrast learning alignment task and an L2 squared difference alignment task; the multi-level image feature alignment module 903 is further configured to aggregate the first alignment feature vectors output by the hybrid expert connector performing the contrast learning alignment task on the first image feature vectors and the third alignment feature vectors output by the hybrid expert connector performing the contrast learning alignment task on the second image feature vectors to obtain a first aggregated feature vector; match the first aggregated feature vector with the text feature vectors, and calculate the contrast loss between the first aggregated feature vector and the text feature vectors; aggregate the fourth alignment feature vectors output by the hybrid expert connector performing the L2 squared difference alignment task on the first image feature vectors and the second aligned image feature vectors output by the hybrid expert connector performing the L2 squared difference alignment task on the second image feature vectors to obtain a second aggregated feature vector; and, match the second aggregated feature vector with the text feature vectors, and calculate the L2 squared difference loss between the second aggregated feature vector and the text feature vectors.
[0180] The specific implementation manners of the training device for the multi-vision model provided by the embodiments of the present application are basically the same as the specific embodiments of the above-mentioned training method for the multi-vision model, and will not be elaborated herein.
[0181] Please refer to Figure 10 , an embodiment of the present application further provides a multimodal task processing device, which can implement the above multimodal task processing method. The device includes a second acquisition module 1001 and a multimodal task processing module 1002. Among them,
[0182] The second acquisition module 1001 is used to acquire multimodal data to be processed for the multimodal task; among them, the multimodal data at least includes image data;
[0183] The multimodal task processing module 1002 is used to input the image data into a preset multimodal task processing model based on visual understanding to obtain the processing result of the multimodal task;
[0184] Among them, the first visual model and the second visual model in the multimodal task processing model respectively learn the image data to obtain corresponding first image feature vectors and second image feature vectors; the mixture-of-experts connector in the multimodal task processing model aligns the first image feature vector and the second image feature vector.
[0185] The specific implementation manner of the multimodal task processing device provided by the embodiment of the present application is basically the same as the specific embodiment of the above multimodal task processing method, and will not be described in detail here.
[0186] An embodiment of the present application further provides a multimodal task processing device. The multimodal task processing device provided by the embodiment of the present application can be a computer device (or an electronic device). Specifically, the computer device can be an intelligent robot, a terminal device equipped with an embodied intelligent system, or a computer device such as a smart phone, a tablet computer, a notebook computer, or a desktop computer. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above model training method for multimodal learning is implemented. In some embodiments, the multimodal task processing device can also be any intelligent terminal such as an in-vehicle computer, a portable PC, and a wearable device.
[0187] Please refer to Figure 11 , Figure 11 illustrates the hardware structure of a computer device in an embodiment. The computer device includes:
[0188] A processor 1101, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0189] The memory 1102 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1102 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1102, and the processor 1101 is called to execute the training method of the multi-vision model in the embodiments of this application;
[0190] The input / output interface 1103 is used to implement information input and output;
[0191] The communication interface 1104 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0192] The bus 1105 transmits information between the various components of the device (such as the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104);
[0193] Among them, the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 achieve communication connections with each other inside the device through the bus 1105.
[0194] The embodiments of this application also provide a vehicle, which is configured with a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned training method of the multi-vision model is implemented.
[0195] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned training method of the multi-vision model is implemented.
[0196] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0197] The embodiments of the present application also provide a computer program product, including a computer program, and the steps implemented when the computer program is executed by a processor are substantially the same as the specific embodiments of the above-mentioned multi-vision model training method, and will not be elaborated herein.
[0198] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0199] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine some steps, or different steps.
[0200] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0202] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0203] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0204] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0205] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0206] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0207] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0208] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of the rights of the embodiments of this application.
Claims
1. A method for training a multi-vision model, characterized in that: The method comprises: Acquire model training data; wherein the model training data at least includes image training data; Inputting the image training data into a first visual model and a second visual model to obtain a first image feature vector output by the first visual model and a second image feature vector output by the second visual model; Inputting the first image feature vector and the second image feature vector into a hybrid expert connector to be trained, training the hybrid expert connector, and obtaining a multimodal task processing model based on visual understanding; Among them, the multimodal task processing model includes the first visual model, the second visual model and a trained hybrid expert connector; the trained hybrid expert connector is used to align the first image feature with the second image feature; the multimodal task processing model is used to process multimodal data including at least image data to obtain a processing result of a multimodal task.
2. The method according to claim 1, characterized in that The step of inputting the first image feature vector and the second image feature vector into a hybrid expert connector to be trained, and training the hybrid expert connector, comprises: Inputting a target image feature vector into a hybrid expert connector to be trained; wherein the target image feature vector is any one of the first image feature vector and the second image feature vector; An alignment task is performed on the target image feature vector based on the hybrid expert connector to align the first image feature vector with the second image feature vector.
3. The method according to claim 2, characterized in that The step of inputting the target image feature vector into the hybrid expert connector to be trained comprises: Adding alignment task parameters to the target image feature vector; wherein the alignment task parameters include contrastive learning alignment task parameters and L2 square difference alignment task parameters; Inputting the target image feature vector to which the contrast learning alignment task parameters are added into the hybrid expert connector to be trained, and inputting the target image feature vector to which the L2 square difference alignment task parameters are added into the hybrid expert connector to be trained; The performing an alignment task on the target image feature vector based on the hybrid expert connector includes: Based on the hybrid expert connector, a contrastive learning alignment task is performed on the target image feature vector to which the contrastive learning alignment task parameters are added, and based on the hybrid expert connector, an L2 square difference alignment task is performed on the target image feature vector to which the L2 square difference alignment task parameters are added.
4. The method according to claim 2, characterized in that: The model training data also includes text training data; the method also includes: Inputting the text training data into a language model to obtain a text feature vector; Matching the aligned feature vector output by the hybrid expert connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the aligned feature vector and the text feature vector; wherein the alignment loss includes contrast loss and L2 square difference loss; The hybrid expert connector is optimized based on the contrast loss and the L2 square difference loss to align the first image feature vector with the second image feature vector.
5. The method according to claim 4, characterized in that The alignment task includes a contrastive learning alignment task and an L2 square difference alignment task; The step of matching the aligned feature vector output by the hybrid expert connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the aligned feature vector and the text feature vector, includes: Matching a first aligned feature vector output by the hybrid expert connector performing the contrastive learning alignment task on the first image feature vector with the text feature vector, and calculating a contrast loss between the first aligned feature vector and the text feature vector; The second aligned feature vector output by the hybrid expert connector performing the L2 square difference alignment task on the second image feature vector is matched with the text feature vector, and the L2 square difference loss between the second aligned feature vector and the text feature vector is calculated.
6. The method according to claim 4, characterized in that The alignment task includes a contrastive learning alignment task and an L2 square difference alignment task; The step of matching the aligned feature vector output by the hybrid expert connector performing the alignment task on the target image feature vector with the text feature vector, and calculating the alignment loss between the aligned feature vector and the text feature vector, includes: Aggregate a first aligned feature vector output by the hybrid expert connector performing the contrastive learning alignment task on the first image feature vector and a third aligned feature vector output by the hybrid expert connector performing the contrastive learning alignment task on the second image feature vector to obtain a first aggregated feature vector; Matching the first aggregated feature vector with the text feature vector, and calculating a contrast loss between the first aggregated feature vector and the text feature vector; Aggregate a fourth aligned feature vector output by the hybrid expert connector performing the L2 square difference alignment task on the first image feature vector with a second aligned image feature vector output by the hybrid expert connector performing the L2 square difference alignment task on the second image feature vector to obtain a second aggregated feature vector; The second aggregated feature vector is matched with the text feature vector, and an L2 squared difference loss between the second aggregated feature vector and the text feature vector is calculated.
7. A multimodal task processing method, characterized in that: The method comprises: Acquire multimodal data to be processed by a multimodal task; wherein the multimodal data at least includes image data; Inputting the image data into a preset multimodal task processing model based on visual understanding to obtain a processing result of the multimodal task; Among them, the first visual model and the second visual model in the multimodal task processing model respectively learn the image data to obtain the corresponding first image feature vector and second image feature vector; the hybrid expert connector in the multimodal task processing model aligns the first image feature vector with the second image feature vector.
8. A multimodal task processing device, characterized in that: The multimodal task processing device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the training method of the multi-vision model described in any one of claims 1 to 6, and / or implements the multimodal task processing method described in claim 7.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the multi-vision model training method described in any one of claims 1 to 6, and / or implements the multimodal task processing method described in claim 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the multi-vision model training method described in any one of claims 1 to 6, and / or implements the multimodal task processing method described in claim 7.
Citation Information
Cited By
Multi-modal video sequence segmentation method based on text-guided hybrid expert mechanism
CN121725393A