Multi-modal model training method, model, image recognition method and related product

Through the bidirectional multimodal spatial consistency prompt adjustment (BMSPT) method, the conceptual understanding limitations of multimodal models in dynamic and open worlds are solved. Through cross-modal mapping and alignment loss optimization, the adaptability and accuracy of the model are enhanced and the user experience is improved.

CN120494125APending Publication Date: 2025-08-15THE HONG KONG UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510482755.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2025-04-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing multimodal models have limitations in conceptual understanding in dynamic and open worlds, resulting in reduced model performance and user trust, mainly due to the limitations of one-way cross-modal semantic alignment.

Method used

Bidirectional multimodal spatial consistency prompt adjustment (BMSPT) method is used to generate multimodal prompt features, cross-modal mapping and reverse mapping are performed, and cross-modal mapping and reverse cross-modal mapping are optimized by combining modal alignment loss and content alignment loss to enhance the model's understanding ability and dynamic interaction between modals.

Benefits of technology

It improves the adaptability and accuracy of the multimodal model, improves user experience and satisfaction, and promotes dynamic and fine-grained alignment between text and image modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494125A_ABST
    Figure CN120494125A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal model training method, a model, an image recognition method and a related product, and relates to the technical field of machine learning, and the multi-modal model training method comprises the steps: generating a multi-modal prompt feature based on multi-modal sample data, carrying out the cross-modal mapping of the multi-modal prompt feature, and obtaining target modal prediction information; determining modal alignment loss through the target modal prediction information and the corresponding sample data; content alignment loss is determined through reverse mapping modal information and corresponding sample data, the reverse mapping modal information is generated by target modal prediction information through reverse cross-modal mapping, and the reverse cross-modal mapping is an inverse process of cross-modal mapping; and optimizing learnable parameters involved in cross-modal mapping, reverse cross-modal mapping and prompt information through modal alignment loss and content alignment loss so as to complete training. According to the embodiment of the invention, the performance of the multi-modal model and the use experience of a user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a multimodal model training method, model, image recognition method and related products. Background Art

[0002] With the development of machine learning technology, multimodal models have been widely used. Pre-trained visual language models (VLMs) such as CLIP (Contrastive Language–Image Pre-training) have demonstrated impressive zero-shot learning capabilities across a wide range of downstream tasks. Thanks to multimodal cue adjustment, machines' ability to understand and interact with multimodal content has been significantly improved.

[0003] Multimodal models in related technologies mainly focus on unidirectional cross-modal semantic alignment, which has limitations in concept understanding in dynamic and open worlds, such as visual concept understanding, resulting in degraded model performance and reduced user trust in the model. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to provide a multimodal model training method, model, image recognition method and related products, aiming to improve model performance and user experience.

[0005] To achieve the above objectives, an embodiment of the present application provides a method for training a multimodal model, comprising the following steps:

[0006] Generating a multimodal prompt feature based on multimodal sample data, wherein the multimodal sample data includes visual sample data, text sample data, and prompt information;

[0007] Performing cross-modal mapping on the multimodal prompt features to obtain target modality prediction information;

[0008] Determining a modality alignment loss using the target modality prediction information and sample data corresponding to the target modality;

[0009] Determining a content alignment loss by reverse mapping modality information and sample data corresponding to the reverse mapping modality, wherein the reverse mapping modality information is generated by the target modality prediction information through reverse cross-modal mapping, and the reverse cross-modal mapping is an inverse process of the cross-modal mapping;

[0010] The cross-modal mapping, the reverse cross-modal mapping, and the learnable parameters involved in the prompt information are optimized by using the modality alignment loss and the content alignment loss to complete the training.

[0011] In some embodiments, the multimodal sample data includes visual sample data, the prompt information includes spatial visual prompt information, and before generating the multimodal prompt feature based on the multimodal sample data, the method further includes:

[0012] Adding the spatial visual cue information to each channel of the spatial dimension of visual information to obtain the visual sample data;

[0013] The spatial visual prompt information is aligned with the spatial dimension of the visual information.

[0014] In some embodiments, the multimodal prompt feature includes a visual prompt feature and a textual prompt feature, and performing cross-modal mapping on the multimodal prompt feature to obtain target modality prediction information includes:

[0015] Mapping the visual cue features to a visual-to-text space to obtain text prediction information;

[0016] The text prompt features are mapped from text to visual space to obtain visual prediction information.

[0017] In some embodiments, the multimodal sample data includes visual sample data and text sample data, the target modality prediction information includes text prediction information and visual prediction information, and determining the modality alignment loss using the target modality prediction information and the sample data corresponding to the target modality includes:

[0018] Determining an alignment loss of visual-to-textual space mapping based on the text sample data and the text prediction information, wherein the alignment loss of the visual-to-textual space mapping aims to minimize a bulldozer distance between the text sample data and the text prediction information so that a distribution of the text sample data is aligned with a distribution of the text prediction information;

[0019] An alignment loss for text-to-visual space mapping is determined based on the visual sample data and the visual prediction information, wherein the alignment loss for text-to-visual space mapping aims to minimize a bulldozer distance between the visual sample data and the visual prediction information so that a distribution of the visual sample data is aligned with a distribution of the visual prediction information.

[0020] In some embodiments, the multimodal sample data includes visual sample data and text sample data, the reverse mapping modality information includes reverse conversion of visual information and reverse conversion of text information, and determining the content alignment loss using the reverse mapping modality information and the sample data corresponding to the reverse mapping modality includes:

[0021] determining a visual content alignment loss based on the inversely transformed visual information and the visual sample data, wherein the visual content alignment loss aims to ensure that the inversely transformed visual information is within a predetermined neighborhood of desired visual information, the desired visual information being used to represent any visual image belonging to the same category as the visual sample data;

[0022] A text content alignment loss is determined by using the inversely converted text information and the text sample data, wherein the text content alignment loss aims to minimize an L1 norm distance between the inversely converted text information and the text sample data.

[0023] In some embodiments, optimizing the cross-modal mapping, the reverse cross-modal mapping, and the learnable parameters involved in the prompt information by using the modality alignment loss and the content alignment loss includes:

[0024] The learnable parameters involved in the visual-to-textual space mapping, text-to-visual space mapping, spatial visual cue information and textual cue information are optimized through the alignment loss of visual-to-textual space mapping, alignment loss of text-to-visual space mapping, visual content alignment loss and textual content alignment loss.

[0025] To achieve the above objectives, another aspect of the present application provides a multimodal model, including:

[0026] an image encoding module, configured to convert input visual sample data into visual features, wherein the visual sample data includes spatial visual cue information, and the visual features include visual cue features;

[0027] A visual-to-linguistic space mapping module for mapping the visual cue features to a text embedding space;

[0028] A text encoding module, configured to convert input text sample data into text features, wherein the text sample data includes text prompt information, and the text features include text prompt features;

[0029] A language to visual space mapping module, configured to map the text prompt features to a visual feature space;

[0030] a contrastive learning module, configured to project the feature vectors output by the image encoding module and the text encoding module into a common embedding space, and bring matching image and text pairs closer together in the embedding space, while pushing mismatched image and text pairs further apart in the embedding space;

[0031] The learnable parameters in the visual-to-language space mapping module, the language-to-visual space mapping module, the spatial visual prompt information and the text prompt information are determined by the multimodal model training method of the above-mentioned embodiment.

[0032] To achieve the above objectives, another aspect of the embodiments of the present application provides an image recognition method, comprising the following steps:

[0033] Acquire a visual image to be recognized;

[0034] Inputting the visual image to be recognized into a visual language model to obtain a recognition result of the visual image;

[0035] outputting a recognition result of the visual image;

[0036] The visual language model is obtained by training using the multimodal model training method of the above-mentioned embodiment.

[0037] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the multimodal model training method of the above-mentioned embodiment or the image recognition method of the above-mentioned embodiment.

[0038] To achieve the above-mentioned objectives, another aspect of an embodiment of the present application proposes a computer program product. When the computer program product is run in an electronic device, the electronic device executes the multimodal model training method of the above-mentioned embodiment or the image recognition method of the above-mentioned embodiment.

[0039] The embodiments of the present application include at least the following beneficial effects:

[0040] The present application provides a training method, model, image recognition method and related products of a multimodal model. In an embodiment of the training method of the multimodal model of the present application, first, a multimodal prompt feature is generated based on multimodal sample data, wherein the multimodal sample data includes prompt information; then, the multimodal prompt feature is cross-modally mapped to obtain target modality prediction information; the modality alignment loss is determined by the target modality prediction information and the sample data corresponding to the target modality; the content alignment loss is determined by the reverse mapping modality information and the sample data corresponding to the reverse mapping modality, wherein the reverse mapping modality information is generated by the target modality prediction information through reverse cross-modal mapping, and the reverse cross-modal mapping is the inverse process of cross-modal mapping; the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping and prompt information are optimized through the modality alignment loss and content alignment loss to complete the training. In the implementation mode of the present application, adding prompt information to the multimodal sample data is helpful to enhance the model's understanding ability, and the use of cross-modal mapping can promote two-way dynamic interaction between modalities. The modality alignment loss and content alignment loss can clearly guide the prompt learning to obtain a representation that is both semantically coherent and information complete during the conversion, thereby improving the performance of the multimodal model and the user experience.

[0041] It is understandable that the beneficial effects of the model, image recognition method, electronic device, and computer program product disclosed in this application are the same as the beneficial effects of the training method of the multimodal model, and will not be repeated here.

[0042] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0044] Figure 1 is a flowchart of a multimodal model training method provided in some embodiments of the present application;

[0045] Figure 2 is another flow chart of a multimodal model training method provided in some embodiments of the present application;

[0046] Figure 3 is another flow chart of a multimodal model training method provided in some embodiments of the present application;

[0047] Figure 4 Schematic diagram of a training framework for a multimodal model provided in some embodiments of the present application;

[0048] Figure 5is a flowchart of an image recognition method provided by some embodiments of the present application;

[0049] Figure 6 Schematic diagram of experimental results of the visual performance of the bidirectional BMSPT model and the half BMSPT model provided in some embodiments of the present application;

[0050] Figure 7 Schematic diagram of experimental results of the diffusion model migration performance of the bidirectional BMSPT model provided by some embodiments of the present application;

[0051] Figure 8 This is a schematic block diagram of modules of a multimodal model training device provided in some embodiments of the present application;

[0052] Figure 9 is a schematic block diagram of modules of an image recognition device provided by some embodiments of the present application;

[0053] Figure 10 is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application;

[0054] Figure 11 This is a schematic diagram of the hardware structure of another electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the reference to "embodiment" in this article means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0056] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0057] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0059] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0060] In order to make the inventive concept of the present application easy to understand, before explaining the embodiments of the present application in detail, the English abbreviations (terms) / related concepts involved in the embodiments of the present application are first explained. The English abbreviations (terms) / related concepts involved in the embodiments of the present application are subject to the following explanations.

[0061] CLIP: CLIP stands for Contrastive Language-Image Pre-training (CLIP). It is a multimodal pre-training model in machine learning that aims to associate image content with natural language. The core idea of the CLIP model is to learn to map image content to a textual representation space by training it with a large amount of images and their associated textual descriptions.

[0062] Generalization: Machine learning models often perform well on training data, but the key is to see how well they perform on unseen data, a skill called generalization. Good generalization means that the model performs well not only on the training data but also on test data or other new datasets.

[0063] With the development of machine learning technology, multimodal models have been widely used. For example, pre-trained vision-language models like CLIP have demonstrated impressive zero-shot learning capabilities across a wide range of downstream tasks. Thanks to multimodal cue adaptation, machines' ability to understand and interact with multimodal content has been significantly improved.

[0064] Pre-trained vision-language models, such as CLIP, have demonstrated remarkable zero-shot performance on a variety of downstream tasks by encoding large-scale datasets into a unified representation space. During inference, CLIP traditionally relies on hand-crafted textual cues (e.g., "a photo of [class]") to align the output category with pre-trained knowledge. However, this reliance on static cues limits the flexibility and adaptability of the system, as they are highly sensitive to semantic changes. Recent advances in natural language processing (NLP) have inspired the development of dynamic, learnable cues to maintain semantic integrity between sentences and labels. These learnable cues allow pre-trained models to dynamically adapt to new tasks without retraining. Visual cue tuning (VPT) has gained prominence in computer vision and multimodal learning tasks by applying visual cues to visual transformers and adjusting very few learnable parameters. For example, by directly introducing task-specific parameters into the input sentence of transformer layers, the VPT strategy can fully exploit the modeling adaptability potential of the image encoder. However, this local attention-based approach may limit the natural expressiveness of the visual modality. For example, if a picture of a zebra is input, it is very likely to classify the zebra as a hyena due to the characteristics of local attention.

[0065] To address these issues to some extent, a new trend has emerged in the field of prompt tuning: combining textual prompt tuning (TPT) and visual prompt tuning (VPT). The goal is to synergistically enhance multimodal prompts across textual and visual domains. Methods combining textual and visual prompt tuning often employ a one-way alignment strategy, using textual descriptions to guide the visual recognition process. Extensive experiments have verified that this approach presents communication challenges during inter-modal alignment. For example, the process of transferring features from textual data representing a "dog" into visual space is fraught with uncertainty. It is unclear whether the data has maintained its integrity, been undistorted, or been correctly interpreted by the visual system. Due to distortions and misinterpretations, a "dog" represented in textual data may be mistakenly identified as a "bear."

[0066] In view of this, the present application proposes a training method, model, image recognition method and related products of a multimodal model. The scheme first generates multimodal prompt features based on multimodal sample data, wherein the multimodal sample data includes prompt information; then, the multimodal prompt features are cross-modally mapped to obtain target modality prediction information; the modality alignment loss is determined by the target modality prediction information and the sample data corresponding to the target modality; the content alignment loss is determined by the reverse mapping modality information and the sample data corresponding to the reverse mapping modality, wherein the reverse mapping modality information is generated by the target modality prediction information through reverse cross-modal mapping, and the reverse cross-modal mapping is the inverse process of cross-modal mapping; the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping and prompt information are optimized through modality alignment loss and content alignment loss to complete the training. In the implementation mode of the present application, adding prompt information to the multimodal sample data is helpful to enhance the model's understanding ability, and the use of cross-modal mapping can promote two-way dynamic interaction between modalities. The modal alignment loss and content alignment loss can clearly guide the prompt learning to obtain a representation that is both semantically coherent and information complete during the conversion, thereby improving the adaptability and accuracy of the multimodal model, thereby improving the user satisfaction when using products based on the multimodal model.

[0067] The multimodal model proposed in this application, which can be called Bidirectional Multimodal Spatial Consistency Prompt Adjustment (BMSPT), is a new model that can promote dynamic and fine-grained alignment between text and image modalities.

[0068] The multimodal model training method and image recognition method provided in the embodiments of the present application can be applied to the electronic device provided in the embodiments of the present application, wherein the electronic device can be a terminal or a server.

[0069] The terminal may be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto.

[0070] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, as well as big data and artificial intelligence platforms.

[0071] It should be understood that the training proposed in this application can be understood as training in a broad sense, that is, including the pre-training process of traditional machine learning models and / or the training process for specific fields. In other words, the training method of the multimodal model provided in this application can be the pre-training of the multimodal model for learning general feature representations on a large-scale dataset, or it can be the training of the multimodal model for task-specific feature representations on a dataset of a specific task.

[0072] The multimodal model training method, model, image recognition method, device and product disclosed in this application have a wide range of application fields, such as image retrieval, text generation and image generation.

[0073] Specifically, the model (BMSPT) obtained using the multimodal model training method disclosed in this application can quickly and accurately retrieve pictures that meet the description from a large image database based on the natural language query provided by the user (such as "a dog on the beach").

[0074] In specific application scenarios, such as e-commerce, consumers can search for products by entering a text description, such as "red dress," to quickly find the corresponding product image, enhancing the shopping experience. It should be understood that by adding a voice-to-text plug-in, BMSPT can also enable product search via voice. When recommending similar products, BMSPT can automatically recommend products of similar styles based on the images viewed by the user, increasing sales opportunities. In inventory management, merchants can also use text descriptions to quickly locate images of inventory items, improving management efficiency.

[0075] In digital asset management application scenarios, for example, BMSPT can help editors and designers in large media companies efficiently manage and retrieve massive image libraries, saving time and costs; on the other hand, through text descriptions, BMSPT can assist in quickly screening and confirming the copyright ownership of images to avoid infringement risks; and planners can also use BMSPT to quickly find relevant images based on thematic requirements, thereby improving content production efficiency.

[0076] In content review application scenarios, for example, users can use BMSPT to quickly filter out images that do not meet standards based on text descriptions such as "violence" and "pornography" to ensure the security of platform content; further, description tags can be automatically generated for reviewed images to facilitate subsequent management and retrieval; BMSPT can also combine real-time image streams to quickly identify and filter illegal content and maintain the network environment.

[0077] In the field of text generation, BMSPT can combine image input and output relevant text descriptions to achieve image-to-text conversion.

[0078] Specifically, BMSPT can automatically generate accurate and consistent descriptions for large-scale image datasets, reducing the cost of manual labeling; on the other hand, it can enhance the semantic information of images through text descriptions and improve the model training effect; in addition, BMSPT can also realize cross-modal retrieval, that is, realize bidirectional retrieval of images and text, and improve information retrieval efficiency.

[0079] In terms of barrier-free assistance, BMSPT can provide visually impaired people with text descriptions of image content to help them "see" the world; combined with speech synthesis technology, BMSPT can convert image descriptions into speech, further facilitating the use of visually impaired people; in special education, BMSPT can use image-to-text conversion to help students with special needs better understand teaching content.

[0080] In the field of image generation, BMSPT can achieve generation from text to image by reversely guiding the generation model (such as DALL-E).

[0081] Specifically, BMSPT can quickly generate concept sketches based on the text description of the designer or client, shortening the design cycle; it can also generate images of different styles by adjusting the text description, helping designers explore more creative possibilities. Furthermore, BMSPT can generate unique image designs based on the user's personalized needs to meet personalized market demands.

[0082] In advertising applications, BMSPT can quickly generate advertising materials that are consistent with brand concepts, such as posters and promotional images, thereby improving advertising production efficiency; by generating images that are consistent with the brand image, it can enhance brand awareness and influence; BMSPT can also generate advertising images of different styles, conduct market testing, and optimize advertising strategies.

[0083] It should be understood that the above only lists some of the application scenarios of this application. In actual production, the multimodal model training method, model, image recognition method, equipment and products provided by this application will have wider applications.

[0084] The following describes in detail the implementation steps of a multimodal model training method provided by an embodiment of the present application in conjunction with the accompanying drawings.

[0085] Please refer to Figure 1 , Figure 1 A flowchart of a method for training a multimodal model provided for some embodiments of the present application; it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0086] The method of the embodiment of the present application includes the following steps:

[0087] Step 101: generating a multimodal prompt feature based on multimodal sample data, wherein the multimodal sample data includes visual sample data, text sample data, and prompt information;

[0088] Step 102: Perform cross-modal mapping on the multimodal prompt features to obtain target modality prediction information;

[0089] Step 103: determining a modality alignment loss using the target modality prediction information and sample data corresponding to the target modality;

[0090] Step 104: determining a content alignment loss using reverse-mapped modality information and sample data corresponding to the reverse-mapped modality, wherein the reverse-mapped modality information is generated by reverse cross-modal mapping of target modality prediction information, and reverse cross-modal mapping is an inverse process of cross-modal mapping;

[0091] Step 105: Optimize the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping, and prompt information through modality alignment loss and content alignment loss to complete training.

[0092] Steps 101 to 104 shown in the embodiment of the present application are helpful to enhance the model's understanding ability by adding prompt information to the multimodal sample data, and can promote two-way dynamic interaction between modalities through the use of cross-modal mapping. The modality alignment loss and content alignment loss can clearly guide the prompt learning to obtain a representation that is both semantically coherent and information complete during the conversion, thereby improving the adaptability and accuracy of the multimodal model, thereby improving the user's satisfaction when using products based on the multimodal model.

[0093] The specific implementation methods of the above steps are introduced below.

[0094] Before executing step 101 , some processing may be performed on the multimodal sample data.

[0095] Optionally, the multimodal sample data includes visual sample data, and the prompt information includes spatial visual prompt information. The spatial visual prompt information can be added to each channel of the spatial dimension of the visual information to obtain the visual sample data; wherein the spatial visual prompt information is aligned with the spatial dimension of the visual information.

[0096] The BMSPT framework can set basic parameters by developing and refining spatial visual cues and textual cues. We can start with spatial visual cues that are aligned with the entire spatial dimension of the input image. Specifically, we can introduce a spatial visual cue tensor (visual cue information) Scaling by a factor α and integrating it into the entire input image X can be expressed as:

[0097]

[0098] Where ⊕ represents adding the spatial cue tensor to each channel of the image spatial dimension. This enables the image encoder to process the input token X p , generating visual cue features f p =f(X p ,θ f ).

[0099] The above embodiments involve spatial visual cue tensors, which may refer to some additional data used to enhance or describe original visual information. These cue information may be spatial information about the position, direction, size, color, etc. of an object. Typically, visual information is represented by multi-dimensional data. For example, in image processing, common dimensions include height, width, and color channels (such as the three channels of R, G, and B in RGB). In multi-channel data, each channel represents a specific type of information. For example, in an RGB image, there are three channels corresponding to red, green, and blue, respectively. "Each channel" means that operations must be performed on each color channel of the image.

[0100] Spatial visual cues are incorporated into the original visual information. This addition can be simple superposition, fusion, or integration through more complex algorithms. The resulting new visual data is called "visual sample data." This dataset, containing both the original visual information and the added spatial visual cues, can be used for further analysis, processing, or machine learning tasks.

[0101] The spatial dimension here refers to the two-dimensional planar structure of the image, namely the height and width.

[0102] Optionally, the multimodal sample data includes text sample data, and the prompt information includes text prompt information, which can be expressed as: These textual cues are configured to supplement spatial visual cues, through g p =g(Y p ,θ g ) to extract text features, where Y p ={t SOS , P t , t1, t2, ..., t L , c k , t EOS For the explanation of the mathematical symbols in this part, please refer to the description of the CLIP model below.

[0103] It is understandable that in order to effectively adapt to their respective modalities, visual cues (spatial visual cues) and textual cues (textual cues) are jointly trained. This training can use the cross-entropy loss function to guide the development of visual cues and textual cues towards a more refined form.

[0104] It should be understood that when the multimodal sample data includes sound sample data, the prompt information includes audio prompt information, and the audio prompt information can be added to the sound data in a corresponding manner, which will not be described in detail here.

[0105] The above process ensures that the cues are ready for the subsequent bidirectional and spatial consistency learning stages. By adding spatial visual cues to each channel of the spatial dimension of visual information, the integrity of the subsequently processed visual data is guaranteed, enhancing its inherent quality without compromising the extraction of task-relevant features.

[0106] Based on the above operations, for example, when the input image is a zebra, the zebra will no longer be mistakenly classified as a hyena based on the similarity of local features or patch patterns.

[0107] In step 101 , a multimodal prompt feature is generated based on multimodal sample data, wherein the multimodal sample data includes visual sample data, text sample data, and prompt information.

[0108] Visual features, text features or audio features can be obtained based on the acquired multimodal sample data, such as visual sample data, text sample data or sound sample data, through a feature extraction module (such as an image encoder, a text encoder or an audio feature encoder), wherein the multimodal sample data includes learnable prompt information, such as text prompts and some category information or spatial visual prompts and image information, etc.

[0109] In some embodiments, for the extraction of image features, visual sample data can be collected from a variety of sources. These data can come from public datasets (such as ImageNet, COCO, etc.), specific field datasets, or be obtained through field collection, web crawling, etc. Next, a specially designed image encoder (such as ResNet, VGG, Inception, etc.) can be selected for feature extraction. The collected image data is subjected to preprocessing operations such as scaling and cropping, and spatial visual cues are added before being input into the selected encoder. These image data include original images (such as photos, scans, etc.) and learnable cues (such as image-specific annotations, attribute labels, etc.). The convolutional layers, pooling layers and other structures in the encoder are used to extract deep features of the image data. These features can capture the texture information, spatial structure and semantic content of the image.

[0110] To obtain text features, you can choose a specially designed text encoder (such as BERT, GPT, XLNet, etc.) for feature extraction. The collected text data is preprocessed by word segmentation, stop word removal, and text-specific prompt information is added before being input into the selected encoder. These text data include original text descriptions (such as category information) and learnable prompts (such as task-specific prompt words, guide words, etc.). The word embedding layer in the encoder, Transformer encoder and other structures are used to extract deep features of the text data. These features can capture the semantic information, contextual relationships and emotional tendencies of the text.

[0111] To obtain audio features, audio sample data can be collected from a variety of sources. These data can come from public datasets (such as LibriSpeech, AudioSet, etc.), specific domain datasets, or be obtained through field recording, web crawling, etc. Next, a specially designed audio encoder (such as VGGish, wav2vec, AST, etc.) can be selected for feature extraction. The collected audio data is preprocessed by denoising, normalization, and other operations, and audio-specific prompt information is added before being input into the selected encoder. These audio data include original audio signals (such as speech, music, etc.) and learnable prompts (such as audio-specific annotations, emotion tags, etc.). The convolutional layers, recurrent neural networks and other structures in the encoder are used to extract deep features of the audio data. These features can capture the spectral information, temporal structure and emotional characteristics of the audio.

[0112] Learnable prompts can be designed based on specific task requirements. These prompts can be task-specific keywords, phrases, or sentences, guiding the model to focus on task-relevant information. The learnable prompts are then fed into the corresponding encoder along with the original sample data (such as text descriptions, images, or sounds). The encoder's corresponding structure converts them into a unified feature representation, such as a multimodal prompt feature.

[0113] Exemplarily, visual cue features and text cue features may be acquired based on a contrastive language-image pre-training (CLIP) model.

[0114] CLIP is a typical dual-encoder architecture that aligns images and language, enabling the use of language to describe images. Using contrastive learning, CLIP uses the same feature extraction network for image processing and the same feature extraction network for natural language processing. It can be assumed that any input, after passing through the feature extraction network, becomes a feature vector containing the input features. This means that at higher dimensions, as long as the feature vectors describing the same object are similar, different descriptions of the same object can be combined between image and language, effectively describing the image using language.

[0115] CLIP consists of an image encoder that maps image inputs to feature vectors and a text encoder that performs the same operation on text inputs. The image encoder typically uses a convolutional neural network (CNN) that has been pre-trained in the image domain, such as ResNet or ViT (Vision Transformer). The text encoder typically uses a Transformer-based architecture, such as BERT (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pre-trained Transformer).

[0116] Through contrastive learning, CLIP aims to align the image feature space and the text feature space, thereby achieving zero-shot transfer capabilities for application to downstream tasks. A CLIP model can be represented by M = {f, g}, where f and g are the image encoder and text encoder, respectively.

[0117] Their pre-trained parameters are denoted as θ CLIP ={θ f ,θ g}, where θ f and θ g Corresponding to the parameters of image encoder and text encoder respectively. An input image is split into M small blocks and projected to generate block labels. The image encoder f processes these blocks through several transformer blocks to produce a latent visual feature representation f = f(X, θ f ). Meanwhile, the relevant category label y is encapsulated in a text template, such as “a photo of [category]”. This template is expressed as Y = {t SOC , t1, t2, ..., t L , c k , t EOS},in and c kThe word embeddings corresponding to the text template and category label respectively. SOS and t EOS is a learnable embedding of start and end tags. The text encoder encodes Y through multiple transformer blocks to produce latent text features g = g(Y, θ g ). For zero-shot inference, the text features of the text template with category labels {1, 2, …, C} are matched with the image features by computing the cosine similarity.

[0118] It should be understood that during the parameter adjustment process involved in the learnable hint, the parameters of the image encoder and the text encoder are frozen, and only the parameters in the hint are optimized.

[0119] Multimodal prompt features are generated based on multimodal sample data. The multimodal sample data includes prompt information, so that multimodal prompt features can be obtained to provide data support for subsequent cross-modal mapping.

[0120] In step 102, the multimodal prompt features are cross-modally mapped to obtain target modality prediction information.

[0121] Optionally, the multimodal prompt features generated by step 101 can be input into a cross-modal projector (for example, a visual to language projector, a language to visual projector, an audio to language projector, or a language to audio projector, etc.) for cross-modal mapping to obtain prediction information of the corresponding modality.

[0122] Exemplarily, the multimodal prompt features may include visual prompt features, text prompt features, and audio prompt features. When the multimodal prompt features include visual prompt features and text prompt features, the visual prompt features may be mapped from a visual to a textual space to obtain text prediction information, and the text prompt features may be mapped from a textual to a visual space to obtain visual prediction information.

[0123] The Vision-to-Language (VL) projector, also known as the Vision-to-Language mapping module, can be composed of representation, which is used to adapt visual data to the text embedding space while maintaining spatial consistency, thus achieving vision-to-language projection.

[0124] The Language-to-Vision (LV) projector, also known as the Language-to-Vision mapping module, can be composed of Representation is used to map text features into visual feature space to ensure the preservation of spatial integrity and achieve language-to-visual projection.

[0125] Optionally, the main task of the vision-to-language mapping module is to map visual features to text embedding space. Its structure can adopt a multi-layer perceptron (MLP), for example, including the following layers:

[0126] Input layer: Receives visual features from the image encoder.

[0127] Hidden layer: Use one or more fully connected layers and combine them with nonlinear activation functions (such as ReLU or GELU) to ensure that the features have a similar distribution to the text features after transformation.

[0128] Output layer: The dimension of the output matches the text embedding space.

[0129] Additional mechanisms: LayerNorm or residual connections may also be added to improve mapping stability and training convergence speed.

[0130] Optionally, the main task of the language-to-visual mapping module is to map text features back to the visual space. Its structure can adopt a deconvolutional network, for example, including:

[0131] Input layer: receives features from the text encoder.

[0132] Upsampling layer: Use the transpose convolution layer to upsample the features to match the spatial distribution of visual features.

[0133] Output layer: restores the dimensions to be similar to the image features.

[0134] Additional mechanism: BatchNorm or activation functions (such as LeakyReLU) can also be used to enhance the mapping effect.

[0135] It should be understood that in addition to the simple MLP structure, the VL projection module can also adopt other architectures to enhance the feature mapping capability, such as:

[0136] 1) Transformer Encoder layer.

[0137] The Transformer Encoder layer can perform multi-head attention calculation on visual features through the self-attention mechanism, and then perform nonlinear transformation through the feedforward network (FFN), thereby capturing more complex cross-modal relationships.

[0138] 2) Convolutional Neural Network (CNN) module.

[0139] The CNN module targets spatial information and can design a lightweight convolution module to retain local features, which are then mapped to the text space through a fully connected layer.

[0140] 3) Residual Network (ResNet) structure.

[0141] The ResNet structure uses residual connections to enhance feature transfer and combines fully connected layers or self-attention mechanisms to achieve VL mapping.

[0142] Correspondingly, the goal of the LV projection module is to map text features back to visual features with spatial distribution. Its implementation can be diversified. In addition to using two transposed convolutional layers, the following methods can also be considered:

[0143] 1) Combination of upsampling and convolution.

[0144] The combination of upsampling and convolution first uses nearest neighbor upsampling or bilinear interpolation to upsample text features to the target size, and then uses conventional convolutional layers for feature integration and refinement.

[0145] 2) Deconvolutional Network.

[0146] Deconvolution networks are similar to transposed convolutions, but can adopt more complex deconvolution block structures, such as residual deconvolution blocks, to enhance information recovery capabilities.

[0147] 3) Generator structure in generative adversarial network (GAN).

[0148] A GAN generator-like architecture can be adopted, taking the above text hint features as conditional inputs and gradually restoring them to the image size through a series of upsampling and convolution operations.

[0149] 4) Self-attention upsampling module.

[0150] The self-attention mechanism is used to capture long-distance dependencies during the upsampling process, and then combined with the upsampling operation to achieve fine-grained visual feature reconstruction.

[0151] The above examples illustrate possible implementation structures of the cross-modal projection module. In practical applications, flexible selection can be made based on actual conditions, and this application does not impose any restrictions on this.

[0152] The present embodiments meticulously preserve spatial dimensionality during cross-modal interaction. The VL projector adapts visual data to a representation of the text embedding space while maintaining spatial coherence. The LV projector, on the other hand, converts textual insights back into visual representations while ensuring that spatial integrity is preserved.

[0153] To promote semantically consistent and distributionally consistent representations, this application adopts a two-pronged strategy to fine-tune cues in both the VL and LV projections. This approach significantly enhances the adaptability of the framework, especially in cases of imbalanced datasets and limited sample sizes. By aligning visual cues in the visual space V with textual cues in the text space T, cross-modal interaction can be significantly optimized. This is achieved through the following steps.

[0154] In step 103 , a modality alignment loss is determined using the target modality prediction information and sample data corresponding to the target modality.

[0155] The modality alignment loss is used to ensure that the distribution of generated text and visual information is aligned with the respective data distribution. This alignment can be quantified using the Earthmoving Distance (Wasserstein distance), which is favored for its smooth continuity and effectiveness in measuring transfer costs in optimal transfer problems.

[0156] The above embodiments involve the Wasserstein distance, also known as the Earthmover's distance (EMD), a metric used to measure the difference between two probability distributions. In generative models (such as generative adversarial networks (GANs), the Wasserstein distance can be used to measure the difference between the generated distribution and the true distribution.

[0157] In some embodiments, the multimodal sample data includes visual sample data and text sample data, the target modality prediction information includes text prediction information and visual prediction information, and determining the modality alignment loss using the target modality prediction information and the sample data corresponding to the target modality may include:

[0158] Determine an alignment loss for visual-to-textual space mapping based on the text sample data and the text prediction information, wherein the alignment loss for visual-to-textual space mapping aims to minimize a bulldozer distance between the text sample data and the text prediction information so that the distribution of the text sample data is aligned with the distribution of the text prediction information;

[0159] An alignment loss for text-to-visual space mapping is determined based on the visual sample data and the visual prediction information, wherein the alignment loss for text-to-visual space mapping aims to minimize the bulldozer distance between the visual sample data and the visual prediction information so that the distribution of the visual sample data is aligned with the distribution of the visual prediction information.

[0160] Optionally, determining the alignment loss for visual-to-text space mapping based on text sample data and text prediction information may include:

[0161] Calculate the feature distribution mean of real text data (text sample data) in the visual transformation space;

[0162] Calculate the mean of the text features (i.e., text prediction information) obtained by cross-modal mapping (i.e., VL mapping);

[0163] Minimize the Wasserstein distance between the two so that the text distribution obtained by VL mapping is as close as possible to the real text distribution.

[0164] This process can be expressed mathematically as:

[0165]

[0166] in, Represents the mean of the feature distribution of real text data in the visual transformation space; Represents the mean of the text features obtained through VL mapping, represents the modality alignment loss of visual to text space mapping. f(x) represents the visual sample data including spatial visual cue information, represents VL mapping, F V is a 1-Lipschitz function. These functions are approximated by parameterized neural networks so that the weight w is restricted to the closed interval [-c, c], with the goal of being as close as possible to the true Wasserstein distance. Represents the data distribution p data (x) calculates the mathematical expectation of the sample x under (x), in this application, Used to calculate the feature mean of the visual modality, Used to calculate the feature mean of the text modality, y represents the sample extracted from the text sample data, y~p data (y), x represents the sample extracted from the visual sample data, x~P data (x), in order to simplify the representation, the subscripts i and j are omitted in formula (2).

[0167] Combining the CLIP encoder with the VL projector, a mapping That is, the mapping from the visual domain to the text domain, the output (Text prediction information) and target text [P t , y]∈T (text sample data), which is consistent with the Wasserstein distance. This setting theoretically induces The output distribution of closely reflects the empirical distribution p data (y), ensuring Therefore, the function Transform the visual input to a domain that closely matches the text domain T

[0168] Aims to minimize the real text sample {y i} and the converted visual samples The Wasserstein distance between them.

[0169] Optionally, determining the alignment loss of the text-to-visual space mapping based on the visual sample data and the visual prediction information may include:

[0170] Calculate the feature distribution mean of real visual data (visual sample data) in the text transformation space;

[0171] Calculate the mean of the visual features (i.e., visual prediction information) obtained by cross-modal mapping (i.e., LV mapping);

[0172] Ensure that the distribution of visual features is aligned with the true visual data distribution when converting from textual modality back to visual modality.

[0173] This process can be expressed mathematically as:

[0174]

[0175] in, Represents the mean of the feature distribution of real visual data in the text transformation space; Represents the mean of the visual features obtained through LV mapping, represents the modal alignment loss of text-to-visual space mapping, g(y) includes the text sample data of the text prompt information, Represents LV mapping, G T is a 1-Lipschitz function. To simplify the representation, the subscripts i and j are omitted in formula (3).

[0176] Aims to minimize the original visual sample {x i} and the converted text sample The Wasserstein distance between them.

[0177] This application determines the modality alignment loss through the target modality prediction information and the sample data corresponding to the target modality, which can ensure that the distribution of the generated text and visual information is aligned with the respective data distribution.

[0178] In step 104, the content alignment loss is determined by reverse mapping modal information and sample data corresponding to the reverse mapping modality, wherein the reverse mapping modal information is generated by target modality prediction information through reverse cross-modal mapping, and reverse cross-modal mapping is the inverse process of cross-modal mapping.

[0179] The content alignment loss introduced in this application strengthens the protection of cross-modal content integrity. The content alignment loss can be expressed as and y represents samples extracted from text sample data, and x represents samples extracted from visual sample data. This loss is combined with the modality alignment loss of visual and text domains to form a comprehensive objective of BMSPT, further enhancing semantic completeness and modality harmony.

[0180] In some embodiments, the multimodal sample data includes visual sample data and text sample data, the reverse mapping modality information includes reverse conversion of visual information and reverse conversion of text information, and determining the content alignment loss using the reverse mapping modality information and the sample data corresponding to the reverse mapping modality may include:

[0181] Determining a visual content alignment loss by inversely transforming the visual information and the visual sample data, wherein the visual content alignment loss aims to ensure that the inversely transformed visual information is within a preset neighborhood of desired visual information, where the desired visual information is used to represent any visual image belonging to the same category as the visual sample data;

[0182] The text content alignment loss is determined by reversely transforming the text information and the text sample data, wherein the text content alignment loss aims to minimize the L1 norm distance between the reversely transformed text information and the text sample data.

[0183] Optionally, determining the text content alignment loss by inversely transforming text information and text sample data means that for text data, a content alignment loss is implemented to ensure that the transformation from text to visual and back to text preserves the original text content. This is achieved through a cycle consistency framework, where text samples are mapped to the visual domain and then back, with the expectation of minimizing the deviation from the original text, which can be expressed as:

[0184]

[0185] in, Used to measure text content alignment loss, This formula implements a cycle consistency constraint, namely: the input text sample data y is converted into visual features through text-to-vision (LV) mapping; the visual features are then converted back into text features ψ(y) through vision-to-text (VL) mapping, that is, the text information is reversely converted; ||·||1 is used to measure the L1 norm of the deviation.

[0186] The text content alignment loss can ensure that the content of the text will not be lost during the mapping process, that is, the original semantics can be maintained after cross-modal transformation.

[0187] Optionally, a visual content alignment loss is determined by inversely transforming the visual information and visual sample data. In this process, for the visual data, the original image x is not required to be exactly restored; instead, the model allows the restored image to fall within the ε neighborhood of any visual image in the same class. This tolerance is quantified using a modified content alignment loss that aims to accommodate visual variability within the same class, which can be expressed as:

[0188]

[0189] in, It is used to measure the visual content alignment loss, j represents the category variable, which is used to determine whether the restored image falls within the ε neighborhood of the j category to which the original image belongs. Here, Indicator function Defined as:

[0190]

[0191] Among them, ||zx|| calculates the norm (usually the Euclidean norm) between the converted image z and the original image x, ensuring that the converted image is within an acceptable distance defined by ε. These losses are crucial to maintaining high fidelity of content during the conversion process, ensuring that the model achieves the goal of accurate content cross-modal conversion with appropriate allowable variations.

[0192] This application determines the content alignment loss by reverse mapping modality information and sample data corresponding to the reverse mapping modality, which can ensure the consistency of text and visual descriptions during the conversion process.

[0193] In step 105 , the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping, and prompt information are optimized through modality alignment loss and content alignment loss to complete the training.

[0194] Specifically, the learnable parameters involved in the visual-to-textual space mapping, text-to-visual space mapping, spatial visual cue information and textual cue information can be optimized through the alignment loss of visual-to-textual space mapping, the alignment loss of text-to-visual space mapping, the visual content alignment loss and the textual content alignment loss.

[0195] This process can be expressed as:

[0196]

[0197] Among them, λ1 and λ2 control the relative importance of each component. The goals of BMSPT are:

[0198]

[0199] The asterisk (*) in the upper right corner of the parameter in the formula represents the optimal value or optimal solution of the corresponding parameter, which is usually used to indicate the optimal parameter or variable value found in an optimization problem. v and P t Corresponding to the learnable parameters in the visual and textual cues during training, represents the learnable parameters in the Vision-to-Language (VL) projector, Representing learnable parameters in a language-to-vision (LV) projector.

[0200] The embodiments of the present application optimize the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping, and prompt information through modality alignment loss and content alignment loss, which not only solves the challenges brought by unbalanced category-image datasets and dimensionality differences between modalities, but also establishes a powerful framework for deep semantic integration between text and visual domains.

[0201] Please refer to Figure 2 , Figure 2 This is another flowchart of the multimodal model training method provided by some embodiments of the present application; it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0202] The method of the embodiment of the present application includes the following steps:

[0203] Step 201: obtaining visual prompt features and text prompt features based on visual sample data and text sample data, wherein the visual sample data includes spatial visual prompt information and the text sample data includes text prompt information;

[0204] Step 202: mapping the visual prompt features from visual to textual space to obtain text prediction information; mapping the text prompt features from textual to visual space to obtain visual prediction information;

[0205] Step 203: determining a first modality alignment loss based on the text prediction information and the text sample data, and determining a second modality alignment loss based on the visual prediction information and the visual sample data;

[0206] Step 204: determining a first content alignment loss by inversely transforming the visual information and the visual sample data, wherein the inversely transformed visual information is inversely generated by text prediction information through text-to-visual space mapping; and determining a second content alignment loss by inversely transforming the text information and the text sample data, wherein the inversely transformed text information is inversely generated by visual prediction information through visual-to-text space mapping;

[0207] Step 205: Optimize the learnable parameters involved in the visual-to-text space mapping, the text-to-visual space mapping, the spatial visual cue information, and the text cue information through the first modality alignment loss, the second modality alignment loss, the first content alignment loss, and the second content alignment loss to complete the training.

[0208] The specific implementation of each step is introduced below.

[0209] In step 201, visual prompt features and text prompt features can be obtained through a feature extraction module (such as an image encoder and a text encoder) based on the acquired visual sample data and text sample data, wherein the visual sample data and text sample data include learnable prompt information, such as text prompts and some category information or spatial visual prompts and image information, etc.

[0210] Optionally, visual cue features and text cue features may be acquired based on a contrastive language-image pre-training (CLIP) model.

[0211] In step 202, the visual prompt features are mapped from visual to textual space to obtain text prediction information, and the text prompt features are mapped from textual to visual space to obtain visual prediction information.

[0212] Optionally, the visual prompt features and text prompt features generated by step 201 can be input into a cross-modal projector (e.g., a visual-to-language projector or a language-to-visual projector) for cross-modal mapping to obtain prediction information of the corresponding modality.

[0213] The Vision-to-Language (VL) projector, also known as the Vision-to-Language mapping module, can be composed of representation, which is used to adapt visual data to the text embedding space while maintaining spatial consistency, thus achieving vision-to-language projection.

[0214] The Language-to-Vision (LV) projector, also known as the Language-to-Vision mapping module, can be composed of Representation is used to map text features into visual feature space to ensure the preservation of spatial integrity and achieve language-to-visual projection.

[0215] Optionally, the main task of the vision-to-language mapping module is to map visual features to text embedding space, and its structure can adopt a multi-layer perceptron (MLP).

[0216] Optionally, the main task of the language-to-visual mapping module is to map text features back to the visual space, and its structure can adopt a deconvolutional network.

[0217] The above examples illustrate possible implementation structures of the cross-modal projection module. In practical applications, flexible selection can be made based on actual conditions, and this application does not impose any restrictions on this.

[0218] The present embodiments meticulously preserve spatial dimensionality during cross-modal interaction. The VL projector adapts visual data to a representation of the text embedding space while maintaining spatial coherence. The LV projector, on the other hand, converts textual insights back into visual representations while ensuring that spatial integrity is preserved.

[0219] In step 203 , a first modality alignment loss is determined by using the text prediction information and the text sample data, and a second modality alignment loss is determined by using the visual prediction information and the visual sample data.

[0220] Optionally, an alignment loss of visual-to-text space mapping is determined based on the text sample data and the text prediction information, wherein the alignment loss of visual-to-text space mapping aims to minimize the bulldozer distance between the text sample data and the text prediction information so that the distribution of the text sample data and the distribution of the text prediction information are aligned.

[0221] The above embodiments involve the Wasserstein distance, also known as the Earthmover's distance (EMD), a metric used to measure the difference between two probability distributions. In generative models (such as generative adversarial networks (GANs), the Wasserstein distance can be used to measure the difference between the generated distribution and the true distribution.

[0222] An alignment loss for text-to-visual space mapping is determined based on the visual sample data and the visual prediction information, wherein the alignment loss for text-to-visual space mapping aims to minimize the bulldozer distance between the visual sample data and the visual prediction information so that the distribution of the visual sample data is aligned with the distribution of the visual prediction information.

[0223] Optionally, determining the alignment loss for visual-to-text space mapping based on text sample data and text prediction information may include:

[0224] Calculate the feature distribution mean of real text data (text sample data) in the visual transformation space;

[0225] Calculate the mean of the text features (i.e., text prediction information) obtained by cross-modal mapping (i.e., VL mapping);

[0226] Minimize the Wasserstein distance between the two so that the text distribution obtained by VL mapping is as close as possible to the real text distribution.

[0227] Optionally, determining the alignment loss of the text-to-visual space mapping based on the visual sample data and the visual prediction information may include:

[0228] Calculate the feature distribution mean of real visual data (visual sample data) in the text transformation space;

[0229] Calculate the mean of the visual features (i.e., visual prediction information) obtained by cross-modal mapping (i.e., LV mapping);

[0230] Ensure that the distribution of visual features is aligned with the true visual data distribution when converting from textual modality back to visual modality.

[0231] This application determines the first modality alignment loss through text prediction information and text sample data, and determines the second modality alignment loss through visual prediction information and visual sample data, which can ensure that the distribution of generated text and visual information is aligned with their respective data distributions.

[0232] In step 204, a first content alignment loss is determined by inversely converting visual information and visual sample data, wherein the inversely converted visual information is inversely generated by text prediction information through text-to-visual space mapping; and a second content alignment loss is determined by inversely converting text information and text sample data, wherein the inversely converted text information is inversely generated by visual prediction information through visual-to-text space mapping.

[0233] The content alignment loss introduced in this application strengthens the protection of cross-modal content integrity. The content alignment loss can be expressed as and y represents samples drawn from text sample data, and t represents samples drawn from visual sample data. This loss is combined with the modality alignment loss of visual and text domains to form a comprehensive objective of BMSPT, further enhancing semantic completeness and modality harmony.

[0234] Optionally, a text content alignment loss is determined by inversely transforming text information and text sample data. For text data, a content alignment loss is implemented to ensure that the transformation from text to visual and back to text preserves the original text content. This is achieved through a cycle consistency framework, where text samples are mapped to the visual domain and back again, with the goal of minimizing the deviation from the original text.

[0235] Optionally, a visual content alignment loss is determined by inversely transforming visual information and visual sample data. In this process, the original image x is not required to be exactly restored for the visual data; instead, the model allows the restored image to fall within the ε neighborhood of any visual image in the same class. This tolerance is quantified using a modified content alignment loss designed to accommodate visual variability within the same class.

[0236] This application determines the content alignment loss in the above manner, thereby ensuring consistency between text and visual descriptions during the conversion process.

[0237] In step 205: the learnable parameters involved in the visual-to-text space mapping, the text-to-visual space mapping, the spatial visual cue information and the text cue information are optimized by using the first modality alignment loss, the second modality alignment loss, the first content alignment loss and the second content alignment loss to complete the training.

[0238] Optionally, the learnable parameters involved in the visual-to-textual space mapping, the text-to-visual space mapping, the spatial visual cue information and the textual cue information can be optimized by the alignment loss of the visual-to-textual space mapping, the alignment loss of the text-to-visual space mapping, the visual content alignment loss and the textual content alignment loss.

[0239] The embodiments of the present application optimize the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping, and prompt information through modality alignment loss and content alignment loss, which not only solves the challenges brought by unbalanced category-image datasets and dimensionality differences between modalities, but also establishes a powerful framework for deep semantic integration between text and visual domains.

[0240] It should be understood that the above steps 201 to 205 illustrate one implementation method of the embodiment of the present application. In accordance with the implementation principles, those skilled in the art may also apply the embodiments illustrated in steps 101 and 105 to steps 201 to 205.

[0241] The following describes the solution of the embodiment of the present invention in detail with reference to specific application examples:

[0242] Please refer to Figure 3 In an embodiment of the present application, a method for training a multimodal model is provided. This method mainly proposes a solution to the problem that the multimodal model in the related art mainly focuses on one-way cross-modal semantic alignment, which has limitations in, for example, visual concept understanding, resulting in a decline in model performance and a reduction in user trust in the model.

[0243] Specifically, using the contrastive language-image pretraining (CLIP) model as the base model, textual cues and category information are input into the CLIP model's text encoder to obtain text features, where the textual features include textual cues. Image and spatial visual cues are input into the CLIP model's visual encoder to obtain visual features, where the visual features include visual cues. The textual cues and visual cues are input into the LV projector and VL projector, respectively, to obtain the corresponding predicted image and predicted text.

[0244] The modality alignment loss of the LV mapping is calculated by the image and spatial visual cues in the predicted image and the original input visual encoder; the obtained predicted image is reversely input into the visual encoder and VL projector to obtain the reverse generated text, and the content alignment loss of this article is calculated by inputting the reverse generated text and the original text.

[0245] The modality alignment loss of the VL mapping is calculated by using the text prompts and category information in the predicted text and the original input text encoder; the obtained predicted text is reversely input into the text encoder and LV projector to obtain the reverse generated image, and the visual content alignment loss is calculated by inputting the reverse generated image and the original image.

[0246] Afterwards, the learnable parameters of the two projectors and two modal cues are optimized through the modal alignment loss of LV mapping, the content alignment loss of this paper, the modal alignment loss of VL mapping, and the visual content alignment loss. If the optimization result does not meet the preset conditions, the next round of optimization is entered until the preset optimization goal is met.

[0247] The multimodal model training method provided in the present application promotes bidirectional interaction between modalities through two new vision-to-language (VL) and language-to-vision (LV) transfer functions. Furthermore, a two-pronged alignment approach (modality alignment and content alignment) is introduced to explicitly guide the learning of representations that are both semantically coherent and information-complete during the transfer, thereby improving the performance of the multimodal model.

[0248] The above is an introduction to an embodiment of the training method for the multimodal model provided in this application.

[0249] For the training method of the above multimodal model, this application also provides a multimodal model, namely the BMSPT model (BMSPT framework). Figure 4 , Figure 4 This is a schematic diagram of the training framework of the multimodal model provided in some embodiments of the present application.

[0250] BMSPT model, such as Figure 4 Shown, including:

[0251] Block embedding vectors are used to convert local regions in an image (such as image blocks or pixel blocks) into high-dimensional vectors that can capture the visual characteristics of the image blocks.

[0252] Specifically, the block embedding vector divides the image into multiple small blocks and extracts features from each block, resulting in a set of vectors that represent the visual content of each block. Through the embedding vector, the model can encode local information in the image, such as color, texture, and edges. The embedding vector also typically includes position information, so that the model knows the location of each block in the image, which is important for understanding the overall structure of the image.

[0253] The image encoding module is used to convert input visual sample data into visual features, where the visual sample data includes spatial visual cue information, and the visual features include visual cue features.

[0254] Specifically, the image encoding module converts the input image into a high-dimensional vector representation that captures the visual content and semantic information of the image. The image encoding module can extract useful visual features from the image, such as color, texture, shape, and object; integrate the features of local regions into a global image representation to capture the overall content and scene of the image; convert the original pixel values of the image into a higher-dimensional vector, so that the model can learn more complex feature representations in a higher-dimensional space; through pre-training tasks (such as contrastive learning), the vector representation of the image and the vector representation of the text are aligned in the same semantic space, thereby achieving cross-modal association between images and text; and learn universal visual features so that the model can perform well in various downstream tasks (such as image classification, object detection, image captioning, etc.).

[0255] In CLIP, the image encoding module is typically implemented using a convolutional neural network (CNN) or a transformer architecture. CNNs excel at processing local features and hierarchical structures, while transformers excel at processing long-range dependencies and global information. Regardless of the architecture used, the ultimate goal of the image encoding module is to generate a vector that effectively represents the image content, allowing for interactive and comparative learning with the output of the text encoding module.

[0256] A visual-to-linguistic space mapping module is used to map visual cue features to text embedding space.

[0257] Specifically, the visual-to-linguistic space mapping module maps visual information (i.e., the feature representation of the image) into the language space, allowing images and text to be compared and associated in the same vector space. Specifically, by mapping image features to the language space, the module enables visual and language information to be represented in the same semantic space, thereby promoting alignment and mutual understanding between the two different modal data.

[0258] The vision-to-language space mapping module can be implemented using a multi-layer perceptron (MLP), a Transformer Encoder layer, a convolutional neural network (CNN) module, or a residual network (ResNet) structure.

[0259] The text encoding module is used to convert the input text sample data into text features, where the text sample data includes text prompt information, and the text features include text prompt features.

[0260] Specifically, the text encoding module can work in conjunction with the image encoding module to achieve cross-modal understanding and association between images and text. The text encoding module can convert the input text (usually words or subwords) into word vectors, which are numerical representations of the text that can capture the semantic information of the words. By using models such as Transformer or recurrent neural networks (RNN), the text encoding module can understand the contextual meaning of words in a sentence and generate vector representations that contain contextual information. The text encoding module converts text into vectors in a high-dimensional space that can capture the deep semantics of the text, including sentence-level or paragraph-level meaning. In a multimodal model, the vector representation generated by the text encoding module can be compared and aligned with the output of the image encoding module to achieve consistency between visual and language information.

[0261] The encoding module in this paper plays different roles in different application scenarios. For example, in text generation scenarios, the text encoding module provides a representation of the input text, serving as the basis for the generative model. In question-answering systems, the text encoding module understands the question text and matches it with candidate answers. In information retrieval tasks, the text encoding module converts query text and documents into vector representations for calculating similarity and retrieving relevant documents.

[0262] In models like CLIP, the text encoding module typically uses a Transformer architecture, such as pre-trained language models like BERT or GPT. These models effectively capture the complex patterns and semantic information of text through self-attention mechanisms and multi-layer neural network structures. Through joint training, the text encoding module and the image encoding module can learn the deep connections between images and text, providing powerful feature representations for multimodal tasks.

[0263] The language-to-visual space mapping module is used to map textual prompt features to visual feature space.

[0264] Specifically, the language-to-visual space mapping module is used to map language information (typically text or a vector representation of text) into the visual space so that it can interact and compare with images or vector representations of images. Through this mapping, the mapping module enables language and visual information to be represented in the same vector space, facilitating the calculation of similarities or correlations between them. The mapping module is able to transform abstract language concepts into concrete visual features.

[0265] The language to visual space mapping module can be implemented using a deconvolutional network, or by combining upsampling with convolution, the generator structure in a generative adversarial network (GAN), and a self-attention upsampling module.

[0266] A contrastive learning module is used to project the feature vectors output by the image encoding module and the text encoding module into a common embedding space, and to bring matching image and text pairs closer together in the embedding space, while pushing unmatched image and text pairs further apart.

[0267] Specifically, the contrastive learning module is able to understand the correlation between images and texts and learn effective representations by simultaneously considering positive sample pairs (matching images and texts) and negative sample pairs (mismatching images and texts).

[0268] During the training process, the contrastive learning module regards each image-text pair as a positive sample pair. At the same time, it randomly selects mismatched images and texts from other images and texts in the batch to construct negative sample pairs. It can map image and text features to the same embedding space by projection, and calculate the similarity between positive and negative sample pairs in the embedding space, usually using cosine similarity or dot product. The contrastive learning module can use a contrastive loss function (such as NT-Xent loss, i.e. normalized temperature scaled cross entropy loss) to measure the difference between the similarity output by the model and the true label (positive sample pairs have high similarity and negative sample pairs have low similarity). The loss function encourages the model to bring positive sample pairs closer while pushing negative sample pairs away. The contrastive loss usually also includes a temperature parameter to adjust the smoothness of the similarity distribution to optimize the training process.

[0269] Among them, the learnable parameters in the visual-to-language space mapping module, the language-to-visual space mapping module, the spatial visual prompt information and the text prompt information are determined by the training method of the multimodal model in the above embodiment.

[0270] For the BMSPT model, it can be combined Figure 3As shown in the figure, the textual prompt and category information are input into the text encoder, and the generated textual prompt information is then input into the language-to-visual space mapping module (G). The image and spatial visual prompt information are sequentially input into the block embedding vector and the image encoder, and the generated visual prompt information is then input into the visual-to-language space mapping module (F). The parameters of the text encoder, image encoder, and block embedding vector are frozen. The textual prompt, spatial visual prompt, language-to-visual space mapping module, and visual-to-language space mapping module have learnable parameters.

[0271] Figure 4 The solid line represents the forward signal flow, and the dashed line represents the reverse signal flow.

[0272] The BMSPT model provided in this application integrates the CLIP pre-trained encoder infrastructure to support bidirectional interaction between textual and visual modalities. This framework employs innovative transformation functions and a two-pronged training approach, enabling multimodal cues to be represented with semantic consistency and information integrity during transformation, thereby ensuring efficient multimodal alignment.

[0273] As described above, the core of the BMSPT framework is a bidirectional alignment strategy that accurately preserves spatial dimensions during cross-modal interactions. The vision-to-language (VL) projector, denoted by F, adjusts visual data to the text embedding space while maintaining spatial consistency. Conversely, the language-to-visual (LV) projector, denoted by G, transforms textual insights into visual representations, ensuring the preservation of spatial integrity.

[0274] To promote the development of semantically consistent and distributionally consistent representations, BMSPT adopts a two-pronged strategy to fine-tune cues in both the VL and LV projections. This approach significantly enhances the adaptability of the framework, especially in the context of imbalanced datasets and limited sample sizes. By aligning visual cues in visual space with textual cues in text space, BMSPT optimizes cross-modal interaction.

[0275] It can be understood that the contents of the above method embodiments are applicable to the present model framework embodiments, the functions specifically implemented by the present model framework embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0276] The above is an introduction to the embodiment of the multimodal model provided by this application.

[0277] An implementation of an image recognition method provided by an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0278] Based on the multimodal model training method or multimodal model provided in the above embodiments, the present application also provides an image recognition method, such as Figure 5 As shown, Figure 5 This is a flowchart of the image recognition method provided by some embodiments of the present application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0279] The image recognition method embodiment of the present application includes the following steps:

[0280] Step 501: Acquire a visual image to be recognized;

[0281] Step 502: Input the visual image to be recognized into the visual language model to obtain a recognition result of the visual image;

[0282] Step 503: outputting the recognition result of the visual image;

[0283] The visual language model is trained using the multimodal model training method described above.

[0284] In steps 501 to 503 of the embodiment of the present application, a visual image is acquired and input into a multimodal model trained using the multimodal model training method provided in the above embodiment of the present application to obtain a recognition result of the visual image and output the result. Because the multimodal model is trained / pre-trained using the multimodal model training method provided in the embodiment of the present application, the accuracy of image recognition can be improved.

[0285] This image recognition method can be used in visual image question answering systems, visual image content review, image generation or autonomous driving systems.

[0286] The above is an introduction to the embodiment of the image recognition method provided in this application.

[0287] The BMSPT model provided in this application can also be applied to image editing applications under the Diffusion framework to ensure dynamic and fine-grained alignment between text and image modalities to improve the accuracy of image editing.

[0288] In order to further illustrate the effectiveness of the technical solution of the present application, the performance of the method of the present application on commonly used benchmark data sets is introduced below with reference to the accompanying drawings.

[0289] Exemplarily, CLIP (ViT-B / 16) is used as the pre-trained visual-language backbone to evaluate the effectiveness of BMSPT. The VL projector uses a multi-layer perceptron (MLP) architecture with a single fully connected layer, while the LV projector consists of two transposed convolutional layers, each supplemented with a ReLU activation function. At the same time, the number of text prompts M is set to 4, the hyperparameters α is set to 0.1 and ∈ is set to 0.2. For all tasks, the BMSPT model is trained using a batch size of 4, a learning rate of 0.0025, and SGD as the optimizer. To evaluate the generalization ability of the BMSPT framework, the categories are evenly divided into base and novel sets, and the model is trained only on the base categories and tested on the novel set.

[0290] Please refer to Figure 6 , Figure 6 The feature attention results of images with category labels of bird, cat, dog and watch are given in Figure 1. Among them, the Semi-BMSPT model (semi-BMSPT variant model) uses a unidirectional modality alignment loss function.

[0291] Comparing the bidirectional BMSPT model and the semi-BMSPT variant model shows that the proposed BMSPT focuses on unique visual features that are semantically related to the category label. For example, in the "cat" category, the BMSPT model accurately highlights the cat's eyes and facial features, while the semi-BMSPT variant model, due to the use of a unidirectional loss function, results in less precise and coherent alignment, showing more scattered local features. This result emphasizes the importance of the bidirectional structure, as it enables more accurate and contextually meaningful visual-semantic alignment.

[0292] Please refer to Figure 7 , Figure 7 Schematic diagram of experimental results of the diffusion model migration performance of the bidirectional BMSPT model provided in some embodiments of the present application.

[0293] Figure 7 A set of source images and target prompts corresponding to the source images are shown in Figure 2, as well as a schematic diagram of the migration performance of Stable Diffusion, Hiper model, and BMSPT combined with Stable Diffusion.

[0294] By combining BMSPT with Stable Diffusion, BMSPT can also be used for image processing tasks. Specifically, BMSPT does not fine-tune the diffusion model, but instead constructs an auxiliary alignment loss between the noisy image and the learning hint to retain some personalized information (e.g., the identity of the subject). Figure 7As shown in Figure 3, BMSPT is able to better modify background, texture, and dynamics while preserving the underlying structure. For example, compared to other methods, BMSPT exhibits similar texture and background to the source image on two bananas.

[0295] The above is a verification description of the performance effect of the embodiment of the present application.

[0296] The following describes in detail the implementation of the multimodal model training device provided in the embodiments of the present application with reference to the accompanying drawings.

[0297] Regarding the multimodal model training method provided in the above embodiment, the present application embodiment further provides a multimodal model training device 800 for implementing the above method, such as Figure 8 As shown, Figure 8 This is a schematic block diagram of a module of a multimodal model training device according to an embodiment of the present application. The multimodal model training device 800 includes:

[0298] A generating module 801 generates a multimodal prompt feature based on multimodal sample data, wherein the multimodal sample data includes visual sample data, text sample data, and prompt information;

[0299] A cross-modal mapping module 802 is used to perform cross-modal mapping on multi-modal prompt features to obtain target modality prediction information;

[0300] A first determining module 803 is configured to determine a modality alignment loss using target modality prediction information and sample data corresponding to the target modality;

[0301] a first determining module 804 configured to determine a content alignment loss using reverse mapping modality information and sample data corresponding to the reverse mapping modality, wherein the reverse mapping modality information is generated by reverse cross-modal mapping from target modality prediction information, and the reverse cross-modal mapping is an inverse process of cross-modal mapping;

[0302] The parameter optimization module 805 is used to optimize the learnable parameters involved in cross-modal mapping, reverse cross-modal mapping and prompt information through modality alignment loss and content alignment loss to complete training.

[0303] Correspondingly, the multimodal model training device 800 may also include the same modules to execute the method corresponding to another embodiment of the present application. For example:

[0304] A generating module 801 acquires visual prompt features and text prompt features based on visual sample data and text sample data, wherein the visual sample data includes spatial visual prompt information and the text sample data includes text prompt information;

[0305] a cross-modal mapping module 802 for mapping visual cue features from visual to textual space to obtain text prediction information, and mapping text cue features from textual to visual space to obtain visual prediction information;

[0306] A first determining module 803 is configured to determine a first modality alignment loss based on text prediction information and text sample data, and to determine a second modality alignment loss based on visual prediction information and visual sample data;

[0307] A first determining module 804 is configured to determine a first content alignment loss by inversely converting visual information and visual sample data, wherein the inversely converted visual information is inversely generated by text prediction information through text-to-visual space mapping; and to determine a second content alignment loss by inversely converting text information and text sample data, wherein the inversely converted text information is inversely generated by visual prediction information through visual-to-text space mapping;

[0308] The parameter optimization module 805 is used to optimize the learnable parameters involved in the visual-to-text space mapping, the text-to-visual space mapping, the spatial visual prompt information and the text prompt information through the first modality alignment loss, the second modality alignment loss, the first content alignment loss and the second content alignment loss to complete the training.

[0309] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0310] The implementation of the image recognition device provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0311] Regarding the multimodal model training method provided in the above embodiment, the present application embodiment further provides an image recognition device 900 for implementing the above method, such as Figure 9 As shown, Figure 9 This is a schematic block diagram of the modules of the image recognition device according to an embodiment of the present application. The multimodal model training device 900 includes:

[0312] Acquisition module 901: used to acquire the visual image to be recognized;

[0313] Input module 902: used to input the visual image to be recognized into the visual language model to obtain the recognition result of the visual image;

[0314] Output module 903: used to output the recognition result of the visual image.

[0315] The visual language model is trained using the multimodal model training method described above.

[0316] It can be understood that the contents of the above-mentioned image recognition method embodiment are applicable to the embodiment of the present image recognition device. The functions specifically implemented by the present image recognition device embodiment are the same as those of the above-mentioned image recognition method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned image recognition method embodiment.

[0317] like Figure 10 As shown, the embodiment of the present application further provides an electronic device 1000, which includes a memory 1001, one or more processors 1002 ( Figure 10 Only one is shown) and a computer program stored in memory 1001 and executable on processor 1002. Memory 1001 is used to store software programs and units, and processor 1002 executes the software programs and units stored in memory 1001 to perform various functional applications and data processing to obtain resources corresponding to the above-mentioned preset events. Optionally, processor 1002 implements the above-mentioned multimodal model training method when executing the above-mentioned computer program stored in memory 1001.

[0318] The memory 1001 is a non-transitory computer-readable medium that can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory 1001 may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 1001 may optionally include a memory remotely located relative to the processor 1002, and these remote memories may be connected to the processor via a network.

[0319] It can be understood that the contents of the above-mentioned multimodal model training method embodiment are applicable to the embodiment of this electronic device 1000, and the functions specifically implemented by the embodiment of this electronic device 1000 are the same as those in the above-mentioned multimodal model training method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned multimodal model training method embodiment.

[0320] like Figure 11 As shown, the embodiment of the present application further provides an electronic device 1100, which includes a memory 1101, one or more processors 1102 ( Figure 11Only one is shown) and a computer program stored in memory 1101 and executable by processor 1102. Memory 1101 is used to store software programs and units. Processor 1102 executes the software programs and units stored in memory 1101 to perform various functional applications and data processing to obtain resources corresponding to the aforementioned preset events. Optionally, processor 1102 implements the aforementioned image recognition method by executing the aforementioned computer program stored in memory 1101.

[0321] The memory 1101 is a non-transitory computer-readable medium that can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory 1101 may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 1101 may optionally include a memory remotely located relative to the processor 1102, and these remote memories may be connected to the processor via a network.

[0322] It can be understood that the contents of the above-mentioned image recognition method embodiments are applicable to the embodiments of this electronic device 1100. The functions specifically implemented by the embodiments of this electronic device 1100 are the same as those of the above-mentioned image recognition method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-mentioned image recognition method embodiments.

[0323] Embodiments of the present application also provide a computer program product, comprising a computer program that, when executed by one or more processors, implements the steps of the multimodal model training method described above. The product also includes the multimodal model training apparatus described above or the software application for the electronic device described above.

[0324] Specifically, the computer program product may be a desktop application, such as office software, graphic design software, games, or development tools, etc. The product may also be a mobile application, such as a social application, shopping application, educational application, or entertainment application, etc.

[0325] The computer program product can be expanded by adding new functional modules. For example, for office software, a project management module, a data analysis module, etc. can be added. A plug-in architecture can also be designed to allow third-party developers to create plug-ins, such as adding new filter plug-ins, special effects plug-ins, etc. to graphic design software.

[0326] Computer program products can be designed for operating system platforms such as Windows, MacOS, and Linux, as well as mobile operating system platforms such as Android and iOS. In addition, they can also be designed for the Web platform to achieve cross-platform use.

[0327] Computer program products can be applied to different scenarios such as individual users, corporate users, educational institutions, and government agencies. For example, for corporate users, an enterprise version of the software can be provided, which includes more advanced features and customized services.

[0328] It can be understood that the contents of the above method embodiments are all applicable to this computer program product, the functions specifically implemented by this computer program product embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0329] The multimodal model training method, model, image recognition method, electronic device and computer program product provided in the embodiments of the present application are conducive to enhancing the model's understanding ability by adding prompt information to multimodal sample data, promoting two-way dynamic interaction between modalities through the use of cross-modal mapping, and clearly guiding prompt learning through modal alignment loss and content alignment loss to achieve representations that are both semantically coherent and information complete in the conversion, thereby improving the performance of the multimodal model and the user experience.

[0330] The Bidirectional Multimodal Spatial Consistency Prompt Adjustment (BMSPT) framework provided in the embodiments of the present application is designed to significantly enhance modal alignment in visual-language models. The framework adopts a dual approach to ensure semantic coherence and information integrity across modalities within the original spatial dimension. BMSPT significantly improves the generalization ability of the model on new categories, promotes migration across datasets, and effectively addresses the challenges of datasets with domain shifts. In addition, by integrating BMSPT with a diffusion model, the framework demonstrates its potential to provide innovative solutions to cutting-edge challenges in text-to-image generation and multimodal learning, paving the way for future developments in this field.

[0331] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0332] Although specific embodiments are described herein, those skilled in the art will recognize that many other modifications or alternative embodiments are also within the scope of the present disclosure. For example, any of the functions and / or processing capabilities described in conjunction with a particular device or component may be performed by any other device or component. In addition, although various exemplary implementations and architectures have been described in accordance with the embodiments of the present disclosure, those skilled in the art will recognize that many other modifications to the exemplary implementations and architectures described herein are also within the scope of the present disclosure.

[0333] Some aspects of the present disclosure have been described above with reference to the block diagrams and flow charts of the systems, methods, systems and / or computer program products according to the exemplary embodiments. It should be understood that the combination of one or more blocks in the block diagram and the flow chart and the blocks in the block diagram and the flow chart can be realized by executing computer executable program instructions respectively. Equally, according to some embodiments, some blocks in the block diagram and the flow chart may not need to be executed in the order shown, or may not need to be executed in full. In addition, additional components and / or operations beyond those components and / or operations shown in the blocks in the block diagram and the flow chart may be present in certain embodiments.

[0334] Therefore, the blocks in the block diagrams and flow charts support combinations of means for performing the specified functions, combinations of elements or steps for performing the specified functions, and program instruction means for performing the specified functions. It should also be understood that each block in the block diagrams and flow charts, and combinations of blocks in the block diagrams and flow charts, can be implemented by a dedicated hardware computer system that performs the specific functions, elements, or steps, or a combination of dedicated hardware and computer instructions.

[0335] The program modules, applications, etc. described herein may include one or more software components, including, for example, software objects, methods, data structures, etc. Each such software component may include computer-executable instructions that, in response to execution, cause at least a portion of the functionality described herein (e.g., one or more operations of the illustrative methods described herein) to be performed.

[0336] Software component can be encoded with any one in various programming languages.A kind of exemplary programming language can be low-level programming language, such as the assembly language associated with specific hardware architecture and / or operating system platform.Comprise that the software component of assembly language instruction may need to be converted to executable machine code by assembler before being executed by hardware architecture and / or platform.Another exemplary programming language can be a more advanced programming language, and it can be transplanted across multiple architectures.Comprise that the software component of more advanced programming language may need to be converted to intermediate representation by interpreter or compiler before execution.Other examples of programming language include but are not limited to macro language, shell or command language, job control language, script language, database query or search language or report writing language.In one or more exemplary embodiments, the software component that comprises the instruction of one in the above-mentioned programming language example can be directly executed by operating system or other software component, without first being converted into another form.

[0337] Software components can be stored as files or other data storage structures. Software components of similar types or related functions can be stored together, such as in a specific directory, folder, or library. Software components can be static (e.g., preset or fixed) or dynamic (e.g., created or modified at execution time).

[0338] The embodiments of the present application are described in detail above in conjunction with the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in the relevant technical field without departing from the purpose of the present application.

Claims

1. A method for training a multimodal model, characterized in that: The following steps are involved: Generating a multimodal prompt feature based on multimodal sample data, wherein the multimodal sample data includes visual sample data, text sample data, and prompt information; Performing cross-modal mapping on the multimodal prompt features to obtain target modality prediction information; Determining a modality alignment loss using the target modality prediction information and sample data corresponding to the target modality; Determining a content alignment loss by reverse mapping modality information and sample data corresponding to the reverse mapping modality, wherein the reverse mapping modality information is generated by the target modality prediction information through reverse cross-modal mapping, and the reverse cross-modal mapping is an inverse process of the cross-modal mapping; The cross-modal mapping, the reverse cross-modal mapping, and the learnable parameters involved in the prompt information are optimized by using the modality alignment loss and the content alignment loss to complete the training.

2. The multimodal model training method according to claim 1, characterized in that: The multimodal sample data includes visual sample data, the prompt information includes spatial visual prompt information, and before generating the multimodal prompt feature based on the multimodal sample data, the method further includes: Adding the spatial visual cue information to each channel of the spatial dimension of visual information to obtain the visual sample data; The spatial visual prompt information is aligned with the spatial dimension of the visual information.

3. The multimodal model training method according to claim 1, characterized in that: The multimodal prompt features include visual prompt features and text prompt features, and the multimodal prompt features are cross-modally mapped to obtain target modality prediction information, including: Mapping the visual cue features to a visual-to-text space to obtain text prediction information; The text prompt features are mapped from text to visual space to obtain visual prediction information.

4. The multimodal model training method according to claim 1, characterized in that: The target modality prediction information includes text prediction information and visual prediction information, and determining the modality alignment loss using the target modality prediction information and sample data corresponding to the target modality includes: Determining an alignment loss of visual-to-textual space mapping based on the text sample data and the text prediction information, wherein the alignment loss of the visual-to-textual space mapping aims to minimize a bulldozer distance between the text sample data and the text prediction information so that a distribution of the text sample data is aligned with a distribution of the text prediction information; An alignment loss for text-to-visual space mapping is determined based on the visual sample data and the visual prediction information, wherein the alignment loss for text-to-visual space mapping aims to minimize a bulldozer distance between the visual sample data and the visual prediction information so that a distribution of the visual sample data is aligned with a distribution of the visual prediction information.

5. The multimodal model training method according to claim 1, characterized in that: The reverse mapping modality information includes reverse conversion visual information and reverse conversion text information, and determining the content alignment loss by using the reverse mapping modality information and sample data corresponding to the reverse mapping modality includes: determining a visual content alignment loss based on the inversely transformed visual information and the visual sample data, wherein the visual content alignment loss aims to ensure that the inversely transformed visual information is within a predetermined neighborhood of desired visual information, the desired visual information being used to represent any visual image belonging to the same category as the visual sample data; A text content alignment loss is determined by using the inversely converted text information and the text sample data, wherein the text content alignment loss aims to minimize an L1 norm distance between the inversely converted text information and the text sample data.

6. The multimodal model training method according to claim 1, characterized in that: Optimizing the learnable parameters involved in the cross-modal mapping, the reverse cross-modal mapping, and the prompt information by using the modality alignment loss and the content alignment loss includes: The learnable parameters involved in the visual-to-textual space mapping, text-to-visual space mapping, spatial visual cue information and textual cue information are optimized through the alignment loss of visual-to-textual space mapping, alignment loss of text-to-visual space mapping, visual content alignment loss and textual content alignment loss.

7. A multimodal model, characterized in that include: an image encoding module, configured to convert input visual sample data into visual features, wherein the visual sample data includes spatial visual cue information, and the visual features include visual cue features; A visual-to-linguistic space mapping module for mapping the visual cue features to a text embedding space; A text encoding module, configured to convert input text sample data into text features, wherein the text sample data includes text prompt information, and the text features include text prompt features; A language to visual space mapping module, configured to map the text prompt features to a visual feature space; a contrastive learning module, configured to project the feature vectors output by the image encoding module and the text encoding module into a common embedding space, and bring matching image and text pairs closer together in the embedding space, while pushing mismatched image and text pairs further apart in the embedding space; The learnable parameters in the visual-to-language space mapping module, the language-to-visual space mapping module, the spatial visual prompt information and the text prompt information are determined by the training method of the multimodal model according to any one of claims 1 to 6.

8. An image recognition method, characterized in that: The following steps are involved: Acquire a visual image to be recognized; Inputting the visual image to be recognized into a visual language model to obtain a recognition result of the visual image; outputting a recognition result of the visual image; The visual language model is obtained by training using the multimodal model training method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the multimodal model training method according to any one of claims 1 to 6 or the image recognition method according to claim 8.

10. A computer program product, characterized in that When the computer program product runs in an electronic device, the electronic device executes the multimodal model training method according to any one of claims 1 to 6 or the image recognition method according to claim 8.

Citation Information

Cited By

  • Multi-modal data multi-model combined training method and system

    CN121144858A