Information processing method based on multi-modal large model and related device
By introducing a dynamic weight generation module to process image patches and text sub-features, the problem of insufficient image understanding capability of multimodal large models is solved, and the task execution capability and semantic alignment accuracy are improved.
Patent Information
- Application Number
- CN202511534953.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-02-10
AI Technical Summary
The image understanding capabilities of multimodal large models need to be further improved in order to enhance their task performance.
A dynamic weight generation module is introduced. The dynamic weight generation module of the multimodal large model processes the visual sub-features and text sub-features of image patches, constructs a weight matrix, fuses visual features and text features, and generates the target processing result.
It improves the image content understanding and task execution capabilities of multimodal large models, and enhances the accuracy and robustness of cross-modal semantic alignment.
Smart Images

Figure CN121502373A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to the technical field of artificial intelligence, image processing, large model, etc. It can be used in application scenarios such as generative search, document intelligent editing, intelligent assistant, virtual assistant, intelligent e-commerce, etc. BACKGROUND
[0002] A multi-modal large model is an artificial intelligence model capable of processing data of multiple different modalities, such as text, images, etc. It has the ability to process and understand multiple data inputs, and can map data of different modalities into a shared space to achieve cross-modal semantic alignment. Through training on large-scale data, the model learns rich feature representations and can capture complex patterns and relationships in the data. Due to its strong representation learning ability, the multi-modal large model exhibits certain generalization ability and transfer learning ability, and is suitable for various downstream tasks such as image description generation, visual question answering, cross-modal retrieval, etc., and has great application potential in multiple fields. SUMMARY
[0003] The present disclosure provides a multi-modal large model-based information processing method and related device.
[0004] According to an aspect of the present disclosure, a multi-modal large model-based information processing method is provided, comprising: pairing a plurality of image blocks of a target image and a plurality of target word pieces in a target text to obtain a plurality of image-text pairs; processing visual sub-features of the image blocks and text sub-features of the target word pieces in each image-text pair using a dynamic weight generation module of a multi-modal large model to obtain a similarity between the image blocks and the target word pieces in each image-text pair; fusing visual features of the target image and text features of the target text based on a weight matrix constructed based on the similarity of each image-text pair to obtain a fused feature; generating a target processing result of the target image based on the fused feature.
[0005] According to another aspect of the present disclosure, a multi-modal large model-based information processing device is provided, comprising: a pairing module configured to pair a plurality of image blocks of a target image and a plurality of target word pieces in a target text to obtain a plurality of image-text pairs; a processing module configured to process visual sub-features of the image blocks and text sub-features of the target word pieces in each image-text pair using a dynamic weight generation module of a multi-modal large model to obtain a similarity between the image blocks and the target word pieces in each image-text pair; a fusion module configured to fuse visual features of the target image and text features of the target text based on a weight matrix constructed based on the similarity of each image-text pair to obtain a fused feature. The first generation module is configured to generate a target processing result of the target image based on the fused features.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to make the computer perform the method according to any of the embodiments of the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.
[0009] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them: Figure 1 is a system architecture diagram of an embodiment of the information processing method and device based on a multi-modal large model according to an embodiment of the present disclosure; Figure 2 is a flowchart of an information processing method based on a multi-modal large model according to an embodiment of the present disclosure; Figure 3 is a flowchart of an information processing method based on a multi-modal large model according to an embodiment of the present disclosure; Figure 4 is a schematic diagram of an information processing method based on a multi-modal large model according to an embodiment of the present disclosure; Figure 5 is a structural schematic diagram of a dynamic weight generation module according to an embodiment of the present disclosure; Figure 6 is a flowchart of generating false negative samples according to an embodiment of the present disclosure; Figure 7 is a structural schematic diagram of an information processing device based on a multi-modal large model according to an embodiment of the present disclosure; Figure 8 This is a block diagram of an electronic device used to implement an embodiment of the information processing method based on a multimodal large model of the present disclosure. Detailed Implementation
[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0012] It should be noted that, unless it is explicitly stated that there is a sequential order of execution between different operations, or that there is a sequential order of execution between different operations in terms of technical implementation, the execution order between multiple operations may not be significant, and multiple operations may be executed simultaneously.
[0013] Multimodal large models enable text-to-image search tasks, such as searching images using text, and also support image content understanding to generate text content. These tasks all require the multimodal large model's ability to understand image content.
[0014] However, in related technologies, the image understanding capabilities of multimodal large models need to be further improved in order to enhance the task execution capabilities of multimodal large models.
[0015] In view of this, embodiments of the present disclosure provide an information processing method based on a multimodal large model to improve the image content understanding capability of the multimodal large model.
[0016] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the information processing method and apparatus based on multimodal large models of the present disclosure can be applied.
[0017] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0018] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various components (such as applications and toolkits) can be installed on terminal devices 101, 102, and 103 and server 105 to enable information communication between them, such as complex task processing applications, browser applications, instant messaging applications, etc.
[0019] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0020] Server 105 can provide various services through its built-in applications. Taking a complex task processing application that provides one-click processing services for complex tasks as an example, server 105 can achieve the following effects when running this application: First, it receives input information from users via terminal devices 101, 102, and 103 through network 104. This input information may include natural language input and / or images. Then, a multimodal large model performs inference analysis on the input information to obtain the target processing result for the image. Specifically, when the input information includes an input image, the target processing result for that input image is specific to that image. When the input information does not include an input image, an image can be retrieved from an image library based on the natural language input in the input information, and the target processing result for that image can be obtained through the multimodal large model.
[0021] Furthermore, the server 105 can also transmit the target processing result back to the terminal devices 101, 102, and 103 via the network 104, so that the terminal devices 101, 102, and 103 can display the received target processing result to the user.
[0022] It should be noted that, in addition to being obtained from terminal devices 101, 102, and 103 via network 104, natural language input can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when starting to process previously reserved pending tasks), it can choose to retrieve this data directly from the local storage. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.
[0023] Because processing complex tasks requires significant computing resources and power, the information processing methods based on multimodal large models provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the information processing device based on the multimodal large model is also generally located in the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also complete the aforementioned calculations performed by the server 105 through complex task processing applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the complex task processing application determines that the terminal device it is on has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the information processing device based on the multimodal large model can also be located in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0024] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0025] like Figure 2 The diagram shown is a flowchart illustrating the information processing method based on a multimodal large model provided in this disclosure, including the following: S201, Pair multiple image blocks of the target image with multiple target words in the target text to obtain multiple image-text pairs.
[0026] The target image is the image that the multimodal large model needs to process, and the target text is the text that the multimodal large model needs to process. The text may come from the input of the terminal device, or it may be the text associated with the target image itself; this disclosure does not limit this aspect.
[0027] Each image patch in the target image is a local region of the target image, which may contain an image portion with complete semantics or be an image patch divided by a window. It is understood that different image patches may overlap, and some image patches may even contain several other image patches with detailed descriptions. This disclosure does not limit the method of dividing image patches; any method that can divide the target image into image patches according to different dimensions and / or levels of detail is applicable to the embodiments of this disclosure.
[0028] Target text typically contains one or more segments of natural language text. Therefore, target text contains multiple lexical units, and certain parts of speech can be extracted as target lexical units. For example, verbs, nouns, and adjectives. In implementation, lexical units that are particularly important to the target text, or descriptive text within the target text targeting different dimensions (such as color, action, etc.), can be determined as target lexical units based on requirements.
[0029] Therefore, it can be understood that image patches divide the target image into descriptions of different dimensions or granularities, while target words divide the target text into descriptions of different dimensions or granularities. This is beneficial for fine-grained comparison of complex images and / or complex text, thereby increasing the ability to understand the content of the target image.
[0030] Each constructed image-text pair includes an image patch and a target word. Each image patch can have its local visual sub-features extracted using feature extraction methods. Similarly, each target word can have its local textual sub-features extracted using feature extraction methods. Here, "local" refers to the content relative to the target image or target text. That is, for the target image, each image patch belongs to the local content, and for the target text, each target word belongs to the local content.
[0031] In order to fully understand the content association between the target image and the target text, the association between each image-text pair can be analyzed, which can be implemented as step S202.
[0032] S202, the dynamic weight generation module of the multimodal large model is used to process the visual sub-features of the image blocks and the text sub-features of the target words in each image-text pair, and obtain the similarity between the image blocks and the target words in each image-text pair.
[0033] During implementation, the similarity between the image blocks and target words in each image-text pair needs to be determined.
[0034] In this context, the same image patch can be paired with multiple target words to construct its own image-text pairs.
[0035] S203, a weight matrix is constructed based on the similarity of each image-text pair, and the visual features of the target image and the text features of the target text are fused to obtain the fused features.
[0036] The similarity of each image-text pair is an element in the weight matrix, used to constrain the semantic alignment of the target image and the target text.
[0037] S204, Based on the fusion features, generate the target image processing result.
[0038] Understandably, this disclosure introduces a dynamic weight generation module on top of the traditional multimodal large model architecture. This module dynamically constructs a weight matrix based on the feature similarity between image patches and target words in the image-text pair. This weight matrix is used to fuse the visual features of the target image and the textual features of the target text. Therefore, dynamically generated weights are used instead of fixed weights in the cross-modal semantic alignment and feature fusion stages. These dynamic weights are obtained through comparative analysis at the image patch and target word granularity, enabling a thorough understanding and exploration of the correlation between the target image and target text, thus allowing for better mutual understanding of the image and text. The fused features obtained based on this better support the task execution capabilities of the multimodal large model and improve the accuracy of the multimodal large model's inference results.
[0039] In some embodiments, the method of obtaining multiple image patches from a target image may include at least one of the following: Method 1) Perform a segmentation operation on the target image to obtain multiple image blocks.
[0040] For example, different sliding windows can be set according to the category of the target image, such as animals, landscapes, and people. The target image can be segmented according to the sliding window to obtain multiple image blocks. In order to fully understand the image content of the target image from different dimensions and granularities, multiple sliding windows of different sizes can be set for the same target image. Each sliding window of a certain size can obtain one or more corresponding image blocks.
[0041] This method of acquiring image patches is simple to operate and can quickly acquire multiple image patches, thereby improving the inference efficiency of multimodal large models and the user experience.
[0042] Method 2) Perform multi-classification on the target image to obtain image regions of each category, which are then used as image blocks.
[0043] For example, classification and regression operations can be performed on the target image to obtain the region range of different categories of objects in the target image. This region range can be represented by a rectangular bounding box. Thus, different categories of objects and their locations can be identified from the target image, resulting in multiple image patches.
[0044] Image blocks obtained through classification have relatively complete semantics for each image block, which facilitates alignment with the semantics of target words and improves the accuracy of content understanding of the target image.
[0045] For the target text, it can be segmented into words, and each segmented word can be used as a target word unit.
[0046] Of course, as explained above, you can also select the word segment with the target part of speech from multiple word segments as the target word segment.
[0047] Furthermore, content understanding can be performed on the target text, selecting the segment that best expresses the core content of the target text from multiple segmentation methods as target lexical units. Target lexical units obtained in this way can also be called core lexical units. Specifically, this can be implemented by processing the target text using a large language model to obtain the core lexical units within the target text. This method leverages the natural language understanding capabilities of the large language model to adaptively and accurately extract core lexical units from the target text, thereby improving the processing capabilities of the multimodal large model.
[0048] In some embodiments, after obtaining multiple image blocks and target words, an image block containing the foreground target can be selected from the multiple divided image blocks for constructing image-text pairs with the target words.
[0049] In other embodiments, in order to fully and meticulously understand the target image, multiple image blocks of the target image and multiple target words in the target text are paired to obtain multiple image-text pairs. This can be implemented by determining each image block and each target word as an image-text pair, thus obtaining multiple image-text pairs.
[0050] In other words, multiple image patches are paired with multiple target words to construct multiple image-text pairs. Each image patch is compared and analyzed with the target words in the target text to comprehensively and meticulously align the semantics of the image and text, thereby improving the accuracy of image content understanding. Through refined content understanding, the task execution capability of multimodal large models can be improved when faced with complex image content and complex text, thereby further enhancing the generalization ability of multimodal large models.
[0051] In some embodiments, the visual sub-features of each image block can be obtained by extracting features from the image block individually. Similarly, the textual sub-features of each target word can be obtained by extracting features from the target word individually.
[0052] To ensure that visual and textual sub-features are closely correlated with features extracted by a multimodal large model, thereby improving the accuracy of the generated weight matrix and better aligning the semantics of different modalities, this embodiment of the disclosure describes how obtaining visual sub-features of image patches and textual sub-features of target words in an image pair can be implemented as follows:Figure 3 As shown: S301, a visual feature extraction module based on a multimodal large model extracts visual features of the target image.
[0053] In practice, the Vision Transformer (ViT) can be used as a visual feature extraction module to extract features from the target image and obtain the visual features of the target image.
[0054] Of course, it is understood that other neural network modules used for extracting visual features are also applicable to the embodiments of this disclosure, and this disclosure is not limited thereto.
[0055] S302, the text feature extraction module based on multimodal large model extracts text features of the target text.
[0056] In implementation, a large language model (LLM) such as BERT (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pre-trained Transformer) can be used as a text feature extraction module to convert the input target text into a high-dimensional semantic vector, thereby obtaining the text features of the target text. This disclosure does not limit the specific structure of the text feature extraction module.
[0057] S303, extract the feature parts of the image patch from the visual features to obtain the visual sub-features of the image patch.
[0058] Among them, the feature part of the image patch, that is, the local visual feature which belongs to the visual feature, reflects the feature expression of the image patch in the target image.
[0059] S304. Extract the feature parts of the target word from the text features to obtain the text sub-features of the target word.
[0060] Similarly, the feature part of the target word, that is, the local text feature that belongs to the text feature, reflects the feature expression of the target word in the target text.
[0061] In this embodiment, by extracting local visual feature representations of image patches from the visual features of the target image as visual sub-features of the image patch, and extracting local text features of target words from the text features of the target text as text sub-features of the target words, the generation of the weight matrix is achieved by relying on the local features of the visual features of the target image and the text features of the target text. This ensures that the visual sub-features and text sub-features are in the same feature space as their global visual and text features, enabling a more reasonable generation of the weight matrix, aligning the semantics of different modalities, and improving the understanding ability of multimodal large models of text and images.
[0062] In implementation, to align the semantics of different modalities from different feature dimensions, embodiments of this disclosure can extract visual features at different levels and at least one level of text features. For example, a text feature extraction module extracts low, medium, and high-level visual features, and semantically aligns them with the text features at each level, thereby optimizing the expressive power of the fused features of the target image and target text. This can be implemented as follows: the image features of the target image output from the first target neural network layer in the visual feature extraction module of the multimodal large model are used as image features; The features of the target text output by the second target neural network layer in the text feature extraction module of the multimodal large model are used as text features; When the first target neural network layer includes multiple layers, the image features output by each first target neural network layer are respectively used as image features; When the second target neural network layer consists of multiple layers, the features output by each second target neural network layer are used as text features.
[0063] In this process, visual features at different levels and text features at corresponding levels are used to construct feature pairs, which are then used to generate fused features for those feature pairs.
[0064] Specifically, the visual feature extraction module contains a first target neural network layer that can be set according to requirements. For example, neural network layers capable of extracting low-level features, intermediate-level features, and high-level features can be used as the first target neural network layers.
[0065] Similarly, for the text feature extraction module, the second target neural network layer it contains can be set according to the requirements. For example, neural network layers that can extract low-level features, intermediate-level features and high-level features can be used as the second target neural network layers respectively.
[0066] Of course, in practice, regarding visual characteristics, the low, medium, and high levels of characteristics have the following differences and meanings, among which: Low-level features: Describe the raw attributes of an image, such as pixel values, brightness, and color. They are suitable for basic analysis, such as image enhancement and compression.
[0067] Intermediate features: Describe the local structure of an image, such as edges, textures, and shapes. They are suitable for object localization and structural analysis, such as object detection and image segmentation.
[0068] High-level features describe the global semantics of an image, such as object categories and scene descriptions. They are primarily used for classification and semantic understanding, such as image classification and scene recognition.
[0069] With the first target neural network layer comprising multiple visual features covering low, medium, and high levels, the visual features at each level can be fused with text features. This enables the semantic alignment of image and text modalities from different dimensions and levels of detail, thereby improving the accuracy of semantic alignment and enhancing the task execution and generalization capabilities of multimodal large-scale models.
[0070] In practice, when visual features include multiple levels, text features can be selected from one level and semantically aligned with visual features at different levels. Alternatively, text features can also include multiple levels, with each level having corresponding visual and text features.
[0071] Therefore, regardless of the number of levels of text features, the visual features and corresponding text features at each level are treated as feature pairs to be processed. For each feature pair, the information processing method provided in this disclosure embodiment can be executed. That is, Figure 2 The process and related technical solutions illustrated, including generating a weight matrix and performing feature fusion, can be understood as an operation performed on a single feature pair. The same operation is performed on each feature pair.
[0072] Of course, in cases involving multiple levels of image features and at least one level of text features, generating the target image processing result based on the fused features can be implemented as follows: Multiple fusion features generated from image features at multiple levels and text features at at least one level are fused together to obtain the target features; Based on the target features, generate the target image processing result.
[0073] like Figure 4 As shown, the visual features at each level and the corresponding text features at the same level are obtained through... Figure 2 The process shown generates a weight matrix for the feature pair, and then performs a fusion operation on the visual and textual features in the feature pair based on this weight matrix to obtain fused features. Each feature pair at each level yields its corresponding fused features, resulting in multiple fused features. These multiple fused features are then further fused to obtain the target feature. The subsequent multimodal large model performs further inference operations based on this target feature to obtain the final target processing result.
[0074] In this embodiment of the disclosure, target features are obtained by fusing different levels of fusion features. Subsequent reasoning is then performed based on the target features to obtain the target processing result. This allows for a full understanding of the relationship between images and text, thereby improving the task execution capability of multimodal large models.
[0075] After constructing multiple image-text pairs based on multiple image patches and multiple target words, a suitable weight matrix can be adaptively generated using the dynamic weight generation module of a multimodal large model. In implementation, a gating network based on a gating mechanism can be used to process the visual sub-features and text sub-features of the image patches in the image-text pairs.
[0076] In other embodiments, to fully explore the relationships between features of different modalities and enable cross-modal semantic alignment, this disclosure embodiment not only introduces a gating mechanism, but the dynamic weight generation module may also include a cross-attention layer and a self-attention layer. Thus, the dynamic weight generation module of the multimodal large model processes the visual sub-features of image blocks and the textual sub-features of target words in each image-text pair to obtain the similarity between image blocks and target words in each image-text pair. This can be implemented by performing the following operations for each image-text pair: Step A1: The visual sub-features and text sub-features of the image-text pair are processed using a cross-attention layer to obtain the first processing result.
[0077] In this implementation, the textual sub-features of the target word can serve as the query features (Q, query) required by the cross-attention layer, while the visual sub-features of the image patch can serve as the key features (K, key) and value features (V, value) required by the cross-attention layer. Based on this approach, the textual sub-features, acting as Q, can efficiently guide attention to key parts of the target image, accurately locating the required information. The visual sub-features, acting as K and V, can provide rich contextual supplementary information without interfering with the query intent, thus contributing to semantic alignment. Moreover, this asymmetric design reduces computational complexity, aligns with task-driven attention mechanisms, and improves the performance of multimodal large models.
[0078] Step A2 involves using a self-attention layer to process text sub-features, resulting in a second processing result.
[0079] In implementation, the query feature (Q) of the self-attention layer can be a textual sub-feature of the target term, and the key feature (K) and value feature (V) of the self-attention layer can be textual sub-features, other predefined features, or visual sub-features of image patches. During implementation, K and V of the self-attention layer can be determined according to the actual task requirements.
[0080] In multimodal large-scale models, using text sub-features as query features (Q) in the self-attention layer can accurately guide the multimodal large-scale model to focus on visual information related to the task objective in the target image, avoiding attention distraction caused by local attributes expressed by image features. As Q, text sub-features, in a semantically driven manner, give the multimodal large-scale model a clear objective, such as locking onto the fur region in an image when "describing cat fur." This design promotes multimodal semantic alignment, conforming to human cognitive habits—having the task objective first and then observing the image. Simultaneously, text sub-features, as Q, leverage their strong semantic expression to guide the multimodal large-scale model to focus on key visual features, improving the efficiency and accuracy of image content understanding.
[0081] Step A3: Perform feature transformation based on gating mechanism on the first and second processing results to obtain the third processing result.
[0082] Step A4: Determine the similarity between the third processing result and the first processing result to obtain the similarity between image blocks and target words in the image-text pair.
[0083] In this embodiment, a cross-attention layer is used to fully explore the semantic relationships between image patches and target words. A self-attention layer leverages text-guided multimodal large-scale models to understand the task, enabling task reasoning and image content analysis. This allows the dynamic weight generation module to fully understand image patches and target words. Furthermore, a gating mechanism is introduced to dynamically adjust the weights of the first and second processing results after feature transformation, thereby better balancing their contributions—that is, balancing the contributions of different modalities—and resolving alignment issues caused by semantic differences. The gating mechanism can flexibly allocate weights according to task requirements, enhancing the robustness of feature fusion and improving the accuracy of semantic alignment, ultimately improving the performance of cross-modal tasks.
[0084] In some embodiments, step A3 described above, which involves performing a feature transformation based on a gating mechanism on the first and second processing results to obtain a third processing result, can be implemented as follows: Step A31: The first processing result and the second processing result are fused to obtain the comprehensive attention feature.
[0085] During implementation, the method for merging the first and second processing results can be set according to actual needs. For example, methods such as splicing or weighted summation can be selected.
[0086] Step A32: Use a gating weight matrix to process the attention-integrated features to obtain intermediate features.
[0087] The gating weight matrix can be a parameter that is adjusted and optimized during the training of the multimodal large model. After training, the gating weight matrix can be regarded as a fixed parameter.
[0088] Furthermore, in other embodiments, the gating weight matrix can be parameters adaptively generated based on the input. For example, a separate gating weight generation network can be configured to generate the gating weight matrix based on the current input and / or historical states. The current input may include at least one of attention synthesis features, a first processing result, and a second processing result.
[0089] Of course, it is understandable that, as a learnable parameter, the gated weight matrix can be fixed through training to meet the task requirements. In order to reduce the complexity of the model and improve efficiency, the gated weight generation network may not be used.
[0090] Step A33: The intermediate features are processed using an activation layer to obtain the third processing result.
[0091] In this embodiment of the disclosure, by fusing the first processing result of the sub-attention layer and the second processing result of the self-attention layer, the dynamic weight generation network based on the gating mechanism can fully understand the feature association between image blocks and corresponding target words, thereby reasonably allocating the weights of different modalities based on the gating weight matrix and improving the accuracy of semantic alignment.
[0092] In some embodiments, the similarity between the third processing result and the first processing result can be determined by calculating cosine similarity, thus obtaining the similarity between image blocks and target words in the image-text pair. Cosine similarity is more suitable for global semantic judgment, such as document classification or semantic retrieval, where directional consistency and semantic consistency are more important.
[0093] Of course, in order to improve the ability to understand content at different dimensions or granularities, similarity can also be obtained by performing element-wise multiplication of the third processing result and the first processing result.
[0094] In this embodiment, element-wise multiplication preserves local information and original dimensions of features, facilitating element-level operations and highlighting locally useful features. This approach is suitable for scenarios relying on local features, particularly those involving image segmentation and target word segmentation. For example, it can aid in aligning image textures or text keywords. The output of pixel-wise multiplication can be directly used to calculate the dynamic weight matrix, providing richer information for the gating mechanism, thereby improving the accuracy of dynamic weights and enhancing the task execution and generalization capabilities of multimodal large models.
[0095] In summary, as Figure 5The diagram shows the structure of the dynamic weight generation module provided in this embodiment. For each image-text pair, the text sub-features of the target word in the image-text pair are used as query Q, and the visual sub-features of the image patch in the image-text pair are used as K and V, input to the cross-attention layer 501 to obtain a first processing result. The text sub-features are input as query Q to the attention layer 502 to obtain a second processing result. The first and second processing results are fused to obtain a comprehensive attention feature. The comprehensive attention feature is multiplied by the gating weight matrix and then passed through the activation layer 503 to obtain a third processing result. The third processing result and the first processing result are multiplied pixel-wise to obtain the similarity between the image patch and the target word. Thus, a weight matrix is constructed based on the similarity of multiple image-text pairs.
[0096] Figure 5 The expression for the dynamic weight generation module shown can be expressed as follows: In expression (1), Q represents the text sub-feature of each target word, K and V are the visual sub-features of the image patch, and W... g This is the gated weight matrix, σ is the sigmoid activation function, and ⊙ represents element-wise multiplication. This method achieves dynamic adjustment, allowing more correlated modal features to ultimately have higher weights.
[0097] In this embodiment, the gating weight matrix is not merely a static parameter, but is adaptively optimized using deep learning methods. During training, the weights are adjusted based on the actual semantic relationship between the image and text, learning deep semantic alignment between features of different modalities in the image and text. Compared to the traditional fixed-weight mechanism used when fusing features of different modalities, the adaptive mechanism of this embodiment improves the semantic alignment capability between images and text by dynamically adjusting the focus of attention. This approach significantly improves the quality of image-text matching, especially in complex contexts and diverse datasets, exhibiting stronger adaptability and robustness.
[0098] This embodiment of the disclosure avoids the shortcomings of traditional methods that simply concatenate or weight features between images and text by employing an adaptive attention mechanism and dynamic feature fusion. In the multimodal large-scale model processing system provided by this approach, image and text features are not only aligned at a surface level but also deeply fused at multiple levels and dimensions. During the feature fusion process, the dynamic attention mechanism allows the system to assign different weights to image and text information at different levels, thereby optimizing the image-text matching effect at each level. This mechanism of dynamically allocating the weight matrix enables the system to determine which "levels" (e.g., details or overall view) and which "dimensions" (e.g., color or action) of image and text features should be more tightly fused, thus producing the most accurate target processing result.
[0099] In practical implementation, traditional attention mechanisms typically use a static weighting method to handle the relationship between images and text, meaning that the weights of the relationship between images and text are fixed for each calculation. However, the embodiments disclosed in this disclosure... Figure 5 The method shown can introduce a gating mechanism into the adaptive attention mechanism, allowing the attention weights to be dynamically adjusted according to the characteristics of the input data and the task requirements.
[0100] In this mechanism, the multimodal large model calculates a gate weight matrix based on preliminary features of the image and text. This matrix controls the contribution of each modality in the attention mechanism and calculates the similarity of each image-text pair. Through the gating mechanism, the multimodal large model can dynamically adjust the attention weights according to similarity and task requirements, achieving more accurate semantic alignment.
[0101] In some embodiments, in addition to the aforementioned multi-dimensional and multi-level semantic alignment, this disclosure also enables the improvement of task processing accuracy and generalization ability of large multimodal models through cross-modal contrastive learning. This can be implemented by optimizing the large multimodal model based on a preset loss value.
[0102] Cross-modal contrastive learning achieves image-text alignment and matching by comparing the similarity between images and text. Existing contrastive learning methods primarily maximize the similarity of positive sample pairs (i.e., related images and text) and minimize the similarity of negative sample pairs. Therefore, the preset loss value may include the global loss between a first global feature of the target image and a second global feature of the target text. The first global feature may be a visual feature at least at the aforementioned level, and the second global feature of the target text may be a text feature at least at the aforementioned level. In this embodiment, through contrastive learning of global features, a multimodal large model can align text and images at the global level.
[0103] Using only global contrastive learning loss methods has several problems in practical applications: 1) It is difficult to effectively distinguish between negative samples and spurious negative samples; 2) It cannot capture the fine-grained semantic differences between images and text; 3) Existing methods rely on global alignment and ignore the importance of local semantics between images and text.
[0104] To overcome these problems, this invention proposes a cross-modal contrastive learning optimization (3CLO) method, which combines an improved contrastive loss function, spurious negative sample generation, local semantic alignment, and a contrastive learning strategy. Therefore, in this embodiment, a preset loss value is also introduced that includes at least one of the following: 1) First loss, used to represent the loss between the sample image and the false negative sample; In this embodiment of the disclosure, during the training phase, the sample image, such as the target image, may include two types of negative samples. One type is called a negative sample, and the other is called a spurious negative sample. For ease of understanding, several terms provided in this embodiment of the disclosure are explained: Negative samples: Negative samples refer to text-image pairs that do not match positive samples. In the contrastive learning framework, positive sample pairs are usually matched text-image pairs, while negative sample pairs are randomly paired text and images. The goal of the model is to learn to bring the representations of positive samples closer together while pushing the representations of negative samples further apart. Negative samples play a crucial role in training, helping the model learn to distinguish between correctly matched and incorrectly matched samples. In this embodiment, the negative samples are true negative samples, without any ambiguity.
[0105] Hard negative samples are the most difficult negative samples to distinguish in the current performance of the model. They are usually identified by mining their similarity scores or loss values online. They are the focus of the model's learning.
[0106] False negative samples are negative samples designed to be highly deceptive and closely approximate positive samples. They may be a subset of hard negative samples, but the emphasis is on their "disguised" nature. Identifying them may involve more complex semantic analysis or iterative model selection.
[0107] This can be understood as follows: negative samples are non-deceptive negative samples, hard negative samples are negative samples with a first degree of deceptivity, and spurious negative samples are negative samples with a second degree of deceptivity. The first degree is less deceptive than the second. Therefore, spurious negative samples are negative samples that are extremely difficult for large multimodal models to distinguish and are prone to errors.
[0108] Traditional contrastive learning cannot distinguish between hard negative samples and spurious negative samples. Spurious negative samples refer to image-text pairs that appear unrelated but are actually semantically connected. To address this issue, this disclosure proposes a spurious negative sample generation mechanism to improve model learning efficiency through adaptive generation of spurious negative samples.
[0109] When implementing, such as Figure 6 As shown, fake negative samples can be generated based on the following methods: S601, determine the similarity between multiple candidate images and multiple candidate texts.
[0110] S602, filter the similarity that meets the target conditions, and construct candidate sample pairs.
[0111] S603: For candidate sample pairs containing sample images, a False Negative Sample Generation Network (FNSGN) is used to process the candidate sample pairs and generate false negative samples.
[0112] This can be understood as follows: First, a preliminary screening is performed based on the initial matching of candidate images and candidate texts, selecting some seemingly unrelated but potentially semantically connected negative sample pairs. For example, completely irrelevant image-text pairs are eliminated. Then, from the remaining image-text pairs that are "somewhat similar but not completely identical" (which can be filtered by whether a similarity threshold meets the target condition), deceptively difficult negative sample pairs are selected as "advanced challenges" for training the model to identify subtle differences, i.e., candidate sample pairs. Then, a False Negative Sample Generation Network (FNSGN) is used to further optimize these samples, generating false negative samples with high semantic relevance, and adding them to the training for comparative learning. In this way, the system can effectively utilize "misjudged" samples to strengthen the model's ability to identify semantic relationships.
[0113] In practice, the specific reasons why the selected candidate samples are difficult for the multimodal large model to understand can be determined. For example, is it the ambiguity of color, or the deceptiveness of details (such as actions and micro-expression differences)? Then, the analyzed reasons guide the fake negative sample generation network to generate fake negative samples. The reasons can be based on manually labeled reasons or analyzed by the multimodal large model.
[0114] In practice, the fake negative sample generation network can be generated using a generator. This network identifies the most crucial, distinguishable detail in a positive sample (such as an action, color, or quantity) as the cause, then intentionally corrects this detail while keeping most other elements unchanged. This creates a very realistic "fake," i.e., a fake negative sample.
[0115] By introducing a fake negative sample generation network, we can identify and effectively utilize the potential semantic relationships between images and text, thereby improving the effect of contrastive learning and avoiding the inefficiency of randomly selecting negative samples in traditional methods.
[0116] After constructing false negative samples of the target image, the first loss can be determined based on these false negative samples to optimize the model parameters of the multimodal large model.
[0117] When implemented, the first loss includes both the first and second items; The first item is positively correlated with the similarity between the sample image and the fake negative sample; The second term is the cumulative value of multiple third terms; each third term is positively correlated with the similarity between the sample image and other fake negative samples. The first loss is determined based on the ratio of the first term to the second term.
[0118] The following expression (2) is an example of this first loss: (2) In expression (2), s(v i ,t f ) represents the sample image v i and fake negative samples t f The similarity between them, s(v i ,t j The similarity between the sample image and other fake negative samples is denoted as . Here, the numerator represents the first term, and the denominator represents the second term.
[0119] By introducing contrastive learning with spurious negative samples, we can effectively utilize "misjudged" samples to enhance the model's ability to identify semantic relationships.
[0120] 2) The second loss is used to represent the loss between image patches and core lexical units.
[0121] Cross-modal contrastive learning methods in related technologies are typically based on global features for comparison, but information in images and text often has multi-scale and multi-level features. Local regions of an image may correspond to certain specific descriptions in the text. To better capture such local relationships, embodiments of this disclosure introduce local semantic alignment and multi-scale contrastive learning.
[0122] Local semantic alignment and multi-scale contrastive learning methods combine local regions of an image (e.g., objects, details, or specific scenes in an image) with relevant descriptions in text (e.g., detailed descriptions, specific object names, etc.) to perform multi-scale contrastive learning of images and text. This enables the model to not only align the semantic relationships between images and text globally, but also to perform more refined matching at the local level, thereby improving its ability to understand complex scenes.
[0123] In practice, a second loss is introduced to achieve local semantic alignment and multi-scale contrastive learning. This second loss includes a fourth and a fifth term. The fourth item is positively correlated with the similarity between image patches of the sample image and individual core words; the method of obtaining image patches of the sample image is the same as that of the target image, and will not be repeated here. The fifth item is the cumulative value of multiple sixth items; each sixth item has a positive correlation with the similarity between the sample image and the core word unit; The second loss is determined based on the ratio between the fourth and fifth items.
[0124] For example, the second loss is exemplified as shown in expression (3): (3) In formula (3), the numerator is the fourth term and the denominator is the fifth term. local It is a visual sub-feature of an image patch, t local These are textual sub-features of the core lexical units.
[0125] The second loss function compares the visual sub-features (i.e., local features, such as the "leg region") of each "small fragmented" image patch in the target image with each "keyword" (local key semantics, such as the word "running") in the text. By calculating how well they "match" (i.e. how high the similarity), matching image patches and core words can be found. For example, the most matching combination of "legs" and "running" can be found, thus mapping them "one-to-one" to optimize the multimodal large model, enabling it to learn subtle differences and improve the task processing capabilities of the multimodal large model.
[0126] In summary, by introducing a first loss based on spurious negative samples, the feature distribution of spurious negative samples is learned, improving the recognition and processing capabilities of the multimodal large-scale model. By introducing a second loss, the multimodal large-scale model learns local detail contrast learning loss, further improving the large-scale model's understanding of image content guided by core lexical units, thus enhancing the multimodal large-scale model's task processing capabilities.
[0127] Finally, the multi-modal contrastive loss function (MCL) in this embodiment combines spurious negative sample generation, local semantic alignment, and global feature comparison to more comprehensively learn the similarity between images and text. This contrastive loss function also includes global feature comparison learning, as shown in the following expression (4): (4) In expression (4), L global It is a global loss, L local This is the second loss, L fn The first loss is based on spurious negative samples, and λ1, λ2, and λ3 are adjustment coefficients. By adjusting the weights of these three components, an optimal balance can be found between global feature alignment, local feature alignment, and negative sample optimization, depending on the specific task requirements.
[0128] Based on the same technical concept, this disclosure also proposes an information processing device 700 based on a multimodal large model, such as... Figure 7 As shown, it includes: The pairing module 701 is used to pair multiple image blocks of the target image with multiple target words in the target text to obtain multiple image-text pairs; The processing module 702 is used to process the visual sub-features of image blocks and the text sub-features of target words in each image-text pair using the dynamic weight generation module of the multimodal large model, so as to obtain the similarity between image blocks and target words in each image-text pair. The fusion module 703 is used to fuse the visual features of the target image and the text features of the target text based on the weight matrix constructed based on the similarity of each image-text pair to obtain the fused features; The first generation module 704 is used to generate the target processing result of the target image based on the fusion features.
[0129] In some embodiments, the dynamic weight generation module includes a cross-attention layer and a self-attention layer; The processing module includes the following units, which perform corresponding operations for each image-text pair: The first processing subunit is used to process the visual sub-features and text sub-features of the image-text pair using a cross-attention layer to obtain the first processing result; The second processing subunit is used to process text sub-features using a self-attention layer to obtain a second processing result; The third processing subunit is used to perform feature transformation based on a gating mechanism on the first and second processing results to obtain the third processing result. The calculation subunit is used to determine the similarity between the third processing result and the first processing result, and to obtain the similarity between the image block and the target word in the image-text pair.
[0130] In some embodiments, the third processing subunit is specifically used to: perform a fusion operation on the first processing result and the second processing result to obtain attention-integrated features; Gated weight matrices are used to process attention-comprehension features to obtain intermediate features; The intermediate features are processed using an activation layer to obtain the third processing result.
[0131] In some embodiments, the computational subunit is specifically used for: The similarity is obtained by performing an element-wise multiplication operation between the third processing result and the first processing result.
[0132] In some embodiments, a first acquisition module is further included, configured to: The visual feature extraction module based on a multimodal large model extracts visual features from the target image; The text feature extraction module based on a multimodal large model extracts text features from the target text. The visual features of an image patch are extracted from the visual features to obtain the visual sub-features of the image patch. The feature parts of the target word are extracted from the text features to obtain the text sub-features of the target word.
[0133] In some embodiments, the image features of the target image output by the first target neural network layer in the visual feature extraction module of the multimodal large model are used as image features; The features of the target text output by the second target neural network layer in the text feature extraction module of the multimodal large model are used as text features; When the first target neural network layer includes multiple layers, the image features output by each first target neural network layer are respectively used as image features; When the second target neural network layer includes multiple layers, the features output by each second target neural network layer are used as text features. In this process, visual features at different levels and text features at corresponding levels are used to construct feature pairs, which are then used to generate fused features for those feature pairs.
[0134] In some embodiments, an optimization module is also included, for: Optimize the multimodal large model based on the preset loss value; The preset loss value includes at least one of the following: The first loss is used to represent the loss between the sample image and the false negative sample; The second loss is used to represent the loss between image patches and core words in the sample image.
[0135] In some embodiments, the first loss includes the first item and the second item; The first item is positively correlated with the similarity between the sample image and the fake negative sample; The second term is the cumulative value of multiple third terms; each third term is positively correlated with the similarity between the sample image and other fake negative samples. The first loss is determined based on the ratio of the first term to the second term.
[0136] In some embodiments, the second loss includes the fourth and fifth items; The fourth item is positively correlated with the similarity between image patches and individual core words in the sample image; The fifth item is the cumulative value of multiple sixth items; each sixth item has a positive correlation with the similarity between the sample image and the core word unit; The second loss is determined based on the ratio between the fourth and fifth items.
[0137] In some embodiments, a second generation module is further included, for: Determine the similarity between multiple candidate images and multiple candidate texts; Filter the similarity that meets the target conditions, and construct candidate sample pairs corresponding to them; For candidate sample pairs containing sample images, a fake negative sample generation network is used to process the candidate sample pairs and generate fake negative samples.
[0138] In some embodiments, a second acquisition module is further included, configured to: The target text is processed using a large language model to obtain the core lexical units in the target text.
[0139] In some embodiments, the preset loss value further includes: The global loss between the first global feature of the target image and the second global feature of the target text.
[0140] In some embodiments, a third acquisition module is further included, for: The target image is segmented to obtain multiple image blocks; The target image is classified into multiple categories to obtain image regions of each category, which are then used as image blocks.
[0141] In some embodiments, the pairing module is specifically used to determine each image block and each target word as a text-image pair, thereby obtaining multiple text-image pairs.
[0142] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0143] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0144] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0145] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0146] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0147] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0148] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as information processing methods based on multimodal large models. For example, in some embodiments, the information processing method based on multimodal large models can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the information processing method based on multimodal large models described above can be performed. Alternatively, in other embodiments, computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform information processing methods based on multimodal large models.
[0149] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0150] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0151] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0153] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0154] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0155] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0156] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An information processing method based on a multimodal large model, comprising: Multiple image blocks in the target image are paired with multiple target words in the target text to obtain multiple image-text pairs; The dynamic weight generation module of the multimodal large model is used to process the visual sub-features of image blocks and the text sub-features of target words in each image-text pair to obtain the similarity between image blocks and target words in each image-text pair. A weight matrix is constructed based on the similarity of each image-text pair. The visual features of the target image and the text features of the target text are then fused to obtain the fused features. Based on the fusion features, the target processing result of the target image is generated.
2. The method according to claim 1, wherein, The dynamic weight generation module includes a cross-attention layer and a self-attention layer; The dynamic weight generation module using a multimodal large model processes the visual sub-features of image blocks and the textual sub-features of target words in each image-text pair to obtain the similarity between image blocks and target words in each image-text pair, including: For each image-text pair, execute the following: The visual sub-features and text sub-features of the image-text pair are processed using the cross-attention layer to obtain the first processing result; The text sub-features are processed using the self-attention layer to obtain a second processing result; A third processing result is obtained by performing a feature transformation based on a gating mechanism on the first and second processing results. The similarity between the third processing result and the first processing result is determined to obtain the similarity between the image block and the target word in the image-text pair.
3. The method according to claim 2, wherein, The step of performing a feature transformation based on a gating mechanism on the first processing result and the second processing result to obtain a third processing result includes: The first processing result and the second processing result are fused to obtain the comprehensive attention feature. The attention-comprehension features are processed using a gating weight matrix to obtain intermediate features; The intermediate features are processed using an activation layer to obtain the third processing result.
4. The method according to claim 2, wherein, Determining the similarity between the third processing result and the first processing result, and obtaining the similarity between image blocks and target words in the image-text pair, includes: The similarity is obtained by performing an element-wise multiplication operation between the third processing result and the first processing result.
5. The method according to any one of claims 1-4, wherein, Obtaining visual sub-features of image patches and textual sub-features of target words in the image pair includes: The visual feature extraction module based on the multimodal large model extracts the visual features of the target image; The text feature extraction module based on the multimodal large model extracts the text features of the target text; The feature portion of the image patch is extracted from the visual features to obtain the visual sub-features of the image patch; The feature portion of the target word is extracted from the text features to obtain the text sub-feature of the target word.
6. The method according to any one of claims 1-5, wherein, The image features of the target image output by the first target neural network layer in the visual feature extraction module of the multimodal large model are used as the image features; The features of the target text output by the second target neural network layer in the text feature extraction module of the multimodal large model are used as the text features; When the first target neural network layer includes multiple layers, the image features output by each first target neural network layer are respectively used as the image features; When the second target neural network layer includes multiple layers, the features output by each second target neural network layer are respectively used as the text features; In this process, visual features at different levels and text features at corresponding levels are used to construct feature pairs, which are then used to generate fused features for those feature pairs.
7. The method according to claim 1, wherein, Also includes: The multimodal large model is optimized based on a preset loss value; The preset loss value includes at least one of the following: The first loss is used to represent the loss between the sample image and the false negative sample; The second loss is used to represent the loss between image blocks and core words in the sample image.
8. The method according to claim 7, wherein, The first loss includes the first item and the second item; The first item has a positive correlation with the similarity between the sample image and the false negative sample; The second term is the cumulative value of multiple third terms; each third term has a positive correlation with the similarity between the sample image and other false negative samples; The first loss is determined based on the ratio of the first term to the second term.
9. The method according to claim 7, wherein, The second loss includes items four and five; The fourth item is positively correlated with the similarity between image blocks and individual core words in the sample image. The fifth item is the cumulative value of multiple sixth items; The sixth item of each item has a positive correlation with the similarity between the sample image and the core word element; The second loss is determined based on the ratio between the fourth and fifth terms.
10. The method of claim 7 or 8, further comprising generating spurious negative samples based on the following method: Determine the similarity between multiple candidate images and multiple candidate texts; Filter the similarity that meets the target conditions, and construct candidate sample pairs corresponding to them; For candidate sample pairs containing the sample images, a fake negative sample generation network is used to process the candidate sample pairs and generate fake negative samples.
11. The method according to claim 7 or 8, further comprising: The target text is processed using a large language model to obtain the core word units in the target text.
12. The method according to claim 7, wherein the preset loss value further includes: The global loss between the first global feature of the target image and the second global feature of the target text.
13. The method according to any one of claims 1-12, further comprising acquiring the plurality of image patches based on at least one of the following methods: The target image is segmented to obtain multiple image blocks; The target image is classified into multiple categories to obtain image regions of each category, which are then used as image blocks.
14. The method according to any one of claims 1-13, wherein, The process of pairing multiple image blocks of the target image with multiple target words in the target text to obtain multiple image-text pairs includes: Each image block and each target word are identified as an image-text pair, thus obtaining the plurality of image-text pairs.
15. An information processing device based on a multimodal large model, comprising: The matching module is used to match multiple image blocks of the target image with multiple target words in the target text to obtain multiple image-text pairs; The processing module is used to process the visual sub-features of image blocks and the text sub-features of target words in each image-text pair using the dynamic weight generation module of the multimodal large model, so as to obtain the similarity between image blocks and target words in each image-text pair. The fusion module is used to fuse the visual features of the target image and the text features of the target text based on the weight matrix constructed based on the similarity of each image-text pair to obtain fused features; The first generation module is used to generate the target processing result of the target image based on the fusion features.
16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-14.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-14.
18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-14.