Multi-modal data processing method and device

By extracting and fusing image features and combining large language model processing technology, the multimodal large language model's performance problem is solved when processing complex instructions for multiple images of correlation relationships, achieving more accurate and efficient processing effects.

CN120181133APending Publication Date: 2025-06-20LENOVO (BEIJING) LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510237514.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When existing multimodal large language models process complex instructions involving text and multiple images, they cannot accurately understand the correlation between multiple images, resulting in poor processing performance.

Method used

The detailed features and global features of the image are extracted by a visual encoder, fused into natural language features, and processed through a large language model in combination with task indication information to understand and perform tasks based on multiple image association relationships.

Benefits of technology

Improves the ability to understand the relationship between multiple images, and enhances the accuracy and performance of processing complex instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181133A_ABST
    Figure CN120181133A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data processing method and device, and the method comprises the steps: obtaining N to-be-processed images and task indication information, the task indication information represents the execution of a target processing task at least based on the incidence relation among the N images, and N is an integer greater than or equal to 2; determining detail features and global features of the N images through a visual encoder; based on fusion features of the detail features and the global features, mapping the fusion features into natural language features through a visual mapper; and processing the natural language features and the text features of the task indication information through a large language model so as to execute the target processing task at least based on the association relationship among the N images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a multimodal data processing method and device. Background Art

[0002] The Multimodal Large Language Model (MLLM) is a technology that combines vision and large language models. Through the multimodal large language model technology, it is possible to process data in multiple modalities including text, images, and audio, and generate natural language outputs.

[0003] Currently, when processing instructions for a single image (such as describing the details of a single image or answering questions related to a single image) through the multimodal large language model technology, corresponding answers can be accurately generated. However, since the large language model cannot understand the correlation between multiple images, the performance of processing complex instructions involving text and multiple images through the multimodal large language model technology is poor, and corresponding task processing results cannot be accurately generated. Summary of the Invention

[0004] On the one hand, this application provides a multimodal data processing method, including:

[0005] Obtain N images to be processed and task indication information, where the task indication information represents performing a target processing task based at least on the correlation between the N images, and N is an integer greater than or equal to 2;

[0006] Determine the detailed features and global features of the N images through a visual encoder;

[0007] Based on the fused features of the detailed features and global features, map the fused features to natural language features through a visual mapper;

[0008] Process the natural language features and the text features of the task indication information through a large language model to perform the target processing task based at least on the correlation between the N images.

[0009] In a possible implementation manner, the process of processing the natural language features and the text features of the task indication information through a large language model to perform the target processing task based at least on the correlation between the N images includes:

[0010] Process the natural language features and the text features of the task indication information through a large language model to obtain intermediate layer features output by the intermediate layer of the large language model;

[0011] Determine the residual features between the intermediate layer features, the natural language features, and the text features;

[0012] Process the intermediate layer features and the residual features through the large language model to perform the target processing task based at least on the association relationship between the N images.

[0013] In another possible implementation manner, the determining the residual features between the intermediate layer features, the natural language features, and the text features includes:

[0014] Convert the intermediate layer features into first features, where the first representation space corresponding to the first features is the same as the representation space corresponding to the natural language features and the text features;

[0015] Determine the residual between the first features, the natural language features, and the text features to obtain initial residual features;

[0016] Perform a conversion process on the initial residual features to obtain the residual features, where the second representation space corresponding to the residual features is the representation space corresponding to the intermediate layer features.

[0017] In another possible implementation manner, the determining the detailed features and the global features of the N images through the visual encoder includes:

[0018] Extract features from the N images through the visual encoder to obtain at least one detailed feature output by at least one intermediate layer of the visual encoder and the global feature output by the output layer of the visual encoder.

[0019] In another possible implementation manner, it further includes:

[0020] Convert the detailed features into second features, where the third representation space corresponding to the second features is the representation space corresponding to the global features;

[0021] Add the second features and the global features to obtain the fused features of the detailed features and the global features.

[0022] In another possible implementation manner, the determining the residual features between the intermediate layer features, the natural language features, and the text features includes:

[0023] Determine the residual features between the intermediate layer features, the natural language features, and the text features through the residual network module associated with the large language model;

[0024] The residual network module is obtained in the following manner:

[0025] Based on the differences between at least two image samples in the image group information, train the residual network module in the model architecture composed of the visual encoder, visual mapper, large language model, and residual network module, so that the large language model can learn the association relationship between at least two image samples in the image group information.

[0026] In another possible implementation, the image group information includes: at least two image samples and the sample association relationship actually possessed by the at least two image samples;

[0027] The training of the residual network module in the model architecture based on the differences between at least two image samples in the image group information includes:

[0028] Based on the differences between at least two image samples in the image group information, with the goal of minimizing the gap between the predicted association relationship of the at least two image samples determined by the large language model and the sample association relationship, train the residual network module in the model architecture;

[0029] The at least two image samples include a first image sample and at least one second image sample, and the second image sample is an image sample generated based on the first image sample through data augmentation or an image generation model.

[0030] In another possible implementation, before training the residual network module in the model architecture, it further includes: using a third image sample as a training sample, with the goal of minimizing the gap between the predicted description information determined by the large language model in the sub-model architecture and the actual description information corresponding to the third image sample, train the visual encoder and visual mapper in the sub-model architecture, and obtain the trained visual encoder and visual mapper. The sub-model architecture is composed of the visual encoder, visual mapper, and large language model in the model architecture.

[0031] In another possible implementation, the image group information further includes: a task indication information sample for indicating the determination of the association relationship between the at least two image samples;

[0032] The training of the residual network module in the model architecture based on the differences between at least two image samples in the image group information, with the goal of minimizing the gap between the predicted association relationship corresponding to the at least two image samples determined by the large language model and the sample association relationship, includes:

[0033] Determine the sample detail features and sample global features of the at least two image samples through the visual encoder;

[0034] The sample fusion feature based on the sample detail features and the sample global features maps the sample fusion feature into a sample natural language feature through the visual mapper;

[0035] Process the sample natural language feature and the sample text feature of the task indication information sample through the large language model to obtain the sample intermediate layer feature output by the middle layer of the large language model;

[0036] Determine the sample residual feature between the sample intermediate layer feature and the sample natural language feature and the sample text feature through the residual network module;

[0037] Based on the sample intermediate layer feature and the sample residual feature, determine the predicted association relationship between the at least two image samples through the large language model;

[0038] If it is determined that the training objective has not been met based on the predicted association relationship corresponding to the at least two image samples and the sample association relationship, adjust the parameters of the residual network module, and return to execute the operation of processing the sample natural language feature and the sample text feature of the task indication information sample through the large language model until the training objective is met. The training objective is to minimize the gap between the predicted association relationship and the sample association relationship.

[0039] In another aspect, the present application also provides a multi-modal data processing device, including:

[0040] A data acquisition unit for acquiring N images to be processed and task indication information, where the task indication information represents performing an object processing task based at least on the association relationship between the N images, and N is an integer greater than or equal to 2;

[0041] A feature determination unit for determining the detail features and global features of the N images through a visual encoder;

[0042] A feature mapping unit for mapping the fusion feature into a natural language feature through a visual mapper based on the fusion feature of the detail feature and the global feature;

[0043] A data processing unit for processing the natural language feature and the text feature of the task indication information through a large language model to perform the object processing task based at least on the association relationship between the N images. Description of the Drawings

[0044] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0045] Figure 1 It is a schematic flowchart of a multi-modal data processing method provided by this application;

[0046] Figure 2 It is another schematic flowchart of a multi-modal data processing method provided by this application;

[0047] Figure 3 It is an example diagram of an implementation framework for a residual network module to determine residual features in this application;

[0048] Figure 4 It is a schematic training flowchart for training a residual network module in this application;

[0049] Figure 5 It is another schematic flowchart of a multi-modal data processing method in this application;

[0050] Figure 6 It is an example diagram of a model framework adopted by the multi-modal data processing method in this application;

[0051] Figure 7 It is a schematic diagram of a composition structure of a multi-modal data processing device provided by this application;

[0052] Figure 8 It is a schematic diagram of a composition architecture of an electronic device provided by this application. Specific Embodiments

[0053] The embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. The terms used in the embodiments part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0054] In the description, claims and the above-mentioned drawings of the present application, terms such as "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of the present application. In addition, the terms "comprising", "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0055] As Figure 1 , a schematic flowchart of a multimodal data processing method provided by the present application is shown. The method of this embodiment may include:

[0056] S101, obtain N images to be processed and task indication information.

[0057] Wherein, N is an integer greater than or equal to 2.

[0058] The task indication information characterizes that a target processing task is performed at least based on the association relationship between the N images.

[0059] For example, the target processing task indicated by the task indication information may be: describe the differences between these N images; generate a story by combining these N images; or, generate a fault handling code by combining these N images, etc., without limitation.

[0060] Wherein, the data type of the task indication information may have various possibilities, without limitation. For example, the task indication information may be task indication information in text form, or task indication information in voice format, etc.

[0061] S102, determine the detail features and global features of the N images through a visual encoder.

[0062] The visual encoder is used to encode visual data such as images into feature representations. In the present application, in order to be able to discover the association relationship between the N images, two types of features of the N images are extracted through the visual encoder, namely detail features and global features.

[0063] Among them, the detailed features of N images characterize the features in details such as the edges, textures, corners, and colors of the N images. For example, when the visual encoder is a neural network, the detailed features can also be called shallow features, which can be the features output by the intermediate layer (generally the relatively front network layer rather than the last output layer) of the visual encoder. Since the detailed features focus on the local detailed information of the N images rather than the global information of the N images, the semantic information of the N images contained in the detailed features is relatively small.

[0064] The global features of N images characterize the global structural information and context information of the N images. Therefore, more semantic information of the N images can be reflected through the global features. For example, when the visual encoder is a neural network, the global features can also be called deep features, which can be the features extracted from several relatively later network layers or the output layer of the visual encoder.

[0065] For example, in a possible implementation, the visual encoder can be used to extract features from the N images to obtain at least one detailed feature output by at least one intermediate layer of the visual encoder and the global feature output by the output layer of the visual encoder.

[0066] S103, based on the fusion feature of the detailed feature and the global feature, map the fusion feature into a natural language feature through a visual mapper.

[0067] Among them, the fusion feature is the feature obtained by fusing the detailed feature and the global feature. Therefore, the fusion feature contains not only the features of the N images that can fully reflect the fine-grained features such as texture and edge, but also can relatively fully reflect the semantic information expressed by the N images.

[0068] Among them, the visual mapper is used to map the visual feature (the fusion feature in this application) into a feature that can be understood by the large language model. Correspondingly, the natural language feature is a feature that can be understood by the large language model and can contain the detailed feature and the global feature. For example, the fusion feature is mapped into the word embedding space corresponding to the large language model through the visual feature to obtain the natural language feature.

[0069] S104, process the natural language feature and the text feature of the task indication information through the large language model to perform the target processing task at least based on the association relationship between the N images.

[0070] For example, splice the natural language feature and the text feature of the task indication information to obtain a spliced feature. Process the spliced feature through the large language model so that the large language model can determine the task result of the target processing task based on the association relationship between the N images.

[0071] Among them, the text features of the task indication information are the features expressed by the text corresponding to the task indication information. For example, if the task indication information is in text form, the features of the task indication information can be directly extracted to obtain the text features. If the task indication information is in voice form, the corresponding text information of the task indication information can be first converted, and the text information can be converted into text features.

[0072] It has been found through research that the correlation relationship between N images is not only related to the semantic information expressed by the N images, but also related to the detailed information of the N images. In this application, the natural language features are the features mapped based on the detailed features and global features of N images, and the global features can reflect the semantic information expressed by the N images. Therefore, the large language model can analyze the correlation relationship between N images based on the natural language features, and naturally can complete the target processing task corresponding to the task indication information based on this correlation relationship.

[0073] As can be seen from the above content, after determining the global features and detailed features of N images through the visual encoder in this application, the visual mapper will map the fusion features of the global features and detailed features of N images into natural language features. Since the global features of N images can more fully reflect the semantic information expressed by the N images, and the detailed features of N images can characterize the detailed information of the N images, the natural language features can contain relatively rich semantic information and detailed information corresponding to the N images, so that the large language model can more accurately mine the correlation relationship between the N images by combining the natural language features, and naturally can more accurately complete the target processing task indicated by the task indication information based on the correlation relationship between the N images.

[0074] It can be understood that based on the characteristics of the large language model itself, the large language model will also pay more attention to the semantic information of the images and may ignore some detailed information. Based on this, in order to further improve the accuracy of the task processing results obtained by the large language model in performing the target processing task, a residual network module can also be constructed in this application. On this basis, in the process of the large language model processing the natural language features and the text features of the task indication information, the natural language features output by the visual mapper and the text information corresponding to the task indication information are deeply interacted and fused with the features output by the middle layer of the large language model, so that the large language model can not only pay attention to the semantic information of the images, but also pay more attention to the detailed information of the images to mine the correlation relationship between N images.

[0075] The following will be combined with Figure 2 for illustration. For example Figure 2 , Fig. shows another schematic flowchart of the multi-modal data processing method provided by the embodiment of the present application. The method of this embodiment may include:

[0076] S201, Obtain N images to be processed and task indication information.

[0077] Among them, the task indication information represents performing a target processing task based at least on the association relationship between the N images.

[0078] N is an integer greater than or equal to 2.

[0079] S202, Determine the detailed features and global features of the N images through a visual encoder.

[0080] S203, Based on the fused features of the detailed features and global features, map the fused features to natural language features through a visual mapper.

[0081] For the above steps S201 to S203, reference can be made to the relevant introductions in the previous embodiments, which will not be elaborated here.

[0082] S204, Process the natural language features and the text features of the task indication information through a large language model to obtain the intermediate features output by the intermediate layer of the large language model.

[0083] Among them, the intermediate layer of the large language model (i.e., the intermediate layer of the large language model) refers to the network layer between the input layer and the output layer in the large language model. Among them, the input layer refers to the first network layer of the large language model and is also the network layer where the large language model receives the natural language features and the text features; the output layer is the network layer where the large language model outputs the final processing result. In this application, any network layer between the input layer and the output layer in the large language model can be used as the intermediate layer of the large language model, so as to use the features output by the intermediate layer of the large model as the intermediate features.

[0084] For example, the features output by an intermediate layer (e.g., the 6th network layer of the large language model) located after the input layer and relatively close to the front in the large language model can be obtained as the intermediate features, which can be specifically set according to needs and are not limited thereto.

[0085] S205, Determine the residual features between the intermediate features and the natural language features and text features.

[0086] For example, determine the concatenated features obtained by concatenating the natural language features and the text features, and determine the residual features between the intermediate features and the concatenated features.

[0087] Among them, the residual features may include the features that the natural language features and text features have and the intermediate features do not have. Based on this, the residual features can make up for the feature information ignored in the process of the large language model processing the natural language features and text features.

[0088] The residual feature can be determined by a residual network module.

[0089] In the present application, there can be multiple possible specific implementations for determining the residual feature, and there is no specific limitation.

[0090] In one possible implementation, the intermediate layer feature can be converted into a first feature, where the first representation space corresponding to the first feature is the same as the representation space corresponding to the natural language feature and the text feature. For example, the first representation space is the representation space corresponding to the concatenated feature obtained by concatenating the natural language feature and the text feature. On this basis, the present application can determine the residual between the first feature and the natural language feature and the text feature to obtain an initial residual feature. Then, the initial residual feature is subjected to a conversion process to obtain the residual feature. Among them, the second representation space corresponding to the residual feature is the representation space corresponding to the intermediate layer feature.

[0091] Among them, the representation space of a feature is a multi-dimensional space used to describe feature data. Features in different representation spaces are mapped to the same representation space to achieve the purpose of feature alignment, so as to facilitate arithmetic processing between different features.

[0092] Among them, there can also be multiple possible specific implementations for determining the initial residual feature between the first feature and the natural language feature and the text feature, and there is no specific limitation. For example, in an alternative manner, the concatenated feature obtained by concatenating the natural language feature and the text feature can be multiplied by the first feature to obtain the initial residual feature.

[0093] Among them, converting the intermediate layer feature into the first feature and converting the initial residual feature into the residual feature can be achieved in multiple ways, as long as it can map the feature to the required representation space, and there is no specific limitation. For example, a linear projector, a non-linear projector, or a neural network model, etc. can be used to map the feature to convert the feature to the corresponding representation space.

[0094] To facilitate understanding of the implementation process of determining the residual feature in the above possible implementation, in combination with Figure 3 a simple explanation is given. Figure 3 Fig. shows a block diagram of an implementation for the residual network module to determine the residual feature.

[0095] From Figure 3 it can be seen that after the natural language feature output by the visual mapper and the text feature corresponding to the task prompt information are concatenated into a concatenated feature, the concatenated feature will be input into the large language model. The large language model has multiple network layers, such as Figure 3 each layer of the large language model is a network layer. On this basis, after the large language model processes the concatenated feature, the present application will obtain the intermediate layer feature output by the intermediate layer of the large language model.

[0096] Take, for example, the case where the residual network module transforms the features into the required representation space through a linear projector in Figure 3 . It can be seen that the residual network module includes: a linear projector 11 for receiving the intermediate layer features, and a guidance module (such as the Guidance module) for performing a target operation on the concatenated feature obtained by concatenating the output of the linear projector 11 with the natural language feature and the text feature. In Figure 3 Take, for example, the case where the guidance module is used to perform a dot product operation on the concatenated feature and the output feature of the linear projector 11. The residual network module further includes a linear projector 12 for linearly mapping the feature output by the guidance module. The output feature of the linear projector 12 will be input into the large language model.

[0097] Based on this, the intermediate layer features will be input into the linear projector 11 in the residual network module. The intermediate layer features are mapped to the representation space corresponding to the concatenated feature through the linear projector 11 to obtain the first feature. The first feature will be multiplied by the concatenated feature obtained by concatenating the natural language feature and the text feature corresponding to the task prompt information to obtain the initial residual feature.

[0098] The initial residual feature is mapped to the representation space corresponding to the intermediate layer feature output by the intermediate layer of the large language model through the linear projector 12 to obtain the residual feature. As Figure 3 shown, the residual feature output by the linear projector 12 will be input into the next network layer after the intermediate layer of the large language model.

[0099] S206. Process the intermediate layer features and the residual features through the large language model to perform the target processing task based at least on the correlation relationship between N images.

[0100] For example, input the intermediate layer feature output by the intermediate layer of the large language model and the residual feature into the network layer after the intermediate layer of the large language model together, so that the large language model can continue to analyze the correlation relationship between N images based on the intermediate layer feature and the residual feature, and perform the target processing task based on this correlation relationship.

[0101] As Figure 3 shown, input the combined feature obtained by adding the intermediate layer feature and the residual feature into the next network layer after the intermediate layer of the large language model, so that each network layer after the intermediate layer of the large language model processes the combined feature, and finally obtains the task result of the target processing task.

[0102] In this embodiment, in the process of processing the natural language features output by the visual mapper and the text features corresponding to the task indication information through the large language model, the residual features between the intermediate layer features output by the large language model and the natural language features and the text features will be determined. The residual features can compensate for the feature information ignored in the process of the large language model processing the natural language features and the text features, that is, compensate for the features missing in the intermediate layer features relative to the natural language features and the text features. On this basis, by further processing the intermediate layer features and the residual features through the large language model, the large language model can pay attention to the semantic information of N images while also paying attention to the detailed information of the N images, thereby further improving the accuracy of the large language model in determining the association relationship between the N images, and naturally improving the accuracy of the task result of the target processing task determined by the large language model.

[0103] In this application, the large language model can be a pre-trained model.

[0104] Furthermore, in order to ensure that the large language model can accurately perform target task processing based on the association relationship between N images, this application also needs to train the residual network module.

[0105] For example, in one implementation, the residual network module in the model architecture composed of the visual encoder, visual mapper, large language model, and residual network module can be trained based on the differences between at least two image samples in the image group information, so that the large language model can learn the association relationship between at least two image samples in the image group information.

[0106] Among them, the image group information includes at least two image samples, and these two image samples are different. Of course, in practical applications, there can be multiple pieces of image group information used to train the residual network module.

[0107] In an optional manner, at least two image samples in the image group information include a first image sample and at least one second image sample. Among them, the second image sample is an image sample generated based on the first image sample through data augmentation or an image generation model.

[0108] Among them, data augmentation of the first image sample can be to perform different transformations on the first image sample to generate an image different from the first image sample. For example, the first image sample can be processed by cropping, scaling, randomly flipping, adjusting brightness, or adjusting contrast, etc. to generate the second image sample.

[0109] Among them, there are various possibilities for the image generation model, which are not specifically limited. Based on the first image sample, the image generation model can generate an image that is different from the first image sample, so as to obtain one or more second image samples.

[0110] As can be seen from the above, there is a certain correlation between the second image sample and the first image sample, and the two are different.

[0111] Furthermore, in order to train the residual network module more effectively and conveniently, the image group information may include: at least two image samples and the sample correlation relationship actually possessed by the at least two image samples. On this basis, the present application can train the residual network module in the model architecture composed of the visual encoder, the visual mapper, the residual network module, and the large language model with the goal of minimizing the gap between the predicted correlation relationship of the at least two image samples determined by the large language model and the sample correlation relationship.

[0112] Among them, with the goal of minimizing the gap between the predicted correlation relationship of the at least two image samples determined by the large language model and the sample correlation relationship, the parameters of the residual network module can be continuously adjusted during the training process, so that after the two image samples are processed by the model architecture, the predicted correlation relationship finally output by the large language model is as consistent as possible with the sample correlation relationship actually corresponding to the two image samples.

[0113] It can be understood that the processing process of the model architecture for the at least two image samples is similar to the processing process of the N images by the visual encoder, the visual mapper, the residual network module, and the large language model before, and will not be elaborated here. Correspondingly, there are various possibilities for the specific implementation of training the residual network module of the model architecture based on the differences between at least two image samples in the image group information, which are not limited.

[0114] Next, taking an implementation manner as an example, the process of training the residual network module will be introduced. As Figure 4 , a training process schematic diagram of training the residual network module in the present application is shown. The method of this embodiment may include:

[0115] S401, obtain at least one piece of image group information.

[0116] Each piece of image group information includes: at least two image samples, the sample correlation relationship actually possessed by the at least two image samples, and a task indication information sample for indicating the determination of the correlation relationship between the at least two image samples.

[0117] S402. For each group of image information, determine the sample detail features and sample global features of at least two image samples in the group of image information through a visual encoder.

[0118] In this application, for the sake of easy distinction, the detail features and global features between at least two image samples are respectively referred to as sample detail features and sample global features. For the specific meanings of the detail features and global features between at least two image samples, reference can be made to the previous introduction about detail features and global features, which will not be elaborated here.

[0119] S403. Based on the sample fusion feature of the sample detail feature and the sample global feature, map the sample fusion feature into a sample natural language feature through a visual mapper.

[0120] Among them, the sample fusion feature is the fusion of the sample detail feature and the sample global feature. For example, add the sample detail feature and the sample global feature to obtain the sample fusion feature.

[0121] In this application, the natural language feature obtained by mapping the sample fusion feature through a visual mapper is referred to as the sample natural language feature. Based on this, it can be known that the sample natural language feature is a natural language feature that can be understood by a large language model.

[0122] S404. Process the sample natural language feature and the sample text feature of the task indication information sample through a large language model to obtain the sample intermediate layer feature output by the intermediate layer of the large language model.

[0123] For example, splice the natural language feature and the sample text feature corresponding to the task indication information sample to obtain a sample splicing feature. Input the sample splicing feature into the large language model and obtain the sample intermediate layer feature output by the intermediate layer of the large language model.

[0124] S405. Determine the sample residual feature between the sample intermediate layer feature and the sample natural language feature and the sample text feature through a residual network module.

[0125] Among them, the specific implementation of the residual network module for determining the sample residual feature is similar to the implementation process of determining the residual feature before, which will not be elaborated here.

[0126] For example, in one implementation, the sample intermediate layer features are converted into first sample features through a residual network module, and the representation space corresponding to the first sample features is the same as the representation spaces corresponding to the sample natural language features and the sample text features (such as sample concatenated features). On this basis, the residuals between the first sample features and the sample natural language features and the sample text features are determined to obtain the sample initial residual features. Then, the sample initial residual features are subjected to transformation processing to obtain sample residual features, and the representation space corresponding to the sample residual features is the representation space corresponding to the sample intermediate layer features.

[0127] S406. Based on the sample intermediate layer features and the sample residual features, determine the predicted association relationship between the at least two image samples through a large language model.

[0128] For example, add the sample intermediate layer features and the sample residual features, input the obtained sample added features into the next network layer after the intermediate layer of the large language model, and process the sample added features through the next network layer and subsequent network layers of the large language model to output the predicted association relationship.

[0129] S407. If it is determined that the training objective has not been met based on the predicted association relationship corresponding to the at least two image samples and the sample association relationship, adjust the parameters of the residual network module, and return to execute the operation of S404 until the training objective is met.

[0130] Among them, the training objective is to minimize the gap between the predicted association relationship and the sample association relationship.

[0131] For example, the function value of the loss function can be calculated based on the predicted association relationship and the sample association relationship corresponding to each image group information. If the function value of the loss function converges or the number of training iterations reaches the set number of times, it is determined that the training objective has been achieved.

[0132] Among them, adjusting the parameters of the residual network module can be to adjust the parameters of each model module in the reference network module. For example, adjust the parameters of the conversion module in the residual network module for converting features to certain representation spaces. For ease of understanding, reference can be made to Figure 3 , adjusting the parameters of the residual network module can at least include adjusting the parameters of the linear projector 11 and the linear projector 12.

[0133] In the above embodiments, in order to further improve the accuracy of the task result determined by the large language model in the model framework for the target processing task, the present application can also train the visual encoder and the visual mapper before training the residual network module in the model framework.

[0134] Since the visual encoder and the visual mapper focus on the features of the image itself, in order to reduce the training process, the visual encoder and the visual mapper can be trained without the participation of the residual network module. Based on this, when training the visual encoder and the visual mapper, only the sub-model architecture composed of the visual encoder, the visual mapper, and the large language model in the above-mentioned model architecture needs to be trained.

[0135] For example, in a possible implementation, the third image sample can be used as the training sample, and the visual encoder and the visual mapper in the sub-model architecture are trained with the goal of minimizing the gap between the predicted description information determined by the large language model in the sub-model architecture and the actual description information corresponding to the third image sample, so as to obtain the trained visual encoder and visual mapper.

[0136] Among them, the specific training method for training the visual encoder and the visual mapper in the sub-model architecture can be any supervised training process, and the specific training process is not limited.

[0137] In any of the above embodiments of the present application, in order to better fuse the detailed features and global features output by the visual encoder and further improve the accuracy of the model architecture in processing the target processing task, in the present application, the detailed features output by the visual encoder can be first converted into second features, and the third representation space corresponding to the second features is the representation space corresponding to the global features output by the visual encoder, so that the converted second features can be feature-aligned with the global features. On this basis, the present application can add the second features and the global features to obtain the fusion features of the detailed features and the global features.

[0138] Taking an implementation manner of obtaining the fusion features as an example, the multi-modal data processing method of the present application will be introduced below. As Figure 5 , a schematic flowchart of another process of the multi-modal data processing method provided by the present application is shown. The method of this embodiment may include:

[0139] S501, obtaining N images to be processed and task indication information.

[0140] Among them, the task indication information represents that the target processing task is executed at least based on the association relationship between the N images, and N is an integer greater than or equal to 2.

[0141] S502, extracting features from the N images through the visual encoder, obtaining at least one detailed feature output by at least one intermediate layer of the visual encoder and the global feature output by the output layer of the visual encoder.

[0142] Among them, the intermediate layer of the visual encoder can be a network layer between the input layer and the output layer of the visual encoder. In practical applications, at least one intermediate layer that needs to obtain detailed features can be selected in combination with the number of network layers of the visual encoder, and there is no specific limitation.

[0143] For example, assume that the visual encoder has a total of 18 network layers, and it is necessary to obtain detailed features from two intermediate layers of the visual encoder. In order to ensure that different shallow features can be obtained more comprehensively, the detailed features output by the 6th layer and the 12th layer of the visual encoder can be obtained.

[0144] S503. For each detailed feature, convert the detailed feature into a second feature.

[0145] Among them, the third representation space corresponding to the second feature is the representation space corresponding to the global feature.

[0146] It can be understood that since the detailed features output by different intermediate layers of the visual encoder are different, the second features converted from different detailed features are also different.

[0147] Among them, there are various specific implementations for converting the detailed feature to the representation space corresponding to the global feature, which is similar to the process of converting the intermediate layer feature into the first feature. For example, the detailed feature can be converted into a second feature through a linear projector, a non - linear projector, and a neural network model.

[0148] In a possible implementation manner, for the detailed features output by different intermediate layers of the visual encoder, different feature converters can be set, and the detailed features output by the corresponding intermediate layers are converted through the feature converters. For example, at least one intermediate layer that outputs detailed features in the visual encoder is respectively connected to a feature converter, and the feature converter can convert the input detailed feature into the representation space corresponding to the global feature. The feature converter can be a linear projector or a neural network model, etc., which will not be elaborated here.

[0149] For the sake of easy understanding, combined with Figure 6 description. Figure 6 Fig. shows an example diagram of the model framework adopted by the solution of this application.

[0150] In addition to including a visual encoder, a visual mapper, a large - language model, and a residual network module, the model framework includes a feature conversion module.

[0151] The feature conversion module is used to convert at least one detailed feature output by at least one intermediate layer of the visual encoder into the representation space corresponding to the global feature output by the output layer of the visual encoder. For example Figure 6As shown, taking the case where the visual encoder has two intermediate layer output detailed features as an example, therefore, the feature conversion module may include two feature converters. In Figure 6 , taking the feature converter as a linear projector as an example. From Figure 6 , it can be seen that the detailed features output by two different intermediate layers of the visual encoder are respectively input into two linear projectors, which are linear projector 21 and linear projector 22 respectively. It is shown that projector 21 and linear projector 22 can map the input detailed features to the representation space where the global features output by the visual encoder are located.

[0152] S504, perform feature addition on each second feature and the global feature to obtain a fused feature of each detailed feature and the global feature.

[0153] Still combined with Figure 6 for illustration, the features respectively converted by linear projector 21 and linear projector 22 and the global features output by the visual encoder are all in the same representation space. Therefore, these three-way features can be added to obtain a fused feature, and this fused feature will be input into the visual mapper.

[0154] S505, map the fused feature to a natural language feature through the visual mapper.

[0155] S506, splice the natural language feature and the text feature corresponding to the task indication information, and input the spliced feature obtained by splicing into a large language model to obtain an intermediate layer feature output by the intermediate layer of the large language model.

[0156] S508, convert the intermediate layer feature into a first feature.

[0157] Among them, the first representation space corresponding to the first feature is the same as the representation space corresponding to the spliced feature obtained by splicing the natural language feature and the text feature.

[0158] S509, determine the residual between the first feature and the spliced feature to obtain an initial residual feature.

[0159] S510, perform conversion processing on the initial residual feature to obtain a residual feature.

[0160] Among them, the second representation space corresponding to the residual feature is the representation space corresponding to the intermediate layer feature.

[0161] S511, process the intermediate layer feature and the residual feature through the large language model to perform the target processing task at least based on the association relationship between the N images.

[0162] After inputting the spliced feature into the large language model, the specific processing of the large language model and the residual network module can refer to the relevant introduction in the previous embodiments. InFigure 6 The large language model and the residual network module are the same as the previous Figure 3 architecture. Therefore, regarding Figure 6 the specific implementation of the large language model processing the spliced features, obtaining the residual features through the residual network module, and the large language model processing the residual features and the intermediate layer features, reference can be made to Figure 3 the relevant introduction, which will not be elaborated here.

[0163] It can be understood that in the case of converting at least one detailed feature output by at least one intermediate layer of the visual encoder into a second feature through the feature conversion module, the present application can also train the feature conversion module, where the feature conversion module can be trained synchronously with the residual network module. For example, based on the differences between at least two image samples in the image group information, with the goal of minimizing the gap between the predicted association relationship and the sample association relationship of at least two image samples determined by the large language model, the residual network module and the feature conversion module in the model architecture are trained.

[0164] Among them, the specific process of training the residual network module and the feature conversion module can be referred to Figure 4 as shown. Only in S403, it is necessary to first use the feature conversion module to convert the sample detailed features into second sample features, fuse the second sample features and the sample global features, and input the obtained sample fused features into the visual mapper. On this basis, when adjusting the parameters of the residual network module in step S407, it is also necessary to adjust the parameters of the feature conversion module, such as adjusting the parameters of each feature converter (such as a linear projector) in the feature conversion module. In addition, after adjusting the parameters of the residual network module and the feature conversion module, it is necessary to return to execute step S402 until the training goal is met, which will not be elaborated here.

[0165] Corresponding to a multi-modal data processing method provided by the present application, the present application also provides a multi-modal data processing device. As Figure 7 shown, it shows a schematic structural diagram of a composition of the multi-modal data processing device provided by the present application. It can be seen from Figure 7 that the multi-modal data processing device may include:

[0166] A data acquisition unit 701, configured to acquire N images to be processed and task indication information, where the task indication information represents performing a target processing task based at least on the association relationship between the N images, and N is an integer greater than or equal to 2;

[0167] A feature determination unit 702, configured to determine the detailed features and global features of the N images through a visual encoder;

[0168] A feature mapping unit 703, configured to map the fused feature into a natural language feature through a visual mapper based on the fused feature of the detailed feature and the global feature;

[0169] A data processing unit 704, configured to process the natural language feature and the text feature of the task indication information through a large language model, and perform the target processing task at least based on the association relationship between the N images.

[0170] In a possible implementation, the data processing unit includes:

[0171] A feature obtaining subunit, configured to process the natural language feature and the text feature of the task indication information through a large language model, and obtain an intermediate layer feature output by an intermediate layer of the large language model;

[0172] A residual determination subunit, configured to determine a residual feature between the intermediate layer feature and the natural language feature and the text feature;

[0173] A feature processing subunit, configured to process the intermediate layer feature and the residual feature through the large language model, and perform the target processing task at least based on the association relationship between the N images.

[0174] In another possible implementation, the residual determination subunit includes:

[0175] A first conversion subunit, configured to convert the intermediate layer feature into a first feature, where a first representation space corresponding to the first feature is the same as a representation space corresponding to the natural language feature and the text feature;

[0176] An initial residual determination subunit, configured to determine a residual between the first feature and the natural language feature and the text feature, and obtain an initial residual feature;

[0177] A residual conversion subunit, configured to perform a conversion process on the initial residual feature to obtain the residual feature, where a second representation space corresponding to the residual feature is a representation space corresponding to the intermediate layer feature.

[0178] In another possible implementation, the residual determination subunit is specifically configured to determine a residual feature between the intermediate layer feature and the natural language feature and the text feature through a residual network module associated with the large language model;

[0179] Wherein, the residual network module is obtained through the following manner:

[0180] Based on the differences between at least two image samples in the image group information, train the residual network module in the model architecture composed of the visual encoder, visual mapper, large language model, and residual network module, so that the large language model can learn the correlation relationship between at least two image samples in the image group information.

[0181] In another possible implementation, the image group information includes: at least two image samples and the sample correlation relationship actually possessed by the at least two image samples;

[0182] Based on the differences between at least two image samples in the image group information, training the residual network module in the model architecture composed of the visual encoder, visual mapper, large language model, and residual network module specifically means training the residual network module in the model architecture with the goal of minimizing the gap between the predicted correlation relationship of the at least two image samples determined by the large language model and the sample correlation relationship based on the differences between at least two image samples in the image group information;

[0183] Among them, the at least two image samples include a first image sample and at least one second image sample, and the second image sample is an image sample generated based on the first image sample through data augmentation or an image generation model.

[0184] In another possible implementation, the multimodal data processing device further includes: a pre-training unit, which is used to, before training the residual network module in the model architecture, use a third image sample as a training sample, and with the goal of minimizing the gap between the predicted description information determined by the large language model in the sub-model architecture and the actual description information corresponding to the third image sample, train the visual encoder and visual mapper in the sub-model architecture to obtain the trained visual encoder and visual mapper. The sub-model architecture is composed of the visual encoder, visual mapper, and large language model in the model architecture.

[0185] In another possible implementation, the image group information further includes: a task indication information sample for indicating the determination of the correlation relationship between the at least two image samples;

[0186] The device may further include: a residual training unit, which is used to train the residual network module in the model architecture in the following manner:

[0187] Determine the sample detail features and sample global features of the at least two image samples through the visual encoder;

[0188] The sample fusion feature based on the sample detail feature and the sample global feature is mapped to a sample natural language feature by the visual mapper;

[0189] The sample natural language feature and the sample text feature of the task indication information sample are processed by the large language model to obtain a sample intermediate layer feature output by the intermediate layer of the large language model;

[0190] The sample residual feature between the sample intermediate layer feature and the sample natural language feature and the sample text feature is determined by the residual network module;

[0191] Based on the sample intermediate layer feature and the sample residual feature, the prediction association relationship between the at least two image samples is determined by the large language model;

[0192] If it is determined that the training objective has not been met based on the prediction association relationship corresponding to the at least two image samples and the sample association relationship, the parameters of the residual network module are adjusted, and the operation of processing the sample natural language feature and the sample text feature of the task indication information sample by the large language model is returned until the training objective is met, where the training objective is to minimize the gap between the prediction association relationship and the sample association relationship.

[0193] In another possible implementation, the feature determination unit includes:

[0194] The feature determination subunit is configured to extract features from the N images through the visual encoder, and obtain at least one detail feature output by at least one intermediate layer of the visual encoder and a global feature output by the output layer of the visual encoder.

[0195] In another possible implementation, the multimodal data processing device further includes:

[0196] The detail conversion unit is configured to convert the detail feature into a second feature, and the third representation space corresponding to the second feature is the representation space corresponding to the global feature;

[0197] The feature fusion unit is configured to add the second feature and the global feature to obtain a fusion feature of the detail feature and the global feature.

[0198] An electronic device is further provided in an embodiment of the present application. As Figure 8 shown, a schematic structural diagram of a composition of the electronic device is shown. The electronic device includes at least a processor 801 and a memory 802;

[0199] The processor 801 is configured to execute the multimodal data processing method described in any of the above embodiments;

[0200] The memory 802 is used to store the programs required for the processor to perform operations

[0201] It can be understood that the electronic device may further include a display unit 803 and an input unit 804.

[0202] Of course, the electronic device may also have Figure 8 more or fewer components, which are not limited herein.

[0203] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any of the multimodal data processing methods provided in the embodiments of the present application.

[0204] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device is enabled to implement any of the multimodal data processing methods provided in the embodiments of the present application.

[0205] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0206] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0207] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0208] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. A multimodal data processing method, comprising: Obtaining N images to be processed and task indication information, wherein the task indication information represents executing a target processing task based at least on an association relationship between the N images, where N is an integer greater than or equal to 2; Determine the detail features and global features of the N images through a visual encoder; Based on the fusion features of the detail features and the global features, mapping the fusion features into natural language features through a visual mapper; The natural language features and the text features of the task instruction information are processed by a large language model to perform the target processing task based at least on the association relationship between the N images.

2. The multimodal data processing method according to claim 1, wherein the processing of the natural language features and the text features of the task instruction information by a large language model to perform the target processing task based at least on the association relationship between the N images comprises: Processing the natural language features and the text features of the task instruction information through a large language model to obtain intermediate layer features output by an intermediate layer of the large language model; Determining residual features between the intermediate layer features and the natural language features and the text features; The intermediate layer features and the residual features are processed by the large language model to perform the target processing task based at least on the association relationship between the N images.

3. The multimodal data processing method according to claim 2, wherein the determining of the residual features between the intermediate layer features and the natural language features and the text features comprises: Converting the intermediate layer feature into a first feature, wherein a first representation space corresponding to the first feature is the same as a representation space corresponding to the natural language feature and the text feature; Determine the residual between the first feature and the natural language feature and the text feature to obtain an initial residual feature; The initial residual feature is converted to obtain the residual feature, and the second representation space corresponding to the residual feature is the representation space corresponding to the intermediate layer feature.

4. The multimodal data processing method according to claim 1, wherein determining the detail features and global features of the N images by a visual encoder comprises: The visual encoder is used to perform feature extraction on the N images to obtain at least one detail feature output by at least one intermediate layer of the visual encoder and a global feature output by an output layer of the visual encoder.

5. The multimodal data processing method according to claim 1 or 4, further comprising: Convert the detail feature into a second feature, wherein a third representation space corresponding to the second feature is a representation space corresponding to the global feature; The second feature and the global feature are added to obtain a fusion feature of the detail feature and the global feature.

6. The multimodal data processing method according to claim 2, wherein determining the residual features between the intermediate layer features and the natural language features and the text features comprises: Determining residual features between the intermediate layer features and the natural language features and the text features through a residual network module associated with the large language model; The residual network module is obtained in the following way: Based on the difference between at least two image samples in the image group information, the residual network module in the model architecture consisting of the visual encoder, the visual mapper, the large language model and the residual network module is trained so that the large language model can learn the association relationship between at least two image samples in the image group information.

7. The multimodal data processing method according to claim 6, wherein the image group information comprises: at least two image samples and a sample association relationship actually possessed by the at least two image samples; The method of training the residual network module in the model architecture consisting of the visual encoder, the visual mapper, the large language model and the residual network module based on the difference between at least two image samples in the image group information comprises: Based on the difference between at least two image samples in the image group information, training the residual network module in the model architecture with the goal of minimizing the gap between the predicted association relationship of the at least two image samples determined by the large language model and the sample association relationship; The at least two image samples include a first image sample and at least one second image sample, where the second image sample is an image sample generated based on the first image sample through data enhancement or an image generation model.

8. The multimodal data processing method according to claim 7, before training the residual network module in the model architecture, further comprising: Taking the third image sample as the training sample and aiming at minimizing the gap between the predicted description information determined by the large language model in the sub-model architecture and the actual description information corresponding to the third image sample, the visual encoder and the visual mapper in the sub-model architecture are trained to obtain the trained visual encoder and the visual mapper, wherein the sub-model architecture is composed of the visual encoder, the visual mapper and the large language model in the model architecture.

9. The multimodal data processing method according to claim 7 or 8, wherein the image group information further comprises: A task instruction information sample for indicating determining an association relationship between the at least two image samples; The method of training the residual network module in the model architecture based on the difference between at least two image samples in the image group information with the goal of minimizing the gap between the predicted association relationship corresponding to the at least two image samples determined by the large language model and the sample association relationship comprises: Determining sample detail features and sample global features of the at least two image samples by the visual encoder; Based on the sample fusion features of the sample detail features and the sample global features, mapping the sample fusion features into sample natural language features through the visual mapper; Processing the sample natural language features and the sample text features of the task instruction information sample by the large language model to obtain sample intermediate layer features output by the intermediate layer of the large language model; Determine, by means of the residual network module, a sample residual feature between the sample intermediate layer feature, the sample natural language feature, and the sample text feature; Based on the sample intermediate layer features and the sample residual features, determining the predicted association relationship between the at least two image samples through the large language model; If it is determined based on the predicted association relationship corresponding to the at least two image samples and the sample association relationship that the training objective has not been met, adjust the parameters of the residual network module, and return to executing the operation of processing the sample natural language features and the sample text features of the task instruction information sample through the large language model until the training objective is met, and the training objective is to minimize the gap between the predicted association relationship and the sample association relationship.

10. A multimodal data processing device, comprising: A data acquisition unit, used to obtain N images to be processed and task indication information, wherein the task indication information represents the execution of a target processing task based at least on the association relationship between the N images, where N is an integer greater than or equal to 2; A feature determination unit, configured to determine detail features and global features of the N images through a visual encoder; A feature mapping unit, configured to map the fused features of the detail features and the global features into natural language features through a visual mapper based on the fused features of the detail features and the global features; A data processing unit is used to process the natural language features and the text features of the task instruction information through a large language model to perform the target processing task based at least on the association relationship between the N images.

Citation Information

Cited By

  • Long text processing method and device, equipment and medium

    CN121960382A