Dialogue Generation Method and Device Based on Multimodal Large Language Model

By building a multimodal large language model based on image feature enhancement, using multi-layer aggregation module and in-module and inter-module enhancement modules, the problem of insufficient utilization of visual features in the multimodal large language model is solved, visual information perception and image key area attention capabilities are improved, and performance is achieved.

CN119938874BActive Publication Date: 2025-07-18XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436346.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing multimodal large language model does not fully utilize the potential of visual features and does not consider the correlation between intramodel and intermodal, resulting in complex answering questions accurately.

Method used

A multimodal large language model based on image feature enhancement is constructed. Through multi-layer aggregation module and in-module and inter-module enhancement module, the multi-modal large language model is fine-tuned, including pre-trained image encoder, text encoder, word segmenter, multi-layer aggregation module, in-module and inter-module enhancement module, multi-layer perceptron module and pre-trained large language model to enhance image features and promote visual feature interaction.

Benefits of technology

The visual information perception ability and image key area attention ability of the multimodal large language model are significantly improved, and the performance is continuously and greatly improved, demonstrating scalability and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938874B_ABST
    Figure CN119938874B_ABST
Patent Text Reader

Abstract

The present invention discloses a dialogue generation method and device based on a multimodal large language model, which relates to the field of dialogue generation and includes: obtaining a query statement and an image and inputting them into a fine-tuned multimodal large language model. The image is input into a pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through a multi-layer aggregation module to extract low-level image features and high-level image features; the query statement is input into a text encoder to obtain text features; the above features are input into an intra-modal and inter-modal enhancement module for enhancement. The enhanced image features are concatenated along the channels and then projected through a multi-layer perceptron module to obtain visual tokens; the query statement is input into a pre-trained tokenizer for tokenization to obtain text tokens; the visual tokens and text tokens are input into a trained large language model to generate an answer statement. The present invention solves the problem that existing MLLMs do not consider intra-modal and inter-modal correlations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal dialogue generation, and specifically relates to a dialogue generation method and device based on a multimodal large language model. Background Art

[0002] The emergence of large language models (LLMs) has catalyzed the remarkable development of multimodal large language models (MLLMs, integrated with visual perception capabilities). These enhanced capabilities have broadened the applicability of MLLMs, facilitating the emergence of new technologies and the establishment of benchmarks.

[0003] Generally, the architecture of an MLLM contains three main components:

[0004] 1) A pre-trained visual encoder designed to extract features from input images.

[0005] 2) A connector that projects image features from the visual space to the text space.

[0006] 3) A pre-trained large language model that generates responses based on visual inputs and accompanying instructions.

[0007] Although a large amount of research has been dedicated to optimizing the design of the connector and tuning the large language model, the research on fully leveraging the potential of visual features is relatively limited. Most existing methods involve feeding visual features extracted from specific layers of the visual encoder into the connector. Then the output visual tokens are concatenated with text tokens and fed into the large language model to generate responses. However, most existing MLLMs simply pass the output features of the image encoder to the connector without considering intra-modal and inter-modal correlations, thus complicating the accurate answering of questions. Summary of the Invention

[0008] The purpose of this application is to propose a dialogue generation method and device based on a multimodal large language model for the above-mentioned technical problems.

[0009] In a first aspect, the present invention provides a dialogue generation method based on a multimodal large language model, including the following steps:

[0010] Construct and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model, where the multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained tokenizer, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model;

[0011] Obtain the query statement and its corresponding image and input them into the fine-tuned multimodal large language model. The image is input into the pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through a multi-layer aggregation module to extract low-level image features and high-level image features; the query statement is input into the text encoder to obtain text features; the text features, low-level image features, high-level image features, and selected image features are input into the in-module and inter-module enhancement module for enhancement to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features; the enhanced low-level image features, enhanced high-level image features, and enhanced selected image features are concatenated along the channel and then projected through a multi-layer perceptron module to obtain visual tokens; the query statement is input into the pre-trained tokenizer for tokenization to obtain text tokens; the visual tokens and text tokens are input into the trained large language model to generate an answer statement.

[0012] Preferably, the multi-scale encoded features are represented as , where represents the encoded features output by the i-th encoding layer in the image encoder, represents the set of real numbers, represents the height of the encoded features, represents the width of the encoded features, represents the dimension of the encoded features, represents the total number of encoding layers in the image encoder; in the multi-layer aggregation module, first, the encoded features output by the first 1 / 2N encoding layers and the last 1 / 2N encoding layers in the multi-scale encoded features are averaged respectively to obtain low-level query features and high-level query features , represents the number of visual tokens. The low-level query features are subjected to attention calculation with the first multi-scale encoded features stacked by the encoded features output by the first 1 / 2N encoding layers in the multi-scale encoded features to obtain low-level image features, as shown in the following formula:

[0013] ;

[0014] where represents the first multi-scale encoded features stacked by the encoded features output by the first 1 / 2N encoding layers in the multi-scale encoded features, , , represent the low-level query weight matrix, low-level key weight matrix, and low-level value weight matrix respectively, represents the dimension of the first multi-scale encoded features, represents the low-level image features, represents the Softmax function, and T represents the transposed matrix;

[0015] Perform attention calculation on the second multi-scale encoded feature, which is formed by stacking the encoded features output from the latter 1 / 2N encoding layers in the multi-scale encoded features and the high-level query features, to obtain high-level image features, as shown in the following formula:

[0016] ;

[0017] Wherein, represents the second multi-scale encoded feature formed by stacking the encoded features output from the latter 1 / 2N encoding layers in the multi-scale encoded features, , , respectively represent the high-level query weight matrix, the high-level key weight matrix, and the high-level value weight matrix, represents the dimension of the second multi-scale encoded feature, represents the high-level image feature.

[0018] Preferably, in the intra-module and inter-module enhancement module, perform intra-module enhancement and inter-module enhancement on the low-level image feature, the high-level image feature, and the selected image feature respectively in combination with the text feature, as shown in the following formula:

[0019] ;

[0020] ;

[0021] Wherein, , , , , , respectively represent the first correlation weight matrix, the second correlation weight matrix, the third correlation weight matrix, the fourth correlation weight matrix, the fifth correlation weight matrix, and the sixth correlation weight matrix, represents the Sigmoid function, and T represents the transpose matrix;

[0022] Take the average value using the following formula to obtain the enhanced low-level image feature, the enhanced high-level image feature, and the enhanced selected image feature:

[0023] ;

[0024] Wherein, represents the low-level image feature, the high-level image feature, or the selected image feature; if is the low-level image feature, then represents the global average low-level image feature obtained by taking the average of the low-level image feature, represents the low-level image feature after intra-module enhancement, represents the low-level image feature after inter-module enhancement, represents enhanced low-level image features; if is a high-level image feature, then represents the global average high-level image feature obtained by averaging the high-level image features, represents the high-level image feature after in-module enhancement, represents the high-level image feature after inter-module enhancement, represents enhanced high-level image features; if is a selected image feature, then represents the global average selected image feature obtained by averaging the selected image features, represents the low-level image feature after in-module enhancement, the high-level image feature after in-module enhancement, or the selected image feature after in-module enhancement, represents the low-level image feature after inter-module enhancement, the high-level image feature after inter-module enhancement, or the selected image feature after inter-module enhancement, represents enhanced selected image features.

[0025] Preferably, the image encoder includes a ViT model, the large language model includes an LLaMA3-8B model, and the multi-layer perceptron module includes two consecutive multi-layer perceptrons.

[0026] Preferably, the selected image feature is the encoded feature output by the last encoding layer of the image encoder.

[0027] Preferably, during the fine-tuning process of the multi-modal large language model, the parameters of the pre-trained image encoder, the pre-trained text encoder, and the pre-trained tokenizer are frozen, and the parameters of the multi-layer aggregation module, the in-module and inter-module enhancement modules, the multi-layer perceptron, and the pre-trained large language model are fine-tuned.

[0028] In a second aspect, the present invention provides a dialogue generation device based on a multi-modal large language model, including:

[0029] a model construction module configured to construct and fine-tune a multi-modal large language model based on image feature enhancement to obtain a fine-tuned multi-modal large language model, where the multi-modal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained tokenizer, a multi-layer aggregation module, an in-module and inter-module enhancement module, a multi-layer perceptron module, and a pre-trained large language model;

[0030] A generation module, which is configured to obtain a query statement and its corresponding image and input them into a fine-tuned multimodal large language model. The image is input into a pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through a multi-layer aggregation module to extract low-level image features and high-level image features. The query statement is input into a text encoder to obtain text features. The text features, low-level image features, high-level image features, and selected image features are input into an intra-modal and inter-modal enhancement module for enhancement to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features. The enhanced low-level image features, enhanced high-level image features, and enhanced selected image features are concatenated along the channel and then projected through a multi-layer perceptron module to obtain visual tokens. The query statement is input into a pre-trained tokenizer for tokenization to obtain text tokens. The visual tokens and text tokens are input into a trained large language model to generate an answer statement.

[0031] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0033] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] (1) The dialogue generation method based on a multimodal large language model proposed by the present invention constructs a multimodal large language model based on image enhancement, and effectively aggregates high-level image features and low-level feature features by using a multi-layer aggregation module, thereby enhancing the ability of the MLLM to perceive multi-level detailed visual information.

[0036] (2) The multimodal large language model constructed by the dialogue generation method based on a multimodal large language model proposed by the present invention uses an intra-modal and inter-modal enhancement module to promote intra-modal and inter-modal interactions of visual features, and significantly improves the ability of the MLLM to focus on key regions of images.

[0037] (3) The dialogue generation method based on the multi-modal large language model proposed by the present invention fine-tunes the parameters of the multi-layer aggregation module, the intra-modal and inter-modal enhancement module, the multi-layer perceptron, and the pre-trained large language model, continuously and significantly improves the performance in various benchmark tests, and can demonstrate its scalability, versatility, and effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a schematic flowchart of the dialogue generation method based on the multi-modal large language model for the embodiments of the present application;

[0040] Figure 2 It is a schematic diagram of the multi-modal large language model of the dialogue generation method based on the multi-modal large language model for the embodiments of the present application;

[0041] Figure 3 It is a schematic diagram of the image features extracted from different layers in CLIP;

[0042] Figure 4 It is a schematic diagram of the attention mapping diagram between the image features extracted from CLIP and different query statements;

[0043] Figure 5 It is a schematic diagram of the multi-layer aggregation module of the dialogue generation method based on the multi-modal large language model for the embodiments of the present application;

[0044] Figure 6 It is a schematic diagram of the intra-modal and inter-modal enhancement module of the dialogue generation method based on the multi-modal large language model for the embodiments of the present application;

[0045] Figure 7 It is a schematic diagram of the specific calculation process of intra-modal enhancement and inter-modal enhancement of the dialogue generation method based on the multi-modal large language model for the embodiments of the present application;

[0046] Figure 8 It is a schematic diagram of the dialogue generation device based on the multi-modal large language model for the embodiments of the present application;

[0047] Figure 9 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0049] Figure 1 A dialogue generation method based on a multimodal large language model provided by an embodiment of the present application is shown, including the following steps:

[0050] S1. Construct and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model. The multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained tokenizer, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model.

[0051] In a specific embodiment, the image encoder includes a ViT model, and the large language model includes an LLaMA3-8B model.

[0052] In a specific embodiment, the multi-layer perceptron module includes two layers of multi-layer perceptrons.

[0053] In a specific embodiment, during the fine-tuning process of the multimodal large language model, the parameters of the pre-trained image encoder, the pre-trained text encoder, and the pre-trained tokenizer are frozen, and the parameters of the multi-layer aggregation module, the intra-modal and inter-modal enhancement module, the multi-layer perceptron, and the pre-trained large language model are fine-tuned.

[0054] Specifically, referring to Figure 2 , an embodiment of the present application proposes a multimodal large language model based on image feature enhancement. Based on the pre-trained large language model, this multimodal large language model adopts a multimodal input method composed of text and images and enhances the image features, so that the pre-trained large language model can generate more accurate answer sentences.

[0055] First, the input image is encoded into multi-scale encoded features by an image encoder based on the ViT model. This encoding process is expressed as:

[0056] ;

[0057] Among them, represents the function corresponding to the image encoder, represents the encoded feature output by the i-th encoding layer in the image encoder, N is the total number of layers of the image encoder, and d is the dimension of the encoded feature.

[0058] Next, the extracted multi-scale encoded features are processed by the proposed Multi-level Aggregation Module (MAM) to derive high-level image features and low-level image features:

[0059] ;

[0060] where are the high-level image feature and the low-level image feature respectively, is the number of image tokens, and the function represents the function corresponding to the Multi-level Aggregation Module (MAM).

[0061] Take the encoded features output by the last encoding layer in the image encoder as the selected image features, input the query statement into the text encoder to obtain text features; combine the text features and enhance the low-level image features, high-level image features, and selected image features through the Intra-modal and Inter-modal Enhancement (IEM) module to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features.

[0062] Then concatenate the enhanced image features along the channel dimension , and project them using a multi-layer perceptron module:

[0063] ;

[0064] where represents the concatenation operation along the channel dimension d, is a multi-layer perceptron module composed of two layers of multi-layer perceptrons (MLP), with an input dimension of 3d and an output dimension of D, is a visual token.

[0065] Use a pre-trained tokenizer to tokenize the query statement T to obtain text tokens . Finally, input the visual tokens and text tokens into a pre-trained large language model (LLM) to generate an answer statement:

[0066] ;

[0067] where represents the answer statement, represents the pre-trained large language model.

[0068] During the fine-tuning process of the multi-modal large language model, the parameters of the multi-layer aggregation module, the intra-modal and inter-modal enhancement module, the multi-layer perceptron, and the pre-trained large language model are mainly adjusted, while the parameters of the pre-trained image encoder, the pre-trained text encoder, and the pre-trained tokenizer are frozen.

[0069] S2. Obtain the query statement and its corresponding image and input them into the fine-tuned multi-modal large language model. The image is input into the pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through the multi-layer aggregation module to extract low-level image features and high-level image features. The query statement is input into the text encoder to obtain text features. The text features, low-level image features, high-level image features, and selected image features are input into the intra-modal and inter-modal enhancement module for enhancement to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features. The enhanced low-level image features, enhanced high-level image features, and enhanced selected image features are concatenated along the channel and then projected through the multi-layer perceptron module to obtain visual tokens. The query statement is input into the pre-trained tokenizer for tokenization to obtain text tokens. The visual tokens and text tokens are input into the trained large language model to generate an answer statement.

[0070] In a specific embodiment, the selected image feature is the encoded feature output by the last encoding layer of the image encoder.

[0071] In a specific embodiment, the multi-scale encoded features are represented as , where represents the encoded feature output by the i-th encoding layer in the image encoder, represents the set of real numbers, represents the height of the encoded feature, represents the width of the encoded feature, represents the dimension of the encoded feature, represents the total number of encoding layers in the image encoder; in the multi-layer aggregation module, first, the average of the encoded features output by the first 1 / 2N encoding layers and the last 1 / 2N encoding layers in the multi-scale encoded features is taken respectively to obtain the low-level query feature and the high-level query feature , represents the number of visual tokens. The low-level query feature is subjected to attention calculation with the first multi-scale encoded feature stacked by the encoded features output by the first 1 / 2N encoding layers in the multi-scale encoded features to obtain the low-level image feature, as shown in the following formula:

[0072] ;

[0073] Where denotes the first multi-scale encoded feature stacked by the encoded features output from the first 1 / 2N encoded layers in the multi-scale encoded features, , , respectively denote the low-level query weight matrix, the low-level key weight matrix, and the low-level value weight matrix, denotes the dimension of the first multi-scale encoded feature, denotes the low-level image feature, denotes the Softmax function, and T denotes the transposed matrix;

[0074] Perform attention calculation on the high-level query feature and the second multi-scale encoded feature stacked by the encoded features output from the last 1 / 2N encoded layers in the multi-scale encoded features to obtain the high-level image feature, as shown in the following formula:

[0075] ;

[0076] where, denotes the second multi-scale encoded feature stacked by the encoded features output from the last 1 / 2N encoded layers in the multi-scale encoded features, , , respectively denote the high-level query weight matrix, the high-level key weight matrix, and the high-level value weight matrix, denotes the dimension of the second multi-scale encoded feature, denotes the high-level image feature.

[0077] As is well known, features from different layers of a neural network contain different levels of information. As Figure 3 shown, low-level features capture fine details, such as edges, while high-level features contain more abstract and semantic information, which is crucial for semantic perception of images. Secondly, both intra-module and inter-module interactions can significantly improve the quality of image features by filtering out irrelevant information and amplifying relevant details. Figure 4 illustrates the attention mapping between image features and various query statements. In particular, the attention map between the image features and the image CLS token highlights the foreground objects, while the attention map between the image features and the query statement focuses on the objects referred to in the query statement. For example, when the query statement is "What color is the cat in the image?", the attention map will focus on the area containing the cat. Similarly, when the query statement is "How many dogs are there in the image?", the attention map will highlight the area containing the dogs. In addition, the attention map between the image CLS token and the image features shows that higher attention scores are concentrated in the areas of cats and dogs, while lower attention scores appear in the background areas. This phenomenon is beneficial because it enables the model to focus on the key areas to provide accurate answers.

[0078] The following introduces the specific calculation processes of the multi-layer aggregation module and the intra-module and inter-module enhancement modules. Refer to Figure 5 , given the multi-scale encoded features , where N is the total number of layers of the image encoder, the multi-layer aggregation (MAM) module aims to extract high-level image features and low-level image features . First, by averaging the encoded features of the first half and the second half of the layers in the multi-scale encoded features, the low-level query feature and the high-level query feature are obtained respectively. This creates representative query features that summarize the information from the low-level and high-level respectively.

[0079] Next, calculate the attention maps between the low-level query feature or the high-level query feature and the first multi-scale encoded feature or the second multi-scale encoded feature stacked by its corresponding multiple encoded features, to identify which layers are relevant to the low-level query feature or the high-level query feature. This is crucial for focusing on the encoded layers in the image features. This process enhances the original query features by emphasizing the important encoded layers identified by the attention maps. The low-level query feature and the high-level query feature are collectively referred to as query features, and the first multi-scale encoded feature and the second multi-scale encoded feature are collectively referred to as multi-layer encoded features.

[0080] In a specific embodiment, in the intra-module and inter-module enhancement module, the intra-module enhancement and inter-module enhancement are respectively performed on the low-level image feature, the high-level image feature, and the selected image feature by combining the text features, as shown in the following formula:

[0081] ;

[0082] ;

[0083] where, , , , , , respectively represent the first correlation weight matrix, the second correlation weight matrix, the third correlation weight matrix, the fourth correlation weight matrix, the fifth correlation weight matrix, and the sixth correlation weight matrix, represents the Sigmoid function, and T represents the transpose matrix;

[0084] The following formula is used to take the average value to obtain the enhanced low-level image feature, the enhanced high-level image feature, and the enhanced selected image feature:

[0085] ;

[0086] where, Represents low-level image features, high-level image features, or selected image features; if is low-level image features, then represents the global average low-level image features obtained by averaging the low-level image features, represents the low-level image features after in-module enhancement, represents the low-level image features after inter-module enhancement, represents the enhanced low-level image features; if is high-level image features, then represents the global average high-level image features obtained by averaging the high-level image features, represents the high-level image features after in-module enhancement, represents the high-level image features after inter-module enhancement, represents the enhanced high-level image features; if is selected image features, then represents the global average selected image features obtained by averaging the selected image features, represents the low-level image features after in-module enhancement, the high-level image features after in-module enhancement, or the selected image features after in-module enhancement, represents the low-level image features after inter-module enhancement, the high-level image features after inter-module enhancement, or the selected image features after inter-module enhancement, represents the enhanced selected image features.

[0087] Specifically, referring to Figure 6 , the in-module and inter-module enhancement (IEM) module aims to promote in-module and inter-module interactions between different levels of image features to enhance the image features. These levels include high-level image features, low-level image features, and selected image features. For simplicity, high-level image features will be taken as an example, and the operations for other level features are similar. High-level image features, low-level image features, and selected image features are collectively referred to as image features, global average high-level image features, global average low-level image features, and global average selected image features are collectively referred to as global average image features, and enhanced high-level image features, enhanced low-level image features, and enhanced selected image features are collectively referred to as enhanced image features.

[0088] As Figure 6 shown, the input includes high-level image features from the MAM , text features (i.e., the text features extracted after encoding the query statement using the text encoder) and global average high-level image features (i.e., the global average high-level image features obtained by averaging ).

[0089] As Figure 7As shown, first, the correlation mapping between global tokens and visual tokens is calculated using matrix multiplication. Then, the visual tokens are enhanced through element-wise multiplication based on the correlation mapping. This process can be understood as focusing on significant features by considering image-to-image (intra-modal) and image-to-text (inter-modal) correlations. The intra-modal and inter-modal enhancement formulas are as follows:

[0090] ;

[0091] ;

[0092] where, if is the high-level image feature, then represents the global average high-level image feature obtained by averaging the high-level image features, represents the high-level image feature after intra-modal enhancement, represents the high-level image feature after inter-modal enhancement;

[0093] The average of the enhanced feature and the feature before enhancement is taken to obtain the enhanced high-level image feature, as shown in the following formula:

[0094] ;

[0095] where, represents the enhanced high-level image feature.

[0096] Further deploying and applying the fine-tuned multi-modal large language model (denoted as vMLLM) can highlight the capabilities of text recognition, visual perception, hallucination, mathematical knowledge, chart understanding, map understanding, and poster understanding.

[0097] Next, the fine-tuned multi-modal large language model proposed in the embodiments of this application will be compared with other advanced technologies.

[0098] Table 1 shows the comparison results of the fine-tuned multi-modal large language model proposed in the embodiments of the present application with the state-of-the-art technologies, which include DenseConnector, VILA, and MG-LLaVA. It is worth noting that the fine-tuned multi-modal large language model (vMLLM) proposed in the embodiments of the present application is based on the 1.5 architecture developed from the powerful LLaVA. The experimental results show that vMLLM has achieved significant performance improvements compared to LLaVA-1.5 in various benchmark tests. Specifically, when using the same LLaMA3-8B language model, the performance of vMLLM on the MMB, MM-Vet, and MathVista benchmark tests has increased by 8.2%, 8.1%, and 13.4% respectively. These enhancements highlight the effectiveness of the method of the present invention in enhancing visual features. In addition, the performance of vMLLM is consistently better than the most recent state-of-the-art models. For example, compared with DenseConnector, the performance of vMLLM on the MMB, MMVet, and MathVista benchmark tests has increased by 2.7%, 8.8%, and 12.4% respectively.

[0099] Table 1:

[0100]

[0101] Further referring to Figure 8 and as an implementation of the methods shown in the above figures, the present application provides an embodiment of a dialogue generation device based on a multi-modal large language model. This device embodiment corresponds to Figure 1 the method embodiment shown and can be specifically applied to various electronic devices.

[0102] The embodiments of the present application provide a dialogue generation device based on a multi-modal large language model, including:

[0103] A model construction module 1, configured to construct and fine-tune a multi-modal large language model based on enhanced image features to obtain a fine-tuned multi-modal large language model. The multi-modal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained tokenizer, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model;

[0104] A generation module 2, configured to obtain a query statement and its corresponding image and input them into a fine-tuned multimodal large language model. The image is input into a pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through a multi-layer aggregation module to extract low-level image features and high-level image features. The query statement is input into a text encoder to obtain text features. The text features, low-level image features, high-level image features, and selected image features are input into an intra-modal and inter-modal enhancement module for enhancement to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features. The enhanced low-level image features, enhanced high-level image features, and enhanced selected image features are concatenated along the channel and then projected through a multi-layer perceptron module to obtain visual tokens. The query statement is input into a pre-trained tokenizer for tokenization to obtain text tokens. The visual tokens and text tokens are input into a trained large language model to generate an answer statement.

[0105] Figure 9 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. As Figure 9 shown, the electronic device of this embodiment includes: a processor 901 and a memory 902. Among them, the memory 902 is used to store computer execution instructions. The processor 901 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0106] Optionally, the memory 902 can be either independent or integrated with the processor 901.

[0107] When the memory 902 is independently provided, the electronic device further includes a bus 903 for connecting the memory 902 and the processor 901.

[0108] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 901 executes the computer execution instructions, the above method is implemented.

[0109] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 901, the above method is implemented.

[0110] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.

[0111] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0112] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0113] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 901 to execute some steps of the methods in various embodiments of the present application.

[0114] It should be understood that the above processor 901 can be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor 901 can also be any conventional processor 901, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 901 or by the combination of the hardware and software modules in the processor 901.

[0115] The memory 902 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.

[0116] The bus 903 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 903 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 903 in the accompanying drawings of this application is not limited to only one bus 903 or one type of bus 903.

[0117] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0118] An exemplary storage medium is coupled to the processor 901, so that the processor 901 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 901. The processor 901 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 901 and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0119] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disks, or optical discs and other media that can store program codes.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dialogue generation method based on a multimodal large language model, characterized in that, Including the following steps: Construct and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model. The multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained tokenizer, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model; Obtain a query statement and its corresponding image and input them into the fine-tuned multimodal large language model. The image is input into the pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through the multi-layer aggregation module to extract low-level image features and high-level image features. The query statement is input into the text encoder to obtain text features. The text features, low-level image features, high-level image features, and selected image features are input into the intra-modal and inter-modal enhancement module for enhancement to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features. In the intra-modal and inter-modal enhancement module, the text features are combined to perform intra-modal enhancement and inter-modal enhancement on the low-level image features, high-level image features, and selected image features respectively, as shown in the following formula: Among them, respectively represent the first correlation weight matrix, the second correlation weight matrix, the third correlation weight matrix, the fourth correlation weight matrix, the fifth correlation weight matrix, and the sixth correlation weight matrix. Sigmoid represents the Sigmoid function, and T represents the transposed matrix; Use the following formula to take the average value to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features: Among them, F b represents the low-level image feature, high-level image feature, or selected image feature; if F b is the low-level image feature, then represents the global average low-level image feature obtained by averaging the low-level image feature, represents the low-level image feature after in-module enhancement, represents the low-level image feature after inter-module enhancement, F′ b represents the enhanced low-level image feature; if F b is the high-level image feature, then represents the global average high-level image feature obtained by averaging the high-level image feature, represents the high-level image feature after in-module enhancement, represents the high-level image feature after inter-module enhancement, F′ b represents the enhanced high-level image feature; if F b is the selected image feature, then represents the global average selected image feature obtained by averaging the selected image feature, represents the low-level image feature after in-module enhancement, high-level image feature after in-module enhancement, or selected image feature after in-module enhancement, represents the low-level image feature after inter-module enhancement, high-level image feature after inter-module enhancement, or selected image feature after inter-module enhancement, F′ b represents the enhanced selected image feature; the enhanced low-level image feature, enhanced high-level image feature, and enhanced selected image feature are concatenated along the channel and then projected through the multi-layer perceptron module to obtain visual tokens; the query statement is input into the pre-trained tokenizer for tokenization to obtain text tokens; the visual tokens and text tokens are input into the trained large language model to generate an answer statement.

2. The dialogue generation method based on a multimodal large language model according to claim 1, wherein, The multi-scale encoded feature representation is where represents the encoded feature output by the i-th encoded layer in the image encoder, represents the set of real numbers, h represents the height of the encoded feature, w represents the width of the encoded feature, d represents the dimension of the encoded feature, and N represents the total number of encoded layers in the image encoder; in the multi-layer aggregation module, first, the encoded features output by the first 1 / 2N encoded layers and the encoded features output by the last 1 / 2N encoded layers in the multi-scale encoded features are averaged respectively to obtain the low-level query feature and the high-level query feature N V represents the number of visual tokens. The low-level query feature is subjected to attention calculation with the first multi-scale encoded feature stacked by the encoded features output by the first 1 / 2N encoded layers in the multi-scale encoded features to obtain the low-level image feature, as shown in the following formula: Among them, represents the first multi-scale encoded feature stacked by the encoded features output by the first 1 / 2N encoded layers in the multi-scale encoded features, respectively represent the low-level query weight matrix, the low-level key weight matrix, and the low-level value weight matrix, d l represents the dimension of the first multi-scale encoded feature, F l represents the low-level image feature, Softmax represents the Softmax function, and T represents the transpose matrix; Perform attention calculation on the high-level query features and the second multi-scale encoded features stacked by the encoded features output by the last 1 / 2N encoding layers in the multi-scale encoded features to obtain the high-level image features, as shown in the following formula: Among them, represents a second multi-scale encoded feature formed by stacking encoded features output from the latter 1 / 2N encoded layers in the multi-scale encoded features, respectively represent the high-level query weight matrix, the high-level key weight matrix, and the high-level value weight matrix, d h represents the dimension of the second multi-scale encoded feature, F h represents the high-level image feature.

3. The dialogue generation method based on a multimodal large language model according to claim 1, wherein, The image encoder includes a ViT model, the large language model includes an LLaMA3-8B model, and the multi-layer perceptron module includes two consecutive multi-layer perceptrons.

4. The dialogue generation method based on a multimodal large language model according to claim 1, wherein, The selected image features are the encoded features output by the last encoding layer of the image encoder.

5. The dialogue generation method based on a multimodal large language model according to claim 1, wherein During the fine-tuning process of the multimodal large language model, the parameters of the pre-trained image encoder, pre-trained text encoder, and pre-trained tokenizer are frozen, and the parameters of the multi-layer aggregation module, intra-modal and inter-modal enhancement module, multi-layer perceptron, and pre-trained large language model are fine-tuned.

6. A dialogue generation device based on a multimodal large language model, which is used to implement the dialogue generation method based on a multimodal large language model according to any one of claims 1-5, and is characterized in that Including: A model construction module configured to construct and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model. The multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained tokenizer, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model; A generation module, configured to obtain a query statement and its corresponding image and input them into the fine-tuned multimodal large language model. The image is input into the pre-trained image encoder to obtain multi-scale encoded features and selected image features. The multi-scale encoded features pass through the multi-layer aggregation module to extract low-level image features and high-level image features; the query statement is input into the text encoder to obtain text features; the text features, low-level image features, high-level image features, and selected image features are input into the intra-modal and inter-modal enhancement module for enhancement to obtain enhanced low-level image features, enhanced high-level image features, and enhanced selected image features. The enhanced low-level image features, enhanced high-level image features, and enhanced selected image features are concatenated along the channel and then projected through the multi-layer perceptron module to obtain visual tokens; the query statement is input into the pre-trained tokenizer for tokenization to obtain text tokens. The visual tokens and text tokens are input into the trained large language model to generate an answer statement.

7. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Memory grounded conversational reasoning and question answering for assistant systems

    CN114072832A

  • Pedestrian attribute cross-modal alignment method based on complete attribute identification enhancement

    WO2024114185A1