Dialogue generation method and device based on multi-modal large language model
By building a multimodal large language model based on image feature enhancement, using multi-layer aggregation and in-module and inter-module enhancement modules, the problem that existing models fail to make full use of visual features is solved, and more accurate and efficient dialogue generation is achieved.
Patent Information
- Application Number
- CN202510436346.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
When using visual features, existing multimodal large language models fail to fully consider the correlation between in-model and inter-modes, resulting in complex answering questions accurately.
A multimodal large language model based on image feature enhancement is constructed, high-level and low-level image features are aggregated through multi-layer aggregation modules, and the interaction of visual features is promoted in-module and inter-module enhancement modules is promoted, and visual symbols are finally generated through multi-layer perceptron modules.
The model's ability to pay attention to key areas of the image is significantly improved, the ability to perceive multi-level detailed visual information is enhanced, and the performance of dialogue generation is continuously and greatly improved.
Smart Images

Figure CN119938874A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal dialogue generation, and in particular to a dialogue generation method and device based on a multimodal large language model. Background Art
[0002] The emergence of Large Language Models (LLMs) has catalyzed the significant development of Multimodal Large Language Models (MLLMs, which integrate visual perception capabilities). These enhancements have broadened the applicability of MLLMs and promoted the emergence of new techniques and the establishment of benchmarks.
[0003] Generally speaking, the architecture of MLLM contains three main components:
[0004] 1) A pre-trained visual encoder designed to extract features from the input image.
[0005] 2) A connector that projects image features from visual space to text space.
[0006] 3) A pre-trained large language model that generates responses based on the visual input and accompanying instructions.
[0007] While a large amount of research has been devoted to optimizing the design of connectors and tuning large language models, relatively limited research has been done on fully leveraging the potential of visual features. Most existing approaches involve inputting visual features extracted from a specific layer of a visual encoder into a connector. The output visual tokens are then concatenated with text tokens and input into a large language model to generate a response. However, most existing MLLMs simply pass the output features of the image encoder to the connector without considering intra- and inter-modal correlations, complicating accurate question answering. Summary of the invention
[0008] The purpose of this application is to propose a dialogue generation method and device based on a multimodal large language model to address the above-mentioned technical problems.
[0009] In a first aspect, the present invention provides a method for generating a dialogue based on a multimodal large language model, comprising the following steps:
[0010] Constructing and fine-tuning a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model, wherein the multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained word segmenter, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model;
[0011] The query statement and its corresponding image are obtained and input into a fine-tuned multimodal large language model. The image is input into a pre-trained image encoder to obtain multi-scale coding features and selected image features. The multi-scale coding features are extracted through a multi-layer aggregation module to obtain low-level image features and high-level image features; the query statement is input into a text encoder to obtain text features; the text features, low-level image features, high-level image features and selected image features are input into intra-modal and inter-modal enhancement modules for enhancement to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features; the enhanced low-level image features, enhanced high-level image features and enhanced selected image features are connected along the channel and projected through a multi-layer perceptron module to obtain visual tokens; the query statement is input into a pre-trained word segmenter for word segmentation to obtain text tokens; the visual tokens and text tokens are input into the trained large language model to generate an answer statement.
[0012] As a preference, the multi-scale coding feature is expressed as ,in, represents the encoding features output by the i-th encoding layer in the image encoder, represents the set of real numbers, represents the height of the encoded feature, represents the width of the encoded feature, represents the dimension of the encoded features, Represents the total number of coding layers in the image encoder; in the multi-layer aggregation module, the coding features output by the first 1 / 2N coding layers and the coding features output by the last 1 / 2N coding layers in the multi-scale coding features are averaged to obtain the low-level query features. and advanced query features , Represents the number of visual symbols. The low-level query feature is stacked with the first multi-scale coding feature output by the first 1 / 2N coding layers in the multi-scale coding feature to perform attention calculation to obtain the low-level image feature, as shown in the following formula:
[0013] ;
[0014] in, Represents the first multi-scale coding feature formed by stacking the coding features output by the first 1 / 2N coding layers in the multi-scale coding feature. , , Respectively represent the low-level query weight matrix, low-level key weight matrix, and low-level value weight matrix, represents the dimension of the first multi-scale encoded feature, represents low-level image features, represents the Softmax function, T represents the transposed matrix;
[0015] The high-level query feature is stacked with the encoding features output by the last 1 / 2N encoding layers in the multi-scale encoding feature to perform attention calculation to obtain the high-level image feature, as shown in the following formula:
[0016] ;
[0017] in, Represents the second multi-scale coding feature formed by stacking the coding features output by the last 1 / 2N coding layers in the multi-scale coding feature. , , They represent the advanced query weight matrix, advanced key weight matrix, and advanced value weight matrix respectively. represents the dimension of the second multi-scale encoded feature, Represents high-level image features.
[0018] Preferably, in the intra-module and inter-module enhancement modules, the low-level image features, the high-level image features and the selected image features are respectively enhanced intra-module and inter-module in combination with the text features, as shown in the following formula:
[0019] ;
[0020] ;
[0021] in, , , , , , denote the first correlation weight matrix, the second correlation weight matrix, the third correlation weight matrix, the fourth correlation weight matrix, the fifth correlation weight matrix, and the sixth correlation weight matrix, respectively, represents the Sigmoid function, T represents the transposed matrix;
[0022] The following formula is used to take the average value to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features:
[0023] ;
[0024] in, represents low-level image features, high-level image features or selected image features; if is a low-level image feature, then represents the global average low-level image feature obtained by averaging the low-level image features, represents the low-level image features after in-mold enhancement, represents the low-level image features after inter-modal enhancement, represents the enhanced low-level image features; if is a high-level image feature, then represents the global average high-level image feature obtained by averaging high-level image features, represents the high-level image features after in-mold enhancement, represents the high-level image features after inter-modal enhancement, represents enhanced high-level image features; if is the selected image feature, then represents the global average selected image feature obtained by averaging the selected image features, represents low-level image features after in-mold enhancement, high-level image features after in-mold enhancement, or selected image features after in-mold enhancement, represents low-level image features after inter-modal enhancement, high-level image features after inter-modal enhancement, or selected image features after inter-modal enhancement, Indicates the selected image features that are enhanced.
[0025] Preferably, the image encoder includes a ViT model, the large language model includes an LLaMA3-8B model, and the multi-layer perceptron module includes two layers of multi-layer perceptrons connected in sequence.
[0026] Preferably, the selected image features are encoding features output by the last encoding layer of the image encoder.
[0027] Preferably, during the fine-tuning process of the multimodal large language model, the parameters of the pre-trained image encoder, the pre-trained text encoder and the pre-trained word segmenter are frozen, and the parameters of the multi-layer aggregation module, the intra-modal and inter-modal enhancement modules, the multi-layer perceptron and the pre-trained large language model are fine-tuned.
[0028] In a second aspect, the present invention provides a dialog generation device based on a multimodal large language model, comprising:
[0029] A model building module is configured to build and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model, wherein the multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained word segmenter, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model;
[0030] The generation module is configured to obtain a query statement and its corresponding image and input them into a fine-tuned multimodal large language model, the image is input into a pre-trained image encoder to obtain multi-scale coding features and selected image features, the multi-scale coding features are extracted through a multi-layer aggregation module to obtain low-level image features and high-level image features; the query statement is input into a text encoder to obtain text features; the text features, low-level image features, high-level image features and selected image features are input into intra-module and inter-module enhancement modules for enhancement to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features; the enhanced low-level image features, enhanced high-level image features and enhanced selected image features are connected along the channel and projected through a multi-layer perceptron module to obtain visual symbols; the query statement is input into a pre-trained word segmenter for word segmentation to obtain text symbols; the visual symbols and text symbols are input into the trained large language model to generate an answer statement.
[0031] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0032] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0033] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] (1) The dialogue generation method based on a multimodal large language model proposed in the present invention constructs a multimodal large language model based on image enhancement, and uses a multi-layer aggregation module to effectively aggregate high-level image features and low-level feature features, thereby enhancing the ability of MLLM to perceive multi-level detailed visual information.
[0036] (2) The multimodal large language model constructed by the dialogue generation method based on the multimodal large language model proposed in the present invention utilizes intra-modal and inter-modal enhancement modules to promote the intra-modal and inter-modal interaction of visual features, significantly improving the ability of MLLM to focus on key areas of the image.
[0037] (3) The dialogue generation method based on a multimodal large language model proposed in the present invention continuously and significantly improves the performance in various benchmark tests by fine-tuning the parameters of the multi-layer aggregation module, intra-modal and inter-modal enhancement modules, the multi-layer perceptron, and the pre-trained large language model, and can demonstrate its scalability, versatility, and effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 A flowchart of a method for generating a dialogue based on a multimodal large language model according to an embodiment of the present application;
[0040] Figure 2 A schematic diagram of a multimodal large language model of a dialogue generation method based on a multimodal large language model according to an embodiment of the present application;
[0041] Figure 3 Schematic diagram of image features extracted from different layers in CLIP;
[0042] Figure 4 Schematic diagram of the attention map between image features extracted from CLIP and different query statements;
[0043] Figure 5 A schematic diagram of a multi-layer aggregation module of a conversation generation method based on a multimodal large language model according to an embodiment of the present application;
[0044] Figure 6 A schematic diagram of an intra-modal and inter-modal enhancement module of a dialog generation method based on a multimodal large language model according to an embodiment of the present application;
[0045] Figure 7 A schematic diagram of a specific calculation process of intra-modal enhancement and inter-modal enhancement of a dialogue generation method based on a multimodal large language model according to an embodiment of the present application;
[0046] Figure 8 A schematic diagram of a conversation generation device based on a multimodal large language model according to an embodiment of the present application;
[0047] Fig. 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] Figure 1 A method for generating a dialogue based on a multimodal large language model provided by an embodiment of the present application is shown, comprising the following steps:
[0050] S1, constructs and fine-tunes a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model. The multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained word segmenter, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module and a pre-trained large language model.
[0051] In a specific embodiment, the image encoder includes a ViT model, and the large language model includes a LLaMA3-8B model.
[0052] In a specific embodiment, the multi-layer perceptron module includes two layers of multi-layer perceptrons.
[0053] In a specific embodiment, during the fine-tuning process of the multimodal large language model, the parameters of the pre-trained image encoder, the pre-trained text encoder, and the pre-trained word segmenter are frozen, and the parameters of the multi-layer aggregation module, the intra-modal and inter-modal enhancement modules, the multi-layer perceptron, and the pre-trained large language model are fine-tuned.
[0054] Specifically, refer to Figure 2 The embodiment of the present application proposes a multimodal large language model based on image feature enhancement. Based on the pre-trained large language model, the multimodal large language model adopts a multimodal input method consisting of text and image, and enhances the image features, so that the pre-trained large language model generates more accurate answer sentences.
[0055] First, the input image , using the image encoder based on the ViT model to encode into multi-scale coding features. This encoding process is expressed as:
[0056] ;
[0057] in, Represents the function corresponding to the image encoder, represents the encoded features output by the i-th encoding layer in the image encoder, N is the total number of layers of the image encoder, and d is the dimension of the encoded features.
[0058] Next, the extracted multi-scale encoding features Processed by the proposed Multi-level Aggregation Modul (MAM) module to derive high-level image features and low-level image features:
[0059] ;
[0060] in are high-level image features and low-level image features, respectively. is the number of image symbols, the function Represents the function corresponding to the Multi-layer Aggregation (MAM) module.
[0061] The encoding features output by the last encoding layer in the image encoder are used as selected image features, and the query statement is input into the text encoder to obtain text features; the low-level image features, high-level image features and selected image features are enhanced by intra-modal and inter-modal enhancement (IEM) modules in combination with the text features to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features.
[0062] The enhanced image features are then concatenated along the channel dimension , and projected using a multi-layer perceptron module:
[0063] ;
[0064] in, represents the connection operation along the channel dimension d, It is a multilayer perceptron module consisting of two layers of multilayer perceptrons (MLP), with an input dimension of 3d and an output dimension of D. It is a visual symbol.
[0065] Use the pre-trained tokenizer to tokenize the query sentence T to obtain text tokens Finally, the visual tokens and text tokens are input into the pre-trained Large Language Model (LLM) to generate the answer sentence:
[0066] ;
[0067] in, Indicates the answer sentence, Represents a pre-trained large language model.
[0068] In the process of fine-tuning the multimodal large language model, the parameters of the multi-layer aggregation module, intra-modal and inter-modal enhancement modules, multi-layer perceptron and pre-trained large language model are mainly adjusted, while the parameters of the pre-trained image encoder, pre-trained text encoder and pre-trained word segmenter are frozen.
[0069] S2, obtain the query statement and its corresponding image and input them into the fine-tuned multimodal large language model, input the image into the pre-trained image encoder to obtain multi-scale coding features and selected image features, and extract the multi-scale coding features through the multi-layer aggregation module to obtain low-level image features and high-level image features; input the query statement into the text encoder to obtain text features; input the text features, low-level image features, high-level image features and selected image features into the intra-modal and inter-modal enhancement modules for enhancement to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features; the enhanced low-level image features, enhanced high-level image features and enhanced selected image features are connected along the channel and projected through the multi-layer perceptron module to obtain visual symbols; input the query statement into the pre-trained word segmenter for word segmentation to obtain text symbols; input the visual symbols and text symbols into the trained large language model to generate the answer statement.
[0070] In a specific embodiment, the selected image features are encoding features output by the last encoding layer of the image encoder.
[0071] In a specific embodiment, the multi-scale coding feature is expressed as ,in, represents the encoding features output by the i-th encoding layer in the image encoder, represents the set of real numbers, represents the height of the encoded feature, represents the width of the encoded feature, represents the dimension of the encoded features, Represents the total number of coding layers in the image encoder; in the multi-layer aggregation module, the coding features output by the first 1 / 2N coding layers and the coding features output by the last 1 / 2N coding layers in the multi-scale coding features are averaged to obtain the low-level query features. and advanced query features , Represents the number of visual symbols. The low-level query feature is stacked with the first multi-scale coding feature output by the first 1 / 2N coding layers in the multi-scale coding feature to perform attention calculation to obtain the low-level image feature, as shown in the following formula:
[0072] ;
[0073] in, Represents the first multi-scale coding feature formed by stacking the coding features output by the first 1 / 2N coding layers in the multi-scale coding feature. , , Respectively represent the low-level query weight matrix, low-level key weight matrix, and low-level value weight matrix, represents the dimension of the first multi-scale encoded feature, represents low-level image features, represents the Softmax function, T represents the transposed matrix;
[0074] The high-level query feature is stacked with the encoding features output by the last 1 / 2N encoding layers in the multi-scale encoding feature to perform attention calculation to obtain the high-level image feature, as shown in the following formula:
[0075] ;
[0076] in, Represents the second multi-scale coding feature formed by stacking the coding features output by the last 1 / 2N coding layers in the multi-scale coding feature. , , They represent the advanced query weight matrix, advanced key weight matrix, and advanced value weight matrix respectively. represents the dimension of the second multi-scale encoded feature, Represents high-level image features.
[0077] It is well known that features from different layers of a neural network contain different levels of information. Figure 3 As shown in Figure 2, low-level features capture fine details such as edges, while high-level features contain more abstract and semantic information, which is crucial for semantic perception of images. Second, both intra-modal and inter-modal interactions can significantly improve the quality of image features by filtering out irrelevant information and amplifying relevant details. Figure 4 The attention maps between image features and various query statements are illustrated. In particular, the attention map between image features and image CLS tokens highlights foreground objects, while the attention map between image features and query statements focuses on objects referenced in the query statements. For example, when the query statement is "What color is the cat in the image?", the attention map focuses on the region containing the cat. Similarly, when the query statement is "How many dogs are there in the image?", the attention map highlights the region containing the dog. In addition, the attention map between image CLS tokens and image features shows that higher attention scores are concentrated in the cat and dog regions, while lower attention scores appear in the background regions. This phenomenon is advantageous as it enables the model to focus on key areas to provide accurate answers.
[0078] The following is a detailed description of the calculation process of the multi-layered aggregation module and the intra- and inter-mold reinforcement modules. Figure 5 , given the multi-scale encoding features , N is the total number of layers of the image encoder, and the multi-layer aggregation (MAM) module is designed to extract high-level image features and low-level image features First, the low-level query features are obtained by averaging the encoding features of the first and second half layers of the multi-scale encoding features. and advanced query features This creates representative query features that summarize information from low-level and high-level contexts respectively.
[0079] Next, an attention map is calculated between the first multi-scale coding features or the second multi-scale coding features stacked with the low-level query features or the high-level query features and their corresponding multiple coding features to identify which layers are related to the low-level query features or the high-level query features. This is crucial for focusing on the coding layers in the image features. This process enhances the original query features by emphasizing the important coding layers identified by the attention map. The low-level query features and the high-level query features are collectively referred to as query features, and the first multi-scale coding features and the second multi-scale coding features are collectively referred to as multi-layer coding features.
[0080] In a specific embodiment, in the intra-modal and inter-modal enhancement modules, the low-level image features, the high-level image features and the selected image features are respectively enhanced intra-modally and inter-modally in combination with the text features, as shown in the following formula:
[0081] ;
[0082] ;
[0083] in, , , , , , denote the first correlation weight matrix, the second correlation weight matrix, the third correlation weight matrix, the fourth correlation weight matrix, the fifth correlation weight matrix, and the sixth correlation weight matrix, respectively, represents the Sigmoid function, T represents the transposed matrix;
[0084] The following formula is used to take the average value to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features:
[0085] ;
[0086] in, represents low-level image features, high-level image features or selected image features; if is a low-level image feature, then represents the global average low-level image feature obtained by averaging the low-level image features, represents the low-level image features after in-mold enhancement, represents the low-level image features after inter-modal enhancement, represents the enhanced low-level image features; if is a high-level image feature, then represents the global average high-level image feature obtained by averaging high-level image features, represents the high-level image features after in-mold enhancement, represents the high-level image features after inter-modal enhancement, represents enhanced high-level image features; if is the selected image feature, then represents the global average selected image feature obtained by averaging the selected image features, represents low-level image features after in-mold enhancement, high-level image features after in-mold enhancement, or selected image features after in-mold enhancement, represents low-level image features after inter-modal enhancement, high-level image features after inter-modal enhancement, or selected image features after inter-modal enhancement, Indicates the selected image features that are enhanced.
[0087] Specifically, refer to Figure 6 , the intra-modal and inter-modal enhancement (IEM) module aims to promote intra-modal and inter-modal interactions between image features at different levels to enhance image features. These levels include high-level image features, low-level image features, and selected image features. For simplicity, high-level image features will be taken as an example, and the operations of features at other levels are similar. High-level image features, low-level image features, and selected image features are collectively referred to as image features, the global average high-level image features, the global average low-level image features, and the global average selected image features are collectively referred to as global average image features, and the enhanced high-level image features, enhanced low-level image features, and enhanced selected image features are collectively referred to as enhanced image features.
[0088] like Figure 6 As shown, the input includes high-level image features from MAM , text features (i.e., the text features extracted after encoding the query using the text encoder) and the global average high-level image features (ie The global average high-level image features obtained by averaging).
[0089] like Figure 7As shown in Figure 1, the correlation map between global symbols and visual symbols is first calculated using matrix multiplication. Then the visual symbols are enhanced by element-by-element multiplication based on the correlation map. This process can be understood as focusing on salient features by considering image-to-image (intra-modal) and image-to-text (inter-modal) correlations. The intra-modal and inter-modal enhancement formulas are as follows:
[0090] ;
[0091] ;
[0092] Among them, if is a high-level image feature, then represents the global average high-level image feature obtained by averaging high-level image features, represents the high-level image features after in-mold enhancement, Represents high-level image features after inter-modal enhancement;
[0093] The enhanced features and the features before enhancement are averaged to obtain enhanced high-level image features, as shown in the following formula:
[0094] ;
[0095] in, Represents enhanced high-level image features.
[0096] Further deployment and application of the fine-tuned multimodal large language model (denoted as vMLLM) can highlight the capabilities of text recognition, visual perception, hallucination, mathematical knowledge, diagram understanding, map understanding, and poster understanding.
[0097] The fine-tuned multimodal large language model proposed in the embodiments of the present application is compared with other advanced technologies.
[0098] Table 1 shows the comparison results of the fine-tuned multimodal large language model proposed in the embodiment of the present application with the state-of-the-art technologies, including DenseConnector, VILA and MG-LLaVA. It is worth noting that the fine-tuned multimodal large language model (vMLLM) proposed in the embodiment of the present application is based on the powerful LLaVA 1.5 architecture. The experimental results show that vMLLM has achieved significant performance improvements over LLaVA-1.5 in various benchmarks. Specifically, when using the same LLaMA3-8B language model, vMLLM's performance on the MMB, MM-Vet and MathVista benchmarks is improved by 8.2%, 8.1% and 13.4%, respectively. These enhancements highlight the effectiveness of the method of the present invention in enhancing visual features. In addition, the performance of vMLLM is consistently better than the most recent state-of-the-art models. For example, compared with DenseConnector, vMLLM's performance on the MMB, MMVet and MathVista benchmarks is improved by 2.7%, 8.8% and 12.4%, respectively.
[0099] Table 1:
[0100]
[0101] Further references Figure 8 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a dialogue generation device based on a multimodal large language model. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0102] The embodiment of the present application provides a dialog generation device based on a multimodal large language model, including:
[0103] A model building module 1 is configured to build and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model, wherein the multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained word segmenter, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model;
[0104] The generation module 2 is configured to obtain a query statement and its corresponding image and input them into a fine-tuned multimodal large language model, the image is input into a pre-trained image encoder to obtain multi-scale coding features and selected image features, and the multi-scale coding features are extracted through a multi-layer aggregation module to obtain low-level image features and high-level image features; the query statement is input into a text encoder to obtain text features; the text features, low-level image features, high-level image features and selected image features are input into the intra-module and inter-module enhancement modules for enhancement to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features; the enhanced low-level image features, enhanced high-level image features and enhanced selected image features are connected along the channel and projected through a multi-layer perceptron module to obtain visual symbols; the query statement is input into a pre-trained word segmenter for word segmentation to obtain text symbols; the visual symbols and text symbols are input into the trained large language model to generate an answer statement.
[0105] Fig. 9 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Fig. 9 As shown, the electronic device of this embodiment includes: a processor 901 and a memory 902; wherein the memory 902 is used to store computer-executable instructions; the processor 901 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0106] Optionally, the memory 902 may be independent or integrated with the processor 901 .
[0107] When the memory 902 is independently provided, the electronic device further includes a bus 903 for connecting the memory 902 and the processor 901 .
[0108] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 901 executes the computer execution instructions, the above method is implemented.
[0109] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 901, the above method is implemented.
[0110] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0111] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.
[0112] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0113] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 901 to perform some steps of the methods of various embodiments of the present application.
[0114] It should be understood that the processor 901 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 901 may be any conventional processor 901, etc. The steps of the method disclosed in the invention may be directly embodied in the hardware processor 901 for execution, or may be executed by a combination of hardware and software modules in the processor 901.
[0115] The memory 902 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.
[0116] The bus 903 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 903 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 903 in the drawings of the present application is not limited to only one bus 903 or one type of bus 903.
[0117] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0118] An exemplary storage medium is coupled to the processor 901, so that the processor 901 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 901. The processor 901 and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor 901 and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0119] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating dialogue based on a multimodal large language model, characterized in that: The following steps are involved: Constructing and fine-tuning a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model, wherein the multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained word segmenter, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model; A query statement and its corresponding image are obtained and input into the fine-tuned multimodal large language model, the image is input into the pre-trained image encoder to obtain multi-scale coding features and selected image features, the multi-scale coding features are extracted through the multi-layer aggregation module to obtain low-level image features and high-level image features; the query statement is input into the text encoder to obtain text features; the text features, low-level image features, high-level image features and selected image features are input into the intra-module and inter-module enhancement modules for enhancement to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features; The enhanced low-level image features, enhanced high-level image features and enhanced selected image features are connected along the channel and projected through the multi-layer perceptron module to obtain visual symbols; the query statement is input into the pre-trained word segmenter for word segmentation to obtain text symbols; The visual symbols and text symbols are input into the trained large language model to generate a response sentence.
2. The method for generating a dialogue based on a multimodal large language model according to claim 1, characterized in that: The multi-scale coding feature is expressed as ,in, represents the encoding features output by the i-th encoding layer in the image encoder, represents the set of real numbers, represents the height of the encoded feature, represents the width of the encoded feature, represents the dimension of the encoded features, represents the total number of coding layers in the image encoder; in the multi-layer aggregation module, the coding features output by the first 1 / 2N coding layers and the coding features output by the last 1 / 2N coding layers in the multi-scale coding features are first averaged to obtain the low-level query features and advanced query features , represents the number of visual symbols, and the first multi-scale coding feature formed by stacking the low-level query feature and the coding features output by the first 1 / 2N coding layers in the multi-scale coding feature is subjected to attention calculation to obtain the low-level image feature, as shown in the following formula: ; in, represents a first multi-scale coding feature formed by stacking coding features output by the first 1 / 2N coding layers in the multi-scale coding feature, , , Respectively represent the low-level query weight matrix, low-level key weight matrix, and low-level value weight matrix, represents the dimension of the first multi-scale encoded feature, represents low-level image features, represents the Softmax function, T represents the transposed matrix; The high-level query feature and the second multi-scale coding feature formed by stacking the coding features output by the last 1 / 2N coding layers in the multi-scale coding features are subjected to attention calculation to obtain the high-level image feature, as shown in the following formula: ; in, represents a second multi-scale coding feature formed by stacking coding features output by the last 1 / 2N coding layers in the multi-scale coding feature, , , They represent the advanced query weight matrix, advanced key weight matrix, and advanced value weight matrix respectively. represents the dimension of the second multi-scale encoded feature, Represents high-level image features.
3. The method for generating a dialogue based on a multimodal large language model according to claim 1, characterized in that: In the intra-modal and inter-modal enhancement modules, the low-level image features, high-level image features and selected image features are respectively enhanced intra-modally and inter-modally in combination with the text features, as shown in the following formula: ; ; in, , , , , , denote the first correlation weight matrix, the second correlation weight matrix, the third correlation weight matrix, the fourth correlation weight matrix, the fifth correlation weight matrix, and the sixth correlation weight matrix, respectively, represents the Sigmoid function, T represents the transposed matrix; The following formula is used to take the average value to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features: ; in, represents the low-level image feature, high-level image feature or selected image feature; if is the low-level image feature, then represents the global average low-level image feature obtained by averaging the low-level image features, represents the low-level image features after in-mold enhancement, represents the low-level image features after inter-modal enhancement, represents the enhanced low-level image features; if is the high-level image feature, then represents the global average high-level image feature obtained by averaging the high-level image features, represents the high-level image features after in-mold enhancement, represents the high-level image features after inter-modal enhancement, represents enhanced high-level image features; if is the selected image feature, then represents the global average selected image feature obtained by averaging the selected image features, represents low-level image features after in-mold enhancement, high-level image features after in-mold enhancement, or selected image features after in-mold enhancement, represents low-level image features after inter-modal enhancement, high-level image features after inter-modal enhancement, or selected image features after inter-modal enhancement, Indicates the selected image features that are enhanced.
4. The method for generating a dialogue based on a multimodal large language model according to claim 1, characterized in that: The image encoder includes a ViT model, the large language model includes an LLaMA3-8B model, and the multilayer perceptron module includes two layers of multilayer perceptrons connected in sequence.
5. The method for generating dialogue based on a multimodal large language model according to claim 1, characterized in that: The selected image features are encoding features output by the last encoding layer of the image encoder.
6. The method for generating a dialogue based on a multimodal large language model according to claim 1, characterized in that: During the fine-tuning process of the multimodal large language model, the parameters of the pre-trained image encoder, the pre-trained text encoder and the pre-trained word segmenter are frozen, and the parameters of the multi-layer aggregation module, the intra-modal and inter-modal enhancement modules, the multi-layer perceptron and the pre-trained large language model are fine-tuned.
7. A dialogue generation device based on a multimodal large language model, characterized in that: include: A model building module is configured to build and fine-tune a multimodal large language model based on image feature enhancement to obtain a fine-tuned multimodal large language model, wherein the multimodal large language model includes a pre-trained image encoder, a pre-trained text encoder, a pre-trained word segmenter, a multi-layer aggregation module, an intra-modal and inter-modal enhancement module, a multi-layer perceptron module, and a pre-trained large language model; A generation module is configured to obtain a query statement and its corresponding image and input them into the fine-tuned multimodal large language model, the image is input into the pre-trained image encoder to obtain multi-scale coding features and selected image features, the multi-scale coding features are extracted through the multi-layer aggregation module to obtain low-level image features and high-level image features; the query statement is input into the text encoder to obtain text features; the text features, low-level image features, high-level image features and selected image features are input into the intra-module and inter-module enhancement modules for enhancement to obtain enhanced low-level image features, enhanced high-level image features and enhanced selected image features; The enhanced low-level image features, enhanced high-level image features and enhanced selected image features are connected along the channel and projected through the multi-layer perceptron module to obtain visual symbols; the query statement is input into the pre-trained word segmenter for word segmentation to obtain text symbols; The visual symbols and text symbols are input into the trained large language model to generate a response sentence.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Memory grounded conversational reasoning and question answering for assistant systems
CN114072832A
Systems and methods for learning unified representations of language, image, and point cloud for three-dimensional recognition
US20240160917A1
Pedestrian attribute cross-modal alignment method based on complete attribute identification enhancement
WO2024114185A1
Knowledge fusion multi-modal interaction method and apparatus based on improved alignment method
WO2025025290A1
Multi-modal knowledge-based question answering method and system for 5g message
WO2025060773A1