Multi-modal large language model dialogue generation method based on natural language understanding

By adopting a multimodal large language model based on natural language understanding in the dialogue system, processing and fusion of multiple modal information, the problem of insufficient ability of dialogue systems to understand user intentions and emotional emotions in the existing technology is solved, and more natural and relevant dialogue replies are achieved, and the user experience is improved.

CN119989268APending Publication Date: 2025-05-13BEIJING LIANPING TECH CO LTD
View PDF 0 Cites 15 Cited by

Patent Information

Application Number
CN202510071190.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process and integrate multiple modal information, resulting in insufficient ability of dialogue systems to understand user intentions and emotions, and the generated replies are unnatural, low correlation, and poor user experience.

Method used

The multimodal large language model dialogue generation method based on natural language understanding is adopted, and the dynamic weighted fusion strategy of the multi-head attention mechanism is used to uniformly represent multimodal information such as text, images and sound, and ensure that the reply content is related to the dialogue history based on a long context processing algorithm.

Benefits of technology

Improves the ability to understand user intentions, maintains context coherence, and generates replies more natural and relevant, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989268A_ABST
    Figure CN119989268A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and provides a multi-modal large language model dialogue generation method based on natural language understanding, comprising the following steps: receiving multi-modal information input by a user through a multi-modal large language model, the multi-modal information comprising a plurality of modal data; preprocessing modal data in the multi-modal information, and extracting to obtain multi-modal features; performing unified fusion feature representation on the extracted multi-modal features based on a dynamic weighted fusion strategy of a multi-head attention mechanism; determining the dialogue state of the long context based on a long context processing algorithm, and ensuring that the generated reply content is associated with the dialogue history; and according to the unified fusion feature representation and the dialogue state, a natural language is generated through an RAG retrieval enhancement generation technology for replying. According to the method, input of multiple modes can be processed and understood, the ability of understanding the intention of the user is improved, the continuity of the context is maintained, and the generated reply is more natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal large language model dialogue generation method based on natural language understanding. Background Art

[0002] With the rapid development of artificial intelligence technology, natural language processing (NLP) technology and large language model (LLM) technology have been widely used in various dialogue systems, such as intelligent assistants, intelligent customer service, digital humans, etc. Intelligent dialogue systems are widely used in finance, medical care, education, media and other fields. The dialogue systems under current technology usually rely on predefined rules or content templates to answer questions, and also use simple machine learning models, which limits the naturalness and flexibility of the dialogue. In addition, in traditional large model dialogue systems, the interaction mainly relies on text input and text intent recognition. However, human communication usually includes multiple modalities, such as expressions, body language, images and sounds.

[0003] In order to more accurately understand user intentions and provide a more natural interactive experience, it is crucial to develop an AI large-model dialogue system that can process and fuse multiple modal information. Text input alone can no longer meet the accuracy requirements of human interaction. If non-text information data is encountered, such as user facial expression images, user videos or sound input, the model often cannot obtain effective information, which limits the ability of large-model dialogues to understand user intentions and emotions. In addition, existing technologies often ignore the coherence of the long context of the dialogue, resulting in the generated responses may be inconsistent with the previous text or not highly relevant. These limitations lead to poor user experience and may even give irrelevant answers. Over time, users may not continue to choose to use the AI ​​model to help them solve related needs. Therefore, it is necessary to provide a multimodal large language model dialogue generation method based on natural language understanding to solve the above problems. Summary of the invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a multimodal large language model dialogue generation method based on natural language understanding, which is an AI large model for intelligent dialogue generation that can comprehensively process multimodal inputs such as text, images and sounds, and can understand and generate human language more naturally, achieve more fluent and intelligent dialogue interaction, so as to improve the understanding ability and interaction naturalness of the dialogue system. The AI ​​large model includes main components such as user intention recognition, sentiment analysis, knowledge understanding and response generation, and is suitable for a variety of application scenarios such as intelligent customer service, intelligent educational robots, and smart home assistants to solve the problems existing in the above-mentioned background technology.

[0005] The present invention is implemented as follows: a multimodal large language model dialogue generation method based on natural language understanding, the method comprising the following steps:

[0006] Receiving multimodal information input by a user through a multimodal large language model, wherein the multimodal information includes a plurality of modal data;

[0007] Preprocess the modal data in the multimodal information to extract multimodal features;

[0008] A dynamic weighted fusion strategy based on the multi-head attention mechanism is used to unify the extracted multimodal features into a unified fusion feature representation.

[0009] Determine the long-context conversation state based on the long-context processing algorithm to ensure that the generated reply content is associated with the conversation history;

[0010] Based on the unified fusion feature representation and dialogue state, natural language is generated for reply through RAG retrieval enhancement generation technology;

[0011] Based on the emotional audio features of speech, the expression features of images and the emotional tendencies of texts, the user's current emotional state is comprehensively and collaboratively analyzed to identify the user's intentions.

[0012] As a further solution of the present invention: the multimodal large language model includes a model encoder, a fusion layer and a Transformer core, the model encoder is used to convert raw data of different modalities into specific features; the fusion layer is used to integrate features from encoders of different modalities to create a unified fused feature representation; the Transformer core is used to process and understand the fused multimodal data.

[0013] As a further solution of the present invention: the multimodal information includes text data, image data and sound data. After receiving the multimodal information, a dynamic weighted fusion method is used to design an adaptive weight mechanism. Based on the attention mechanism, the weights of different modalities are learned through model training, and the weights are dynamically allocated between the input modal data. The importance of each modality is adjusted in real time according to the context, thereby optimizing the extraction and fusion of information.

[0014] As a further solution of the present invention: the step of preprocessing the modal data in the multimodal information and extracting the multimodal features specifically includes: denoising and standardizing the modal data, using an enhanced algorithm for feature extraction, adding a multimodal emotion fusion network feature layer, and analyzing the emotional characteristics of different modalities.

[0015] As a further solution of the present invention: for text data, preprocessing includes: removing irrelevant characters, converting to lowercase, word segmentation, removing stop words, stemming and word form restoration, and feature extraction includes: bag-of-words model and TF-IDF word embedding; for image data, preprocessing includes: unifying image size, normalization and data enhancement, and feature extraction includes: using pre-trained convolutional neural network and extracting feature graphs; for sound data, preprocessing includes: denoising, normalizing volume, framing and applying window function, and feature extraction includes: Mel-frequency cepstral coefficients, spectrogram, zero crossing rate and energy features.

[0016] As a further solution of the present invention: the subsequent steps of unifying the fused feature representation of the extracted multimodal features are: combining the global context and historical conversation status to ensure adaptive adjustment of weights, and in order to ensure the accuracy of feature fusion, combining multi-task learning and reinforcement learning strategies, optimizing the weight distribution of different tasks, and adjusting the feature weight distribution after fusion of different modalities through feedback.

[0017] As a further solution of the present invention: the step of determining the long context dialogue state based on the long context processing algorithm to ensure that the generated reply content is associated with the dialogue history specifically includes: using an extended version of the Transformer architecture model to process modal data of different lengths, and in multiple rounds of dialogues, using a memory network mechanism to ensure that information in long conversations is retained and processed, and storing and passing information from previous conversations to subsequent dialogue generation stages.

[0018] As a further solution of the present invention: the method also includes constructing a user personalized portrait-driven response generation algorithm based on the user's historical behavior, interest preferences and habits. In the conversation, the user's personalized portrait is used as one of the inputs of the model, combined with the context information of the conversation, to generate responses to the user's personalized needs; and combined with the adaptive learning strategy of multi-task learning and meta-learning, respond to user feedback and adjust model parameters.

[0019] Another object of the present invention is to provide a multimodal large language model dialogue generation system based on natural language understanding, the system comprising:

[0020] An input processing module, used for receiving multimodal information input by a user through a multimodal large language model, wherein the multimodal information includes a plurality of modal data;

[0021] A feature extraction module is used to preprocess the modal data in the multimodal information and extract multimodal features;

[0022] The feature fusion module is used to unify the fusion feature representation of the extracted multimodal features based on the dynamic weighted fusion strategy of the multi-head attention mechanism;

[0023] A context processing module, which is used to determine the long-context conversation state based on a long-context processing algorithm and ensure that the generated reply content is associated with the conversation history;

[0024] The reply response module is used to generate natural language replies based on the unified fusion feature representation and dialogue state through RAG retrieval enhancement generation technology;

[0025] The emotion recognition module is used to comprehensively and collaboratively analyze the user's current emotional state and identify the user's intention based on the emotional audio features of the voice, the expression features of the image, and the emotional tendency of the text.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] The present invention can process and understand inputs of multiple modalities, improve the ability to understand user intent, maintain contextual coherence, and generate more natural and relevant responses. Through unified feature representation and processing flow, the model can more effectively learn and utilize the relationships and dependencies between different modalities, improving processing efficiency and effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flowchart of a multimodal large language model dialogue generation method based on natural language understanding.

[0029] Figure 2 The present invention is a flowchart for processing modal data in a multimodal large language model dialogue generation method based on natural language understanding.

[0030] Figure 3 This is a schematic diagram of the structure of a multimodal large language model dialogue generation system based on natural language understanding. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments.

[0033] like Figure 1 As shown, an embodiment of the present invention provides a multimodal large language model dialogue generation method based on natural language understanding, and the method includes the following steps:

[0034] S100, receiving multimodal information input by a user through a multimodal large language model, wherein the multimodal information includes a plurality of modal data;

[0035] S200, preprocessing the modal data in the multimodal information to extract multimodal features;

[0036] S300, a dynamic weighted fusion strategy based on a multi-head attention mechanism, which unifies the extracted multimodal features into a unified fusion feature representation;

[0037] S400, determining the long context conversation state based on the long context processing algorithm, and ensuring that the generated reply content is associated with the conversation history;

[0038] S500, based on the unified fusion feature representation and the dialogue state, generates natural language for reply through RAG retrieval enhancement generation technology;

[0039] S600, comprehensively and collaboratively analyzing the user's current emotional state based on the emotional audio features of the voice, the expression features of the image, and the emotional tendency of the text, and identifying the user's intention.

[0040] In an embodiment of the present invention, the multimodal large language model includes a model encoder, a fusion layer and a Transformer core. The model encoder is used to convert raw data of different modalities into specific features, and convert raw data of different modalities (such as images, videos, and audio) into high-level feature representations. These encoders are optimized for each modality to capture the unique information of the modal data; the fusion layer is used to integrate features from encoders of different modalities to create a unified fused feature representation. The key to this layer is that it can process and fuse information of different modalities so that the model can have a comprehensive understanding; the Transformer core is used to process and understand the fused multimodal data, using a large language model (LLM) GPT, Llama series, etc.

[0041] In an embodiment of the present invention, the multimodal information includes text data, image data and sound data. After receiving the multimodal information, a dynamic weighted fusion method is used to design an adaptive weight mechanism (Adaptive Weight Mechanism), and based on the attention mechanism (Attention Mechanism), the weights of different modalities in specific situations are learned through model training, and weights are dynamically allocated between different input modalities. The importance of each modality is adjusted in real time according to the context, thereby optimizing the extraction and fusion of information.

[0042] In an embodiment of the present invention, the step of preprocessing the modal data in the multimodal information and extracting the multimodal features specifically includes: denoising and standardizing the modal data, and extracting features. An enhanced algorithm for feature extraction is used to add a multimodal emotion fusion network feature layer, mainly to analyze the emotional features that different modalities may have, ensure the accuracy of emotion feature recognition, and mine the user's intentions and emotional states, thereby optimizing downstream tasks.

[0043] like Figure 2 As shown, in the embodiment of the present invention, for text data, preprocessing includes: removing irrelevant characters, converting to lowercase, word segmentation, removing stop words, stemming and word form restoration, and feature extraction includes: bag-of-words model and TF-IDF word embedding; for image data, preprocessing includes: unifying image size, normalization and data enhancement, and feature extraction includes: using a pre-trained convolutional neural network and extracting feature maps; for sound data, preprocessing includes: denoising, normalizing volume, framing and applying a window function, and feature extraction includes: Mel-frequency cepstral coefficients, spectrograms, zero crossing rates and energy features.

[0044] In the embodiment of the present invention, when the extracted multimodal features are unified into fused features, a dynamic weighting fusion strategy based on a multi-head attention mechanism is adopted to unify the fusion representation of the preprocessed and extracted multimodal features, and the global context and historical dialogue state are combined to ensure the adaptive adjustment of the weights. In order to ensure the accuracy of feature fusion, the strategies of multi-task learning and reinforcement learning are combined to optimize the weight distribution of different tasks, and then the weight distribution of features after fusion of different modalities is adjusted through feedback. Finally, the generated feature representation is made more accurate and effective.

[0045] In an embodiment of the present invention, the step of determining the long context dialogue state based on the long context processing algorithm and ensuring that the generated reply content is associated with the dialogue history specifically includes: using an extended version of the Transformer architecture model to process inputs of different lengths (including images, videos, texts, audios, etc.), the model can process longer contexts with higher efficiency, and in multi-round dialogues, using a memory network mechanism to ensure that information in long-term conversations is retained and efficiently processed, storing and passing key information from previous dialogues to subsequent dialogue generation stages, and avoiding context interruption or information loss as much as possible.

[0046] In the embodiment of the present invention, a natural language response is generated by a response generation module based on a unified feature representation and a dialog state. Here, the RAG retrieval-augmented generation technology is introduced. When encountering complex or real-time problems, the system can dynamically retrieve external knowledge bases, integrate multimodal features and external knowledge, and thus generate more real-time and informative natural language responses.

[0047] In an embodiment of the present invention, the method also includes adjusting model parameters according to user feedback to improve the quality of conversation, specifically: a personalized user portrait driven response generation algorithm is constructed based on multi-dimensional data such as user historical behavior, interest preferences, habits, etc. In an actual conversation, the system first uses the portrait as one of the inputs of the model, and combines the context information of the current conversation to generate a response that is more in line with the user's personalized needs, thereby improving the user experience; combining multi-task learning (Multi-taskLearning) and meta-learning (Meta-learning) adaptive learning strategies, the system can quickly respond to user feedback and adjust model parameters in a short time to improve the generated conversation quality and the model's self-improvement ability.

[0048] In the embodiment of the present invention, a deep neural network model is used to identify user intentions. The current emotional state of the user is comprehensively and collaboratively analyzed based on the emotional audio features of the voice, the expression features of the image (such as facial expression recognition), and the emotional tendency of the text. In the case of inconsistent modalities in the input, the accuracy of emotion recognition is enhanced, thereby providing more humane answers. For complex queries and real-time update problems, the knowledge base or RAG technology, model incremental iteration and other methods can be used to combine natural language understanding technology to process, and finally generate natural and fluent dialogue output.

[0049] In the embodiment of the present invention, in order to ensure the efficient deployment and operation of the model, the model quantization (Quantization) and distillation (Distillation) and other technologies are used to reduce the computational complexity of the model, and the distributed training and reasoning methods are adopted to ensure that the model runs efficiently on different computing devices. It not only ensures the quality of the generated dialogue, but also optimizes the use of computing resources and improves the user's response speed. Large model implementation case: For example, the user inputs through the dialogue interface: "I feel very sad today, recommend a comedy movie." The multimodal emotion fusion network comprehensively analyzes the user's voice intonation, facial expression (assuming there is one), and text content, and concludes that the user's current emotional state is "depressed". Then, the system uses the user portrait data to determine the user's movie preference, and then determines that the user's request is "movie recommendation" through the intention recognition module, and the emotion analysis module identifies the user's emotion as "sad". The current popular comedy movie data in the knowledge base management module generates a personalized response: "I recommend you to watch the recently highly rated comedy "A Beautiful Day", I hope it will make you happier." It is displayed to the user through the user interaction interface. In addition, the system will continuously optimize the accuracy and personalization of movie recommendations based on user feedback.

[0050] In the embodiment of the present invention, cross-modal feature learning is also performed, specifically including multimodal association learning, context-aware model training and end-to-end optimization.

[0051] When performing multimodal association learning, modal alignment and association analysis are first performed. Modal alignment: This step involves aligning data from different modalities into a unified semantic space so that information from different sources can be directly compared and associated. For example, the "red car" in the text description should match the visual features of the red car shown in the image. Association analysis: Analyze the association between different modalities through deep learning techniques such as collaborative filtering, canonical correlation analysis (CCA), or deep canonical correlation analysis (DCCA). For example, learn the association between keywords in the text and specific objects in the image. Then perform cross-modal embedding. Cross-modal embedding space: Use deep neural networks such as multimodal variational autoencoders (MVAE) or cross-modal Transformers to create a shared embedding space in which text, images, video, and audio data are encoded into vectors with high semantic correlation. Feature fusion: Design a fusion mechanism (for example, through an attention mechanism or a fusion layer) to integrate features from different modalities and enhance the model's ability to capture key information.

[0052] When training a context-aware model, we first model the context sequence: using a sequence model (such as LSTM or Transformer) to process cross-modal inputs, maintain the coherence of the conversation history, and ensure that the generated response takes into account both the current input and the previous communication content. Then there is dynamic adjustment: dynamically adjust the focus of the model according to the context of the conversation, for example, by adjusting the attention weight to prioritize the most relevant modal input.

[0053] When performing end-to-end optimization, first perform joint training: optimize the processing and fusion of all modalities at the same time through a joint training strategy, rather than processing each modality separately. This helps the model better learn how to integrate information from different sources and improve the relevance and accuracy of the final output. Then determine the dialogue generation strategy, conditional generation: when generating dialogue responses, the model considers the integrated multimodal information as a condition to generate more accurate and relevant answers. Feedback learning: use user interaction feedback to fine-tune the model and improve the model's ability to handle complex multimodal scenarios.

[0054] In the embodiment of the present invention, the quality of the training data set: the data set is the basis of the training model, and the quality of the data set directly affects the performance of the model. Therefore, obtaining a high-quality multimodal data set is one of the key points. This requires that the data set contains rich modal information, such as text, images, sounds, etc., and these modal information needs to be closely related to the task. In addition, the data set needs to be properly preprocessed, such as word segmentation, sequence filling, word vector encoding, etc., so that the model can better learn and understand the data. The structure of the model: the structure of the model is a key factor in determining its performance. An encoder-decoder structure is adopted, and an attention mechanism is introduced between the encoder and the decoder. This enables the model to focus on the important parts of the input when generating a response. In addition, the model also uses a bidirectional LSTM and a shared word embedding layer to ensure the context awareness of the model. These designs are all to improve the performance and efficiency of the model. The choice of hyperparameters: the choice of hyperparameters also has an important impact on the performance of the model. For example, sequence length (SEQ_LENGTH), vocabulary size (MAX_NB_WORDS), hidden layer dimension (LATENT_DIM), batch size (BATCH_SIZE) and number of training rounds (EPOCHS) are all hyperparameters that need to be carefully selected. These hyperparameters need to be adjusted according to the characteristics of the task and the characteristics of the data set. Use of pre-trained word vectors: Pre-trained word vectors are a method of mapping words to high-dimensional space, which can capture the semantic information of words. In this algorithm, pre-trained word vectors are used to initialize the word embedding layer, which can greatly improve the performance of the model. Model training and evaluation: Model training and evaluation are important steps in machine learning. This algorithm uses the Adam optimizer and sparse category cross entropy loss function for training, and uses accuracy as an evaluation metric. In addition, the model also uses callback functions such as TensorBoard and ModelCheckpoint to monitor the performance of the model during training and save the best model. In general, the embodiments of the present invention have key points in terms of the quality of the training data set, the structure of the model, the selection of hyperparameters, the use of pre-trained word vectors, and the training and evaluation of the model.

[0055] The present invention can process and understand inputs of multiple modalities, improving the system's ability to understand user intent. Maintaining the coherence of the context, the generated replies are more natural and more relevant. The adaptive learning mechanism enables the system to self-optimize based on user feedback, improving long-term user satisfaction. Flexible and applicable to a wide range of application scenarios, improving the commercial application value of the system, multimodal models can understand and generate multiple types of data, so that they can provide more natural and expressive interaction methods in more application scenarios, such as in virtual assistants, educational technology and media production. The architecture of the present invention supports cross-modal conversion, such as from text to image (text description to generate image), from image to text (image description), etc. This capability is very valuable in the fields of automated content creation, auxiliary design decision-making and enhanced user interaction experience. Through a unified feature representation and processing flow, the model can more effectively learn and utilize the relationships and dependencies between different modalities, improving processing efficiency and effect. It is scalable, and the architectural design allows modules (such as different modal encoders or generators) to be easily added or replaced, so that the model can adapt to new modalities or improve the processing of existing modalities as needed. Data fusion denoising output,When dealing with situations containing noisy or incomplete data, multimodal models,by fusing information from different sources, are better able to resist noise and provide more accurate,output.

[0056] like Figure 3 As shown, an embodiment of the present invention further provides a multimodal large language model dialogue generation system based on natural language understanding, the system comprising:

[0057] An input processing module 100 is used to receive multimodal information input by a user through a multimodal large language model, wherein the multimodal information includes a plurality of modal data;

[0058] The feature extraction module 200 is used to pre-process the modal data in the multimodal information to extract multimodal features;

[0059] The feature fusion module 300 is used to unify the fusion feature representation of the extracted multimodal features based on the dynamic weighted fusion strategy of the multi-head attention mechanism;

[0060] A context processing module 400, for determining the long context conversation state based on a long context processing algorithm, and ensuring that the generated reply content is associated with the conversation history;

[0061] The reply response module 500 is used to generate natural language for replying through RAG retrieval enhancement generation technology according to the unified fusion feature representation and the dialogue state;

[0062] The emotion recognition module 600 is used to comprehensively and collaboratively analyze the user's current emotional state based on the emotional audio features of the voice, the expression features of the image and the emotional tendency of the text, and identify the user's intention.

[0063] The above only describes in detail the preferred embodiments of the present invention, which is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0064] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0065] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0066] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the disclosure in the specification and examples. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

Claims

1. A multimodal large language model dialogue generation method based on natural language understanding, characterized in that: The method comprises the following steps: Receiving multimodal information input by a user through a multimodal large language model, wherein the multimodal information includes a plurality of modal data; Preprocess the modal data in the multimodal information to extract multimodal features; A dynamic weighted fusion strategy based on the multi-head attention mechanism is used to unify the extracted multimodal features into a unified fusion feature representation. Determine the long-context conversation state based on the long-context processing algorithm to ensure that the generated reply content is associated with the conversation history; Based on the unified fusion feature representation and dialogue state, natural language is generated for reply through RAG retrieval enhancement generation technology; Based on the emotional audio features of speech, the expression features of images and the emotional tendencies of texts, the user's current emotional state is comprehensively and collaboratively analyzed to identify the user's intentions.

2. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 1, characterized in that: The multimodal large language model includes a model encoder, a fusion layer and a Transformer core, wherein the model encoder is used to convert raw data of different modalities into specific features; the fusion layer is used to integrate features from encoders of different modalities to create a unified fusion feature representation; The Transformer core is used to process and understand the fused multimodal data.

3. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 1, characterized in that: The multimodal information includes text data, image data and sound data. After receiving the multimodal information, a dynamic weighted fusion method is used to design an adaptive weight mechanism. Based on the attention mechanism, the weights of different modalities are learned through model training, and the weights are dynamically allocated between the input modal data. The importance of each modality is adjusted in real time according to the context, thereby optimizing the extraction and fusion of information.

4. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 3, characterized in that: The step of preprocessing the modal data in the multimodal information to extract the multimodal features specifically includes: denoising and standardizing the modal data, using an enhanced algorithm for feature extraction, adding a multimodal emotion fusion network feature layer, and analyzing the emotion characteristics of different modalities.

5. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 4, characterized in that: For text data, preprocessing includes: removing irrelevant characters, converting to lowercase, word segmentation, removing stop words, stemming and word form restoration, and feature extraction includes: bag-of-words model and TF-IDF word embedding; for image data, preprocessing includes: unifying image size, normalization and data enhancement, and feature extraction includes: using pre-trained convolutional neural networks and extracting feature maps; for sound data, preprocessing includes: denoising, normalizing volume, framing and applying window functions, and feature extraction includes: Mel-frequency cepstral coefficients, spectrograms, zero-crossing rate and energy features.

6. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 1, characterized in that: The subsequent steps of unifying the fused feature representation of the extracted multimodal features are: combining the global context and historical conversation status to ensure adaptive adjustment of weights, and in order to ensure the accuracy of feature fusion, combining multi-task learning and reinforcement learning strategies to optimize the weight distribution of different tasks, and adjusting the feature weight distribution after fusion of different modalities through feedback.

7. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 1, characterized in that: The steps of determining the long-context dialogue state based on the long-context processing algorithm and ensuring that the generated reply content is associated with the dialogue history specifically include: using an extended version of the Transformer architecture model to process modal data of different lengths, using a memory network mechanism in multiple rounds of dialogue to ensure that information in long conversations is retained and processed, and storing and passing information from previous conversations to subsequent dialogue generation stages.

8. The method for generating multimodal large language model dialogue based on natural language understanding according to claim 1, characterized in that: The method also includes constructing a user personalized portrait driven response generation algorithm based on the user's historical behavior, interest preferences and habits, and in the conversation, using the user's personalized portrait as one of the inputs of the model, combined with the context information of the conversation, to generate a response to the user's personalized needs; It also combines multi-task learning with an adaptive learning strategy of meta-learning to respond to user feedback and adjust model parameters.

Citation Information

Cited By

  • Personalized content generation method of content sharing platform

    CN120179913A

  • AI multi-technology fusion intelligent dialogue model construction method and system for old people

    CN120317381A

  • AI multi-technology fusion intelligent dialogue model construction method and system for elderly care

    CN120317381B

  • MLLM-based energy storage battery data automatic retrieval analysis method and system

    CN120407753A

  • Dynamic intention understanding method based on multi-modal fusion

    CN120579009A