AI multi-modal dialogue system based on multi-modal identification
By adopting cross attention and multi-head attention mechanisms in the multimodal dialogue system, combining sparse coupling and symmetric optimization technology, the limitations of cross-modal feature extraction and dynamic interactive modeling are solved, efficient fusion and intention recognition of text and image information are achieved, and the accuracy and response correlation of the multimodal dialogue system are improved.
Patent Information
- Application Number
- CN202510812668.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing multimodal dialogue system has limitations in cross-modal feature extraction and dynamic interaction modeling, and it is difficult to effectively integrate text and image information, resulting in insufficient accuracy and correlation of intention recognition and response.
The AI multimodal dialogue system based on multimodal recognition is adopted, and text and image features are extracted through natural language single-modal depth encoder and image depth encoder respectively, and dynamic association is constructed using cross-attention and multi-head attention mechanisms. Combining sparse coupling and symmetric optimization technology, efficient fusion and intention recognition of cross-modal information are achieved.
It significantly improves the accuracy and response correlation of multimodal intention understanding, solves the modal divide problem, and realizes accurate multimodal information fusion and intention recognition.
Smart Images

Figure CN120336493A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent conversations, and more specifically, to an AI multi-modal conversation system based on multi-modal recognition. Background Art
[0002] With the development of artificial intelligence, traditional text-based conversation systems are facing limitations as they are unable to process multi-modal information such as images and voices. Human communication is inherently multi-modal, and combining visual, auditory, and other information can enhance the naturalness and intelligence level of interactions. Therefore, building a conversation system that integrates multi-modal information such as text and images has become a trend. By introducing image information, the system can more accurately understand the specific things, scenarios, or emotions referred to by the user, and thus provide more accurate and personalized responses, significantly enhancing the human-computer interaction experience.
[0003] However, effectively integrating information from different modalities is a challenging problem. There are significant differences in structure, features, and semantics among different modalities (modality differences), making simple concatenation difficult to effectively integrate. The correlation between multi-modal information has complex dynamics in the spatio-temporal dimension, and the information value of each modality changes continuously in different stages or tasks. Modality interaction involves multiple levels from low-level feature coupling to high-level semantic complementarity, requiring the fusion strategy to have the ability to handle dynamics and hierarchy. Existing multi-modal conversation systems use relatively simple fusion methods when performing modality fusion, such as direct concatenation of feature vectors or simple attention mechanisms. These methods are difficult to fully address the above challenges, resulting in poor fusion effects. The system is prone to deviation when understanding the cross-modal intentions of users, thereby affecting the accuracy and relevance of responses. Especially in application scenarios that require a fine understanding of the complex correspondence between text and images, existing technologies have deficiencies in extracting and fusing cross-modal features at different granularities and levels, and it is difficult to capture the key information sufficient to support accurate intention recognition, resulting in the overall performance of the system being unable to meet the user's needs for intelligent and natural interactions.
[0004] Therefore, how to design an efficient and robust multi-modal information fusion strategy to overcome the gap between modalities, accurately capture the complex correlations between cross-modal information, and effectively use it for intention recognition and response generation is the key problem faced by current AI multi-modal conversation systems. Summary of the Invention
[0005] In order to solve the above technical problems, this application is proposed.
[0006] According to one aspect of the present application, there is provided an AI multimodal dialogue system based on multimodal recognition, which includes: an input acquisition module for acquiring text information input by a user and image data uploaded by the user; a multimodal encoding module for respectively inputting the text information and the image data into a natural language unimodal depth encoder and an image depth encoder to obtain a set of user intention word granularity semantic encoding vectors and a set of image local semantic encoding feature vectors; an early feature fusion module for inputting the set of user intention word granularity semantic encoding vectors and the set of image local semantic encoding feature vectors into an early feature fusion component based on a cross-attention network to obtain a set of enhanced user intention word granularity semantic encoding vectors and a set of enhanced image local semantic encoding feature vectors; a semantic layer feature fusion module for inputting the set of enhanced user intention word granularity semantic encoding vectors and the set of enhanced image local semantic encoding feature vectors into a semantic layer feature fusion component based on a multi-head attention module to obtain a multimodal fusion representation of user intention; an intention recognition module for inputting the multimodal fusion representation of user intention into an intention recognition classifier to obtain an intention recognition result; and an intelligent response generation module for inputting the intention recognition result and the multimodal fusion representation of user intention into a large language model to obtain a response text.
[0007] Compared with the prior art, the AI multimodal dialogue system based on multimodal recognition provided by the present application first extracts text word granularity and image local features respectively, laying a fine-grained foundation for cross-modal interaction. Secondly, through a bidirectional cross-attention mechanism, a dynamic association between text and image is constructed at the feature level. Among them, attention sparsification processing screens out key cross-modal interaction nodes, and the residual enhancement module alleviates the feature mismatch problem caused by modal differences through a global-local symmetric optimization strategy, strengthening the complementarity of low-level features. Subsequently, the dynamic association of cross-modal high-level semantics is captured through multi-head attention, and its hierarchical processing mechanism enables the system to adapt to the multi-modal information value weights at different task stages. This progressive fusion architecture from the feature layer to the semantic layer, combined with sparse coupling and symmetric optimization technologies, not only solves the modal gap problem caused by simple splicing in traditional methods, but also captures cross-modal spatio-temporal associations through a dynamic attention mechanism. Finally, accurate responses are achieved through intention recognition and large model generation. In this way, the limitations of the prior art in cross-modal feature extraction and dynamic interaction modeling are broken through, and the accuracy of multi-modal intention understanding and response relevance are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other objects, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. They are used together with the embodiments of the present application to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 It is a block diagram of an AI multimodal dialogue system based on multimodal recognition according to an embodiment of the present application.
[0010] Figure 2 It is a block diagram of a multimodal encoding module in an AI multimodal dialogue system based on multimodal recognition according to an embodiment of the present application.
[0011] Figure 3 It is a block diagram of an early feature fusion module in an AI multimodal dialogue system based on multimodal recognition according to an embodiment of the present application.
[0012] Figure 4 It is a block diagram of an image cross-modal feature fusion unit in an AI multimodal dialogue system based on multimodal recognition according to an embodiment of the present application. Detailed implementation manners
[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0014] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0015] In view of the problems in the above background art, the present application is proposed. Figure 1 It is a block diagram of an AI multimodal dialogue system based on multimodal recognition according to an embodiment of the present application. Specifically, as Figure 1As shown in the figure, the AI multimodal dialogue system 100 based on multimodal recognition according to an embodiment of the present application includes: an input acquisition module 110, configured to acquire text information input by a user and image data uploaded by the user; a multimodal encoding module 120, configured to input the text information and the image data into a natural language unimodal depth encoder and an image depth encoder respectively to obtain a set of user intention word-level semantic encoding vectors and a set of image local semantic encoding feature vectors; an early feature fusion module 130, configured to input the set of user intention word-level semantic encoding vectors and the set of image local semantic encoding feature vectors into an early feature fusion component based on a cross-attention network to obtain a set of enhanced user intention word-level semantic encoding vectors and a set of enhanced image local semantic encoding feature vectors; a semantic layer feature fusion module 140, configured to input the set of enhanced user intention word-level semantic encoding vectors and the set of enhanced image local semantic encoding feature vectors into a semantic layer feature fusion component based on a multi-head attention module to obtain a user intention multimodal fusion representation; an intention recognition module 150, configured to input the user intention multimodal fusion representation into an intention recognition classifier to obtain an intention recognition result; and an intelligent response generation module 160, configured to input the intention recognition result and the user intention multimodal fusion representation into a large language model to obtain a response text.
[0016] Specifically, the input acquisition module 110 is configured to acquire text information input by a user and image data uploaded by the user. It should be understood that it is difficult to capture the intention, context, or specific reference object expressed by the user in non-verbal forms such as images relying solely on text information. Given that human communication is naturally multimodal, the user may use text and images simultaneously or alternately to express their needs when initiating a conversation (for example, What's the name of the plant in this picture?). Therefore, acquiring the text information input by the user and the image data uploaded by the user can capture all relevant information input by the user, lay a foundation for subsequent in-depth multimodal understanding, and ensure that the system can receive and integrate the original data streams from different modalities, so as to comprehensively and accurately understand the user's true intention and the situation, which is a prerequisite for realizing intelligent, natural, and efficient multimodal conversations.
[0017] In particular, in a possible example, the acquisition method of the input acquisition module is as follows: it can be completed through the interaction interface provided by the system to the user. Specifically, the user inputs their text message through a text input box. When the user finishes inputting and submits (for example, clicks the send button or presses the enter key), the text information input by the user is captured and transmitted as a raw string data stream to the input acquisition module of the system. At the same time, the interface provides an image upload function, such as an upload picture button or area. When the user clicks this function and selects an image file (for example, an image file in.jpg or.png format) from their local device (such as a computer or mobile phone), the user's device reads the raw binary data of the image file and sends it to the input acquisition module on the server side through the network. The input acquisition module is responsible for listening for and receiving these different data streams from the user interface, including text strings and binary data of images. After receiving the data, the input acquisition module performs preliminary processing on these raw data, such as format checking, and then converts it into a format that can be further processed within the system and passes it to the subsequent multimodal encoding module.
[0018] Specifically, the multimodal encoding module 120 is used to input the text information and the image data into a natural language unimodal deep encoder and an image deep encoder respectively to obtain a set of user intention word granularity semantic encoding vectors and a set of image local semantic encoding feature vectors. Correspondingly, according to the background technology described, the key challenge in constructing a multimodal dialogue system lies in dealing with the modality differences and complex associations between different modality information such as text and images. The original text information (such as a character sequence) and image data (such as a pixel matrix) have essential differences in structure, feature space, and semantic expression methods, and cannot be directly compared or fused effectively. Therefore, inputting the text information and the image data into a natural language unimodal deep encoder and an image deep encoder respectively can convert these original and heterogeneous modality data into their respective high-dimensional and semantically rich feature spaces, providing a unified and standardized semantic representation for subsequent multimodal fusion.
[0019] Specifically, in a specific example of the present application Figure 2 is a block diagram of the multimodal encoding module in the AI multimodal dialogue system based on multimodal recognition according to the embodiments of the present application. As Figure 2As shown, the multimodal encoding module 120 includes: a text segmentation processing unit 121, which is used to perform segmentation processing on the text information to obtain a text word sequence; a text semantic encoding unit 122, which is used to input the text word sequence into the natural language unimodal deep encoder to obtain a set of user intended word granularity semantic encoding vectors, wherein the natural language unimodal deep encoder is a semantic encoder including a Bert model; an image block processing unit 123, which is used to perform image block processing on the image data to obtain an image block sequence; an image semantic encoding unit 124, which is used to input the image block sequence into the image depth encoder to obtain a set of image local semantic encoding feature vectors, wherein the image depth encoder is a Vit model including an image block embedding coding layer.
[0020] It should be understood that text information, as an unstructured character sequence, is difficult to be directly understood and processed by deep learning models in its original form to extract deep semantic features. Natural language processing models, especially deep encoders based on Transformer structures, are trained and inferred based on discrete vocabulary units or their subunits. Therefore, word segmentation of the text information to obtain a text word sequence can decompose the continuous text string input by the user into a series of basic units (i.e., words or tokens) with independent semantic meanings or constituting semantic units, forming an ordered sequence, thereby providing data support for the subsequent extraction of semantic encoding features of user intentions at the word granularity.
[0021] In particular, in a possible example, the text segmentation processing unit 121 is implemented as follows: first, preprocessing is performed, such as removing noise in the text (such as standardization of special symbols and blank characters), full-width and half-width conversion, etc.; then, the preprocessed text is segmented according to a preset segmentation rule or dictionary. Commonly used segmentation methods include dictionary-based maximum matching methods, methods based on statistical models (such as hidden Markov models, conditional random fields), or more advanced methods based on deep learning models. For example, for Chinese text, since there are no obvious separators between words, the segmentation unit will identify the word boundaries in the text according to the built-in vocabulary and language model, and segment the continuous Chinese character sequence into independent words. For example, the input text information is: What is the name of the plant in this picture? After the segmentation process, the text word sequence may be ["this", "piece", "picture", "in", "of", "plant", "called", "what", "?"]. The word segmentation method and dictionary used will be selected and set according to the specific language and application scenario, aiming to generate word sequences that can accurately reflect the semantic structure of the text and provide high-quality input for the subsequent semantic encoding step.
[0022] Accordingly, the original text information, after being segmented, results in a discrete sequence of text words. Although word segmentation divides continuous text into word units, these words themselves are still symbolic representations and cannot directly capture their deep semantics, context relationships, and subtle meanings in the expression of user intentions. Different words may have different meanings in different contexts, and simple symbolic representations cannot reflect this context dependence. Therefore, inputting the text word sequence into the natural language unimodal depth encoder to obtain a set of user intention word-level semantic encoding vectors can transform these discrete word symbols into continuous, high-dimensional, and semantically rich vector representations. Through the natural language unimodal depth encoder, the system can learn and extract the deep semantic features of each word or token in the text within the entire sentence and even broader context, thereby obtaining more expressive user intention word-level semantic encoding vectors, providing a basis for subsequent cross-modal fusion and intention recognition.
[0023] In particular, in a possible example, the implementation of the text semantic encoding unit 122 is as follows: For example, for the input text (What's the name of the plant in this picture?), after word segmentation, the word sequence ["this", "picture", "in", "the", "plant", "is", "called", "what", "?"] can be obtained. This word sequence is input into the natural language unimodal depth encoder, which contains a Bert model. The Bert model is a bidirectional encoder based on the Transformer architecture, and its core consists of multiple identical Transformer encoder layers stacked together. Before inputting the word sequence into the Bert model, tokenization will be performed first to map the word sequence to the token sequence in the vocabulary preset by the Bert model. Each token is converted into its corresponding initial token embedding vector, and combined with its position information in the sequence (position encoding) and possible segment information (segment encoding, if dealing with multi-segment text) to form the input representation. These input representations then pass through multiple Transformer encoder layers of the Bert model. In each layer, using the multi-head self-attention mechanism, the model can calculate the association strength between each token in the sequence and all other tokens, so as to learn the contextualized representation of each token in the global context. Then, these contextualized representations are further non-linearly transformed through a feed-forward neural network. After being processed by all Transformer layers, the Bert model outputs a vector sequence of the same length as the input token sequence, where each vector is the final contextualized semantic encoding of the corresponding token. For example, for the input token sequence, "this", "picture", "in", "the", "plant", "is", "called", "what", "?". The Bert model will output a set of 9 vectors, each vector corresponding to the semantic encoding of a token, and the obtained vector sequence is the set of the user intention word granularity semantic encoding vectors. These vectors capture the semantic and syntactic information of each word or token in the entire text, providing a fine-grained text feature representation for subsequent modality fusion. The Bert model is pre-trained on a large-scale text corpus to learn general language understanding capabilities, and then can be fine-tuned on specific tasks (such as multi-modal dialogue in this system) to better adapt to the task requirements.
[0024] It should be understood that the original image data exists in the form of a pixel matrix, and there are significant modal differences between this two-dimensional spatial structure and the linear sequence structure of text information. It is difficult to directly use the original pixel data for subsequent cross-modal fusion and semantic understanding because the pixel-level information is too low-level to directly reflect the semantic content of objects, scenes, or their local regions in the image. To be able to process the image data, the image data needs to be converted into a serialized form that can be aligned or interacted with the text sequence to some extent, so that the image data can be processed by the subsequent encoder to extract features with local semantic meanings.
[0025] Specifically, in a possible example, the implementation of the image chunk processing unit 123 is as follows: First, a preset image chunk size is determined. For example, the size of the image chunk is set to 16x16 pixels. This size is an important hyperparameter, and its setting depends on the specific model design, computing resources, and the requirement for the granularity of local features, and is determined before model training. Then, the image chunk processing unit will divide the input original image into multiple small image chunks without overlap in the horizontal and vertical directions according to this preset chunk size. For example, for an RGB image of 224x224 pixels, if the image chunk size is set to 16x16 pixels, then it will be divided into 224 / 16 = 14 chunks in the horizontal direction and also 14 chunks in the vertical direction, resulting in a total of 14*14 = 196 image chunks. After the division, each image chunk is a pixel matrix of a fixed size (e.g., 16x16x3 if the image is in RGB format). Finally, these two-dimensional image chunks are arranged in a certain order (e.g., from left to right, from top to bottom), and the pixels inside each image chunk are flattened into a one-dimensional vector (e.g., 16x16x3 = 768-dimensional vector), thus forming an ordered sequence composed of multiple image chunk vectors. This sequence composed of the flattened image chunk vectors is the image chunk sequence. For example, for the above 224x224 image, after processing, a sequence of length 196 will be obtained, and each element in the sequence is a 768-dimensional vector.
[0026] Correspondingly, although the image is organized into an image chunk sequence after chunk processing, each chunk is still a flattened representation of the original pixel data. Such low-level pixel features fail to capture high-level visual information such as the semantic content, texture, and shape of the local regions of the image, nor can they reflect the spatial relationships and interactions between the image chunks. To enable the system to understand the image content and perform effective cross-modal fusion with the text information, these original image chunks need to be converted into feature vectors with rich semantic expression capabilities.
[0027] In particular, in a possible example, the implementation of the image semantic encoding unit 124 is as follows: For an image divided into 196 16x16 pixel image patches, the input is a sequence containing 196 flattened image patch vectors, and each vector is, for example, 768-dimensional. The image semantic encoding unit inputs this sequence of image patches into the image depth encoder, which contains a ViT model based on an image patch embedding encoding layer. The ViT model first processes the input sequence of image patches through an image patch embedding encoding layer. This embedding encoding layer is a linear projection layer that maps each flattened image patch vector (e.g., 768-dimensional) to a lower-dimensional embedding space (e.g., setting the embedding dimension to 768 or 1024 dimensions) to obtain image patch embedding vectors. This process converts the original pixel values into a more abstract and compact feature representation. To preserve the spatial location information of the image patches (since the self-attention mechanism of the Transformer is itself position-independent), the ViT model adds learnable or fixed position encodings to the image patch embedding vectors. These image patch embedding vectors with position information are then input into the Transformer encoder part of the ViT model, which consists of multiple stacked self-attention layers and feed-forward neural networks. Through the self-attention mechanism, the model can calculate the associations between each image patch embedding in the sequence and all other image patch embeddings, thereby learning the contextualized representation of each local region in the context of the entire image. After being processed by all the Transformer layers, the ViT model outputs a sequence of 196 image patch vectors, where each image patch vector is the final contextualized semantic encoding of the corresponding image patch. This vector sequence is the set of the image local semantic encoding feature vectors. For example, for an input sequence of 196 image patches, the ViT model will output a set containing 196 vectors, and each vector corresponds to the local semantic encoding of an image patch. The ViT model is pre-trained on a large-scale image dataset (such as ImageNet) to learn the ability to extract visual features, and then can be fine-tuned according to the characteristics of the multi-modal dialogue task.
[0028] Specifically, the early feature fusion module 130 is configured to input the set of user intention word-level semantic encoding vectors and the set of image local semantic encoding feature vectors into an early feature fusion component based on a cross-attention network to obtain an enhanced set of user intention word-level semantic encoding vectors and an enhanced set of image local semantic encoding feature vectors. It should be understood that although the text information and image data have been respectively converted into their respective semantic encoding vector sets through single-modal deep encoders, these vectors have rich semantic information within their respective modalities, but they are still independently generated representations, lacking direct interaction and association between modalities. The user's intention in a multi-modal conversation is often the result of the collaborative expression of text and image information. For example, the text may refer to a specific object in the image, or the image provides the context of the text. Simple concatenation or independent semantic representations cannot capture this cross-modal fine-grained correspondence and mutual influence. Therefore, at this relatively early feature level, it is necessary to design a precise interaction mechanism to enable fine-grained information exchange and mutual enhancement between features of different modalities.
[0029] Specifically, in a specific example of the present application, Figure 3 is a block diagram of an early feature fusion module in an AI multi-modal conversation system based on multi-modal recognition according to an embodiment of the present application. As Figure 3 shown, the early feature fusion module 130 includes: a user intention semantic vector extraction unit 131, configured to extract a first user intention word-level semantic encoding vector from the set of user intention word-level semantic encoding vectors; an image-to-early cross-modal attention calculation unit 132, configured to calculate the early feature cross-modal attention values of each image local semantic encoding feature vector in the set of image local semantic encoding feature vectors relative to the first user intention word-level semantic encoding vector to obtain a set of image-to-early feature cross-modal attention values; an image cross-modal feature fusion unit 133, configured to fuse the set of image local semantic encoding feature vectors based on the set of image-to-early feature cross-modal attention values to obtain an image-to-cross-modal early feature attention interaction encoding vector; and a user intention residual enhancement unit 134, configured to input the image-to-cross-modal early feature attention interaction encoding vector and the first user intention word-level semantic encoding vector into a residual module to obtain an enhanced user intention word-level semantic encoding vector corresponding to the first user intention word-level semantic encoding vector.
[0030] Accordingly, the text semantic encoding unit outputs a set of user intention word-level semantic encoding vectors, which contains the vector representations obtained after deep encoding of each word or token in the text sequence. In the subsequent early feature fusion stage, a cross-attention network-based approach is adopted, where elements of one modality (such as text) need to be used as queries or key-value pairs to attend to elements of another modality (such as images). In the design of some cross-attention mechanisms, a representative vector is extracted from the query modality as the overall query or starting point, which can guide the attention process to focus on the most relevant information in the query modality for interaction with the key-value pair modality. Therefore, a specific vector in the text modality needs to be obtained as the starting point for calculating the early cross-modal attention between the local semantic features of text and images, thereby driving information to flow from the image modality to the text modality (or vice versa, depending on the specific direction of cross-attention) to achieve early interaction and enhancement of features.
[0031] Specifically, in a possible instance, the implementation of the user intention semantic vector extraction unit 131 is as follows: This set is an ordered sequence of vectors corresponding to the token sequence obtained after tokenization of the original text. For example, if the original text (What's the name of the plant in this picture?) is processed to obtain a token sequence, and the natural language single-modal deep encoder outputs a set of vectors {vthis, vpicture, vthe, vplant, vin, vname, vwhat, v?} corresponding to this sequence, where each vtoken is a high-dimensional (e.g., 768-dimensional) semantic encoding vector. The user intention semantic vector extraction unit will select and extract a specific vector from this set of vectors as the first user intention word-level semantic encoding vector. This extracted vector is then used to drive the early cross-modal attention calculation from the image to the text.
[0032] It should be understood that although the text and images have been encoded into sets of semantic vectors respectively, in the early fusion stage, it is necessary to establish cross-modal associations so that the features of one modality can perceive and utilize relevant information of the other modality. In particular, to better align the local features of the image and enhance the part related to the user's intention expressed in the text, the system needs to know which regions in the image are more relevant to the specific vector representing the user's core intention in the text. Therefore, it is necessary to quantify the similarity or degree of association between each local region in the image (represented by its semantic encoded feature vector) and the representative semantic vector extracted from the text, that is, to calculate the early feature cross-modal attention value of each image local semantic encoded feature vector in the set of image local semantic encoded feature vectors with respect to the first user intention word granularity semantic encoded vector. This similarity metric will be converted into attention weights, indicating the importance of each image local feature in response to a specific text intention, thereby guiding the subsequent feature fusion process to achieve guided enhancement of image features by text.
[0033] In particular, in one possible instance, the implementation of the image to the early cross-modal attention calculation unit 132 is as follows: This unit receives two inputs: the first user intention word granularity semantic encoded vector provided by the user intention semantic vector extraction unit (as the query in the attention mechanism, abbreviated as the query vector) and the set of image local semantic encoded feature vectors provided by the image semantic encoding unit (as the keys in the attention mechanism, abbreviated as the key vector set). For example, the first user intention word granularity semantic encoded vector is a high-dimensional vector representing the text intention, and its dimension is set to 768; the set of image local semantic encoded feature vectors contains multiple high-dimensional vectors representing local regions of the image, such as 196 768-dimensional vectors.
[0034] The implementation process is based on the scaled dot-product attention mechanism. First, the input query vector is projected into the query space through a learnable linear transformation (implemented by a weight matrix). At the same time, each image local semantic encoded feature vector in the input key vector set is projected into the key space through another learnable linear transformation (implemented by a weight matrix). The parameters of these two weight matrices for linear transformation are automatically learned and determined during the system training process by optimizing the overall task objective. Their role is to map the features of different modalities into the same compatible representation space for effective similarity comparison.
[0035] Next, calculate the dot product similarity between the projected query vector and each projected key vector to obtain the original attention scores. For example, for a set of 196 image local feature vectors, 196 original attention scores will be obtained, each corresponding to the similarity between an image local feature and the text query vector. To stabilize the calculation and avoid overly large dot product results in higher dimensions, divide each original attention score by the square root of the dimension of the projected key vector. Finally, apply the Softmax function to all the scaled scores. The Softmax function transforms these scores into a set of probability values such that the sum of the probability values corresponding to all image patches is 1. This set of probability values output by the Softmax function is the set of cross-modal attention values of the image towards the early features. This set reflects the degree of attention or importance weight of the specific text intention to each local region of the image.
[0036] Correspondingly, the set of image local semantic encoding feature vectors contains the semantic information of each local region after the image is segmented, while the set of cross-modal attention values of the image towards the early features calculated in the previous step quantifies the association strength or importance between each image local feature and the first user intention word granularity semantic encoding vector. In multimodal interaction, the user's text intention often targets specific objects or regions in the image. Simply using the mean or concatenation of all image local features to represent the image may introduce a large amount of noise information unrelated to the text intention and fail to highlight the visual elements closely related to the text content. Therefore, it is necessary to use the calculated attention values as weights to perform weighted aggregation on the set of image local semantic encoding feature vectors.
[0037] Specifically, in an implementable example of the present application, Figure 4 is a block diagram of an image cross-modal feature fusion unit in an AI multimodal dialogue system based on multimodal recognition according to an embodiment of the present application. As Figure 4 shown, the image cross-modal feature fusion unit 133 includes: an attention weight value calculation sub-unit 1331, configured to input the set of cross-modal attention values of the image towards the early features into the Softmax activation function to obtain a set of cross-modal attention weight values of the image towards the early features; an attention weight value sparsification sub-unit 1332, configured to input the set of cross-modal attention weight values of the image towards the early features into an attention sparsity module based on a preset threshold to obtain a sparse set of cross-modal attention weight values of the image towards the early features; an image cross-modal weighted sub-unit 1333, configured to use the sparse set of cross-modal attention weight values of the image towards the early features as weights to calculate the weighted sum of the set of image local semantic encoding feature vectors to obtain the cross-modal early feature attention interaction encoding vector of the image.
[0038] It should be understood that in the previous step, the similarity or association strength between each local region in the image and a specific vector representing the user's intention in the text was calculated, obtaining a set of cross-modal attention values from the image to early features. These values reflect the original correlation scores of each local image feature with the text intention. However, in order to effectively utilize these scores to fuse local image features in subsequent steps, it is necessary to convert these scores within arbitrary ranges into a form with a clear physical meaning, namely weights, such that these weights can represent how much information each local image feature should contribute during the fusion process, and the sum of these contributions is normalized.
[0039] Specifically, in a possible instance, the implementation of the attention weight value calculation sub-unit 1331 is as follows: This sub-unit receives as input the set of cross-modal attention values from the image to early features output by the previous step. The Softmax activation function is applied to this set of numerical values. The role of the Softmax function is to convert a set of arbitrary real-valued numbers into a probability distribution. The specific calculation method is that for each numerical value in the input set, first calculate its exponential (the power of the natural constant e to this numerical value). Then, add up all these exponential calculation results to obtain a sum. Finally, divide the exponential result of each numerical value by this sum. For example, if the input set is {s1, s2,..., sn}, then in the output weight set {h1, h2,..., hn}, the calculation method for each weight hi is: hi is equal to e to the power of si, and then divided by the sum of e to the power of s1, e to the power of s2 until e to the power of sn. After this calculation process, the output obtained is the set of cross-modal attention weight values from the image to early features. This set also contains the same number of numerical values as the input set (for example, 196 weight values), each numerical value is a non-negative real number between 0 and 1, and the sum of all numerical values in the set is strictly equal to 1. These weight values directly indicate the relative contribution degrees of each local image feature under the drive of the text intention.
[0040] Accordingly, the set of cross-modal attention weight values of the image towards early features obtained through the Softmax function assigns a weight to each local region of the image. However, in an actual multi-modal scenario, the user's text intention may only be highly relevant to a few key regions in the image, while most regions in the image may be background or information unrelated to the intention. Even for non-relevant regions, Softmax may assign a very small non-zero weight to them. Directly using all these non-zero weights for fusion will, on the one hand, introduce noise unrelated to the intention and reduce the fusion efficiency; on the other hand, processing all local features consumes a large amount of computing resources. Therefore, in this application, the set of cross-modal attention weight values of the image towards early features is input into an attention sparsity module based on a preset threshold to obtain a sparse set of cross-modal attention weight values of the image towards early features, so as to identify and retain the weights corresponding to those key image local features that are truly highly relevant to the text intention, while eliminating or suppressing the weights corresponding to those regions with lower weights that are considered unrelated to the text intention.
[0041] Specifically, in a possible example, the attention weight value sparsification subunit 1332 is used to: perform attention sparsification processing on the set of cross-modal attention weight values of the image towards early features according to the following formula to obtain the sparse set of cross-modal attention weight values of the image towards early features, where the formula is: ; where are the cross-modal attention weight values of the image towards early features in the set of cross-modal attention weight values of the image towards early features, is the preset threshold, is the attention sparsification processing, are the cross-modal attention weight values of the image towards early features in the sparse set of cross-modal attention weight values of the image towards early features. Specifically, the preset threshold here is a floating-point number between 0 and 1, and its specific value is determined by optimizing the model performance during training. After processing, only those local image regions with attention weights greater than or equal to the preset threshold will have their corresponding weight values retained (or remain unchanged), while the regions with weight values less than will be ignored and their weights become zero, thus forming a sparse attention weight set.
[0042] It should be understood that the set of the local semantic encoding feature vectors of the image contains the semantic information of all local regions of the image, while the sparse set of the cross-modal attention weight values of the image to the early features obtained in the previous step clearly indicates which local regions in the image are highly relevant to the first user intention word granularity semantic encoding vector (representing the text intention), and quantifies the relative importance of these relevant regions. In order to integrate these scattered local image information into a single and at the same time fusion representation that can highlight the text intention focus, it is necessary to use the sparse attention weight as a guiding mechanism to perform biased aggregation on the local image features. The obtained cross-modal early feature attention interaction encoding vector of the image is a more compact, information more focused, and fusion representation enhanced by the text intention information, providing a more targeted and robust basic feature for subsequent cross-modal understanding or task execution.
[0043] Specifically, in a possible example, the user intention residual enhancement unit 134 includes: a sparse coupling symmetric optimization subunit, configured to perform global-local sparse coupling symmetric optimization on the cross-modal early feature attention interaction encoding vector of the image and the first user intention word granularity semantic encoding vector to obtain a cross-modal early feature attention interaction optimized encoding vector of the image and a first user intention word granularity semantic optimized encoding vector; a first user intention feature enhancement subunit, configured to input the cross-modal early feature attention interaction optimized encoding vector of the image and the first user intention word granularity semantic optimized encoding vector into a residual module to obtain an enhanced user intention word granularity semantic encoding vector corresponding to the first user intention word granularity semantic encoding vector.
[0044] It should be noted that the cross-modal early feature attention interaction encoding vector of the image is a global image semantic attention interaction feature representation that fuses the set of the local semantic encoding feature vectors of the image, while the first user intention word granularity semantic encoding vector is a local text semantic encoding feature representation. When the two are input into the residual module, the first user intention word granularity semantic encoding vector will cause significant sparsification of the semantic space due to interpolation. Although the attention sparse module based on the preset threshold also performs attention weighted sparsification processing on the set of the local semantic encoding feature vectors of the image, there is still a significant sparsity imbalance between the two, thus affecting the expression effect of the enhanced user intention word granularity semantic encoding vector obtained through the residual module. Therefore, it is necessary to perform global-local sparse coupling symmetric optimization on the cross-modal early feature attention interaction encoding vector of the image and the first user intention word granularity semantic encoding vector.
[0045] Based on this, in an achievable example of the present application, the sparse coupling symmetric optimization subunit is used to: perform global sparsity residual modeling on the image towards the cross-modal early feature attention interaction coding vector and the first user intention word granularity semantic coding vector to obtain the image towards the cross-modal early feature attention interaction coding global vector and the first user intention word granularity semantic coding global vector, that is: ; where is the image towards the cross-modal early feature attention interaction coding vector, is the first user intention word granularity semantic coding vector, is pointwise subtraction by position, is the vector two-norm, is pointwise addition by position, is to calculate the square of each eigenvalue in the vector, is the value of the logarithmic function with the natural constant e as the base, is pointwise addition by position, is absolute value calculation, is the image towards the cross-modal early feature attention interaction coding global vector, is the first user intention word granularity semantic coding global vector; that is, based on the residual probability density representation, integration is performed in their respective semantic spaces to obtain the global characteristics of the sparsity difference through spatial integration modeling.
[0046] Taking the image towards the cross-modal early feature attention interaction coding vector and the first user intention word granularity semantic coding vector as the local space coordinate bases, calculate the smoothness mutual constraint factor of the global sparsity residual to obtain the image towards the cross-modal constraint factor and the first user intention constraint factor, that is: ; where is the image towards the cross-modal constraint factor, is the first user intention constraint factor.
[0047] Based on the image towards the cross-modal constraint factor and the first user intention constraint factor, perform symmetric transformation optimization on the image towards the cross-modal early feature attention interaction coding vector and the first user intention word granularity semantic coding vector to obtain the image towards the cross-modal early feature attention interaction optimization coding vector and the first user intention word granularity semantic optimization coding vector, that is: ; where is the image towards the cross-modal early feature attention interaction optimization coding vector, It is the first user intention word granularity semantic optimization encoding vector. In this way, through the residual coupling of the global sparsity on the local probability density and the smoothness constraint, the geometric characteristics of the feature manifold in the semantic space can be guaranteed, so that the problem of sparsity imbalance is corrected through the symmetric transformation of the semantic class space, and the expression effect of the enhanced user intention word granularity semantic encoding vector obtained through the residual module is improved.
[0048] Correspondingly, in the cross-modal interaction process, the user's text intention often needs to be refined and verified through image information. Although the image local feature fusion vector based on the text intention (the image cross-modal early feature attention interaction optimization encoding vector) has been calculated in the previous step, this is still a representation focusing on the image side, while the original text intention vector (the first user intention word granularity semantic optimization encoding vector) may contain more abstract or global semantic information. In order to make the final text intention representation fully integrate the relevant visual evidence and become more accurate, specific and robust, it is necessary to use the image-side interaction information in reverse to enhance the original text intention representation and correct and supplement the original text intention vector.
[0049] Specifically, in a possible example, the implementation of the first user intention feature enhancement sub-unit is as follows: First, the image cross-modal early feature attention interaction optimization encoding vector and the first user intention word granularity semantic optimization encoding vector are concatenated in the feature dimension to form a joint vector with a larger dimension (for example, the dimension after concatenation is 1536). This concatenated vector contains the image information and the original text intention information obtained after the early cross-modal interaction. Then, the concatenated vector is input into a feed-forward neural network. The feed-forward neural network is the main processing path in the residual structure, consisting of one or more fully connected layers (linear transformation) and non-linear activation functions (such as ReLU), and its parameters (weights and biases) are learned during the training process. The role of the feed-forward neural network is to learn a residual representation for the original text intention vector from the joint vector, and the dimension of this residual representation is designed to be the same as the dimension of the first user intention word granularity semantic optimization encoding vector (for example, 768). Finally, the residual vector output by the feed-forward neural network is added element-wise to the original input first user intention word granularity semantic optimization encoding vector. This summation operation is the core of the residual connection. The resulting vector after addition is the enhanced user intention word granularity semantic encoding vector corresponding to the first user intention word granularity semantic encoding vector. In this way, the original text intention vector is enhanced with the relevant information on the image side, generating a more discriminative and context-aware fusion representation.
[0050] In particular, in another specific example of the present application, the early feature fusion module further includes: an image semantic vector extraction unit, configured to extract a first image local semantic coding feature vector from the set of image local semantic coding feature vectors; a text-to-early cross-modal attention calculation unit, configured to calculate the early feature cross-modal attention values of each user intention word granularity semantic coding vector in the set of user intention word granularity semantic coding vectors with respect to the first image local semantic coding feature vector to obtain a set of text-to-early feature cross-modal attention values; a text cross-modal feature fusion unit, configured to fuse the set of user intention word granularity semantic coding vectors based on the set of text-to-early feature cross-modal attention values to obtain a text-to-cross-modal early feature attention interaction coding vector; an image semantic residual enhancement unit, configured to input the text-to-cross-modal early feature attention interaction coding vector and the first image local semantic coding feature vector into a residual module to obtain an enhanced image local semantic coding feature vector corresponding to the first image local semantic coding feature vector. It should be understood that although the original set of image local semantic coding feature vectors contains visual information of each region of the image, this information is purely at the visual level and does not directly incorporate the user's text intention information. However, the user's intention is directed at a specific object or region in the image and is precisely expressed by the text. To enable the local representation of the image to perceive and reflect the visual content related to the user's text intention, thereby enhancing the effectiveness of subsequent cross-modal fusion and understanding, the present application inputs the set of user intention word granularity semantic coding vectors and the set of image local semantic coding feature vectors into an early feature fusion component based on a cross-attention network and performs local enhancement of the image to introduce text information, so that each local feature of the image can learn and absorb the text semantic context related to it, thereby enhancing the expression of the image local semantic coding feature vector. In particular, the specific implementation processes of these three units are the same as those of the user intention semantic vector extraction unit, the image-to-early cross-modal attention calculation unit, the image cross-modal feature fusion unit, and the user intention residual enhancement unit in the early feature fusion module to obtain the enhanced user intention word granularity semantic coding vector.
[0051] Specifically, the semantic layer feature fusion module 140 is configured to input the set of enhanced user intention word granularity semantic encoding vectors and the set of enhanced image local semantic encoding feature vectors into the semantic layer feature fusion component based on the multi-head attention module to obtain the user intention multi-modal fusion representation. Correspondingly, through the processing of the early feature fusion component based on the cross-attention network, a set of enhanced user intention word granularity semantic encoding vectors and a set of enhanced image local semantic encoding feature vectors that incorporate some cross-modal information are obtained. Although the vectors in these sets already contain enhanced information from the other modality, they are still scattered representations corresponding to the word granularity of the text or the local regions of the image. To finally obtain a single multi-modal understanding representation that can comprehensively and integrally represent the user's overall intention and image-related content, it is necessary to effectively converge and fuse these enhanced feature information distributed in different modalities and different local regions, so that the finally generated user intention multi-modal fusion representation will be the product of the deep fusion of text and image information, directly serving subsequent intention understanding or task execution.
[0052] Specifically, in a possible example, the implementation of the semantic layer feature fusion module 140 is as follows: First, the set of enhanced user intention word granularity semantic encoding vectors and the set of enhanced image local semantic encoding feature vectors are concatenated into a unified vector sequence. That is, all M text vectors and all N image vectors are arranged in a certain order to form a sequence with a total length of M + N. This concatenated sequence is used as the input to the multi-head attention module. The multi-head attention module contains a number of parallel attention heads. For example, the number of heads can be preset to 8, and this number is a hyperparameter determined by experimental optimization. Inside each attention head, first, the input vector sequence is mapped to query (Query), key (Key), and value (Value) vector sequences through different linear transformations (i.e., multiplying with different learnable weight matrices and adding biases). The parameters of these linear transformations are learned by the model during training. Then, the matrix product of the query and the transpose of the key is calculated to obtain the attention scores. Next, the attention scores are scaled (divided by the square root of the key vector dimension) and the normalized exponential function (Softmax) is applied to obtain the attention weights. Finally, the attention weight matrix is multiplied by the value vector sequence to obtain the output sequence of this attention head. The length of this output sequence is also M + N, and the dimension of each vector is D divided by the number of heads (if divisible, or the dimension is adjusted by a subsequent linear layer). For example, 768 / 8 = 96 dimensions. The output sequences of all attention heads are concatenated in the last dimension to form a higher-dimensional sequence (e.g., 8 * 96 = 768 dimensions). This concatenated sequence is then passed through a final linear transformation to map its dimension back to the original vector dimension D (e.g., 768 dimensions), obtaining the total output sequence of the multi-head attention module, with a length still of M + N and a dimension of D. To obtain a single user intention multi-modal fusion representation from this sequence containing M + N vectors, an aggregation operation is required. A commonly used method is global average pooling, that is, calculating the average value of all M + N vectors in the multi-head attention module output sequence for each dimension. For example, adding and averaging the first dimension values of all vectors to obtain the first dimension value of the result vector; adding and averaging the second dimension values of all vectors to obtain the second dimension value of the result vector, and so on, until all D dimensions are calculated. The resulting final single vector, that is, the user intention multi-modal fusion representation vector, is the user intention multi-modal fusion representation. It synthesizes all enhanced text word granularity information and image local feature information, and captures the complex interaction relationship between them by the multi-head attention mechanism.
[0053] Specifically, the intent recognition module 150 is configured to input the multi-modal fusion representation of the user intent into an intent recognition classifier to obtain an intent recognition result. It should be understood that through the multi-level cross-modal attention mechanism in the previous steps, the text intent of the user has been deeply fused with the key information of the image, and finally a compact and information-rich single vector, that is, the multi-modal fusion representation of the user intent, has been generated. This vector comprehensively reflects the meaning of the user intent in the image context. In order for the system to make specific judgments or take actions based on this fusion representation, it is necessary to convert this high-dimensional numerical representation into a discrete and understandable intent category. The role of the classifier is to convert complex numerical features into explicit semantic labels, so as to explicitly identify the ultimate purpose or requirement expressed by the user input (text + image), enabling the system to understand the true intent of the user and take corresponding subsequent processing steps. In particular, the intent recognition classifier can be a feed-forward neural network, for example, it can be a multi-layer perceptron including one or more hidden layers. A non-linear activation function is used between the hidden layers, and the output layer is a linear layer.
[0054] Specifically, in a possible example, the implementation of the intention recognition module 150 is as follows: The user intention multi-modal fusion representation vector (user intention multi-modal fusion representation) is input into the intention recognition classifier. For example, the classifier is a feed-forward network with one hidden layer. The input vector first passes through the first linear transformation layer. In this layer, the input vector is multiplied by a learnable weight matrix and then added with a learnable bias vector to obtain an intermediate vector. For example, if the dimension of the hidden layer is set to 512 dimensions, the dimension of the weight matrix is (512×768), and the dimension of the bias vector is 512. Then, a non-linear activation function, such as the rectified linear unit (ReLU), is applied to this intermediate vector to introduce non-linearity. The activated vector is then input into the output linear transformation layer. In this layer, the vector is multiplied by another learnable weight matrix and added with a learnable bias vector. The dimension of the weight matrix of the output layer is (number of intention categories×512), and the dimension of the bias vector is the number of intention categories. The number of intention categories is equal to the total number of user intentions predefined by the system that need to be recognized, such as 100 different intention categories. The output of this linear layer is a vector with a dimension equal to the number of intention categories, and each element represents the raw prediction score of the model for the corresponding intention category. Then, the normalized exponential function (Softmax) is applied to this prediction score vector. The Softmax function converts the scores into a probability distribution, where the value of each element ranges from 0 to 1, and the sum of all elements is 1. Each probability value represents the likelihood that the model predicts the user input belongs to the corresponding intention category. What the intention recognition result specifically contains depends on the design of the system. The most basic is the identified intention category label. It also includes the confidence score corresponding to this prediction label (i.e., the highest probability value output by Softmax). For example, if the user input text is: What's the name of the plant in this picture?, and provides a picture containing a specific plant. After the previous cross-modal fusion step, the user intention multi-modal fusion representation containing the text "plant" and the plant features in the image is generated. When this fusion vector is input into the intention recognition classifier, the classifier will calculate the scores for all predefined intention categories. For example, the predefined intention categories of the system include identifying plants, searching for products, navigation, etc. After Softmax, the model may output probabilities for each category, such as the probability of identifying plants is 0.98, searching for products is 0.01, navigation is 0.005, etc. At this time, the intention recognition result is the label "identifying plants", and it will include the confidence of 0.98. All learnable parameters (weight matrices and bias vectors) of the classifier are automatically determined during the model training phase by learning a large number of user multi-modal input samples with labeled intention categories (such as identifying plants).
[0055] Specifically, the intelligent response generation module 160 is used to input the intention recognition result and the multimodal fusion representation of the user's intention into a large language model to obtain a response text. In particular, the previous steps have successfully fused the user's textual intention with the relevant visual information in the image into a single, high-dimensional multimodal fusion representation of the user's intention through complex cross-modal interaction and feature fusion, and clarified the user's main intention category through the intention recognition classifier. However, these internal representations and classification results are machine-understandable and insufficient to directly communicate with the user in a natural and humanized manner. What the user expects is a response text expressed in natural language that can accurately and fluently respond to the intention and content contained in its multimodal input. Therefore, it is necessary to input the intention recognition result and the multimodal fusion representation of the user's intention into a large language model to obtain a response text and convert it into a text output that conforms to human language habits. It is worth mentioning that the large language model, with its powerful text generation ability and ability to understand complex contexts, can comprehensively consider the user's specific intentions, the relevant details in the image, and the deep meaning after their fusion, thereby generating a coherent, relevant and high-quality response text to achieve effective interaction between the system and the user.
[0056] Specifically, in a possible example, the implementation of the intelligent response generation module 160 is as follows: First, it is necessary to convert the intention recognition result and the multi-modal fusion representation of the user intention into an input sequence form that can be processed by the large language model. The intention recognition result, such as the recognized intention label "recognize plants", will be converted into a series of input tokens predefined by the model, such as a special token representing the intention type, or a text string describing the intention (such as the user hopes to recognize the plants in the picture). These tokens will be mapped to the corresponding sequence of embedding vectors through the token embedding layer of the large language model. The vector of the multi-modal fusion representation of the user intention, as a continuous high-dimensional vector, needs to be processed through a dedicated linear projection layer. This linear projection layer maps the 768-dimensional fusion vector to the embedding vector dimension that matches the internal representation dimension of the large language model (for example, if the hidden state dimension of the large language model is 1024 dimensions, it is mapped to a 1024-dimensional vector through a weight matrix and bias of 768x1024). Then, these transformed and projected sequences of embedding vectors (including the embeddings from the intention recognition result and the multi-modal fusion representation, and possibly also the embeddings of the user's original text) are concatenated in a predetermined order to form a complete input sequence. This sequence is input into the main structure of the large language model, namely the multi-layer Transformer module. In these layers, through the self-attention mechanism and the feed-forward network, the model can fully process all the information in the input sequence, understand the meaning of the intention label, the context information carried by the multi-modal fusion representation, and establish complex associations between them. The output sequence processed by the Transformer layer contains the deep understanding representation of the model for the input context. Based on the representation processed by the Transformer layer, the large language model enters the text generation stage. The model predicts and generates the next token one by one in an autoregressive manner according to the input context and the currently generated token sequence. This process starts from a starting token (such as a special token representing the start of generation) until an end token is generated or the preset maximum length is reached. Predicting the next token is achieved by passing the output of the last layer of the Transformer through a linear layer and the Softmax function to obtain the probability distribution of each token in the vocabulary as the next token by the model, and then selecting the token with the highest probability or by a sampling strategy. The generated response text is to decode and restore this predicted token sequence into natural language.
[0057] For example, if the user input text is "What's the name of the plant in this picture?" and provides a picture containing a specific plant. The intent recognition classifier outputs an intent recognition result of recognizing the plant, and the user intent multi-modal fusion representation vector contains the fusion information of the plant in the text and the image. Convert the intent of recognizing the plant into an embedding, project the fusion vector into an embedding, and concatenate them and input them into a large language model. The large language model processes this information and understands that the user wants to know the name of the plant in the picture. Based on its training knowledge and the input fusion features, the model generates a response text, such as: According to the picture information, this plant is likely to be a tulip. The response text contains a response to the user's intent (such as "According to the picture information"), and specific information recognized based on the multi-modal fusion result (such as "tulip"). The response text aims to provide the information required by the user in natural language.
[0058] It is worth mentioning that all the weight matrices and parameters used in each component such as the natural language unimodal depth encoder (including the BERT model), the image depth encoder (including the ViT model), the early feature fusion component (based on the cross-attention module), the semantic layer feature fusion component (based on the multi-head self-attention module), the intent recognition classifier (feed-forward neural network or multi-layer perceptron), and the large language model in this application are not preset fixed values. These parameters are automatically obtained by the model through learning a large number of labeled data sets during the training process. The training process usually involves performing forward calculations on the training data to obtain output results, and then calculating the difference between the model output and the true target (such as the correct intent label, the expected response text, or the matching features), and this difference is measured by the loss function. Then, using an optimization algorithm (such as gradient descent and its variants), according to the gradient information of the loss function, the weight matrices and parameters inside the model are iteratively updated. By repeatedly performing forward calculations and parameter updates, the model gradually adjusts its internal parameters to minimize the loss function, thereby learning how to effectively encode, fuse, classify, and generate information.
[0059] In summary, the AI multi-modal dialogue system 100 based on multi-modal recognition according to the embodiments of the present application is elucidated. It first extracts text word-level granularity and image local features respectively, laying a fine-grained foundation for cross-modal interaction. Secondly, through the bidirectional cross-attention mechanism, a dynamic association between text and image is constructed at the feature level. Among them, the attention sparsification process filters out key cross-modal interaction nodes, and the residual enhancement module alleviates the feature mismatch problem caused by modal differences through a global-local symmetric optimization strategy, strengthening the complementarity of low-level features. Subsequently, the dynamic association of cross-modal high-level semantics is captured through multi-head attention, and its hierarchical processing mechanism enables the system to adaptively weigh the value of multi-modal information at different task stages. This progressive fusion architecture from the feature layer to the semantic layer, combined with sparse coupling and symmetric optimization techniques, not only solves the modal gap problem caused by simple splicing in traditional methods, but also captures cross-modal spatio-temporal associations through the dynamic attention mechanism, and finally achieves accurate responses through intention recognition and large model generation. In this way, the limitations of the prior art in cross-modal feature extraction and dynamic interaction modeling are broken through, and the accuracy of multi-modal intention understanding and response relevance are significantly improved.
[0060] As described above, the AI multi-modal dialogue system 100 based on multi-modal recognition according to the embodiments of the present application can be implemented in various wireless terminals, such as a server with an AI multi-modal dialogue algorithm based on multi-modal recognition. In a possible implementation, the AI multi-modal dialogue system 100 based on the embodiments of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the AI multi-modal dialogue system 100 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the AI multi-modal dialogue system 100 can also be one of the many hardware modules of the wireless terminal.
[0061] Alternatively, in another example, the AI multi-modal dialogue system 100 based on multi-modal recognition and the wireless terminal can also be separate devices, and the AI multi-modal dialogue system 100 based on multi-modal recognition can be connected to the wireless terminal through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0062] The above have described the various implementations of the present disclosure. The above description is exemplary and not exhaustive. And it is not limited to the disclosed implementations. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described implementations.
Claims
1. An AI multimodal dialogue system based on multimodal recognition, characterized in that, Including: An input acquisition module, configured to acquire text information input by a user and image data uploaded by the user; A multi-modal encoding module, configured to respectively input the text information and the image data into a natural language single-modal depth encoder and an image depth encoder to obtain a set of user-intention word-grained semantic encoding vectors and a set of image local semantic encoding feature vectors; An early feature fusion module, configured to input the set of user-intention word-grained semantic encoding vectors and the set of image local semantic encoding feature vectors into an early feature fusion component based on a cross-attention network to obtain a set of enhanced user-intention word-grained semantic encoding vectors and a set of enhanced image local semantic encoding feature vectors; A semantic layer feature fusion module, configured to input the set of enhanced user-intention word-grained semantic encoding vectors and the set of enhanced image local semantic encoding feature vectors into a semantic layer feature fusion component based on a multi-head attention module to obtain a user-intention multi-modal fusion representation; An intention recognition module, configured to input the user-intention multi-modal fusion representation into an intention recognition classifier to obtain an intention recognition result; An intelligent response generation module, configured to input the intention recognition result and the user-intention multi-modal fusion representation into a large language model to obtain a response text.
2. The AI multimodal dialogue system based on multimodal recognition according to claim 1, characterized in that The multi-modal encoding module includes: a text tokenization processing unit, configured to perform tokenization processing on the text information to obtain a text word sequence; a text semantic encoding unit, configured to input the text word sequence into the natural language single-modal depth encoder to obtain the set of user-intention word-grained semantic encoding vectors, wherein the natural language single-modal depth encoder is a semantic encoder including a Bert model; an image chunking processing unit, configured to perform image chunking processing on the image data to obtain an image chunk sequence; an image semantic encoding unit, configured to input the image chunk sequence into the image depth encoder to obtain the set of image local semantic encoding feature vectors, wherein the image depth encoder is a Vit model including an image chunk embedding encoding layer.
3. The AI multimodal dialogue system based on multimodal recognition according to claim 1, wherein, The early feature fusion module includes: a user intention semantic vector extraction unit for extracting a first user intention word-level semantic encoding vector from the set of user intention word-level semantic encoding vectors; an image-to-early cross-modal attention calculation unit for calculating early feature cross-modal attention values of each image local semantic encoding feature vector in the set of image local semantic encoding feature vectors with respect to the first user intention word-level semantic encoding vector to obtain a set of image-to-early feature cross-modal attention values; an image cross-modal feature fusion unit for fusing the set of image local semantic encoding feature vectors based on the set of image-to-early feature cross-modal attention values to obtain an image-to-cross-modal early feature attention interaction encoding vector; and a user intention residual enhancement unit for inputting the image-to-cross-modal early feature attention interaction encoding vector and the first user intention word-level semantic encoding vector into a residual module to obtain an enhanced user intention word-level semantic encoding vector corresponding to the first user intention word-level semantic encoding vector.
4. The AI multi-modal dialogue system based on multi-modal recognition according to claim 3, characterized in that, The image cross-modal feature fusion unit includes: an attention weight value calculation sub-unit for inputting the set of image-to-early feature cross-modal attention values into a Softmax activation function to obtain a set of image-to-early feature cross-modal attention weight values; an attention weight value sparsification sub-unit for inputting the set of image-to-early feature cross-modal attention weight values into an attention sparsification module based on a preset threshold to obtain a sparse set of image-to-early feature cross-modal attention weight values; and an image cross-modal weighted sub-unit for calculating a weighted sum of the set of image local semantic encoding feature vectors using the sparse set of image-to-early feature cross-modal attention weight values as weights to obtain the image-to-cross-modal early feature attention interaction encoding vector.
5. The AI multimodal dialogue system based on multimodal recognition according to claim 4, characterized in that, The attention weight value sparsification subunit is configured to: perform attention sparsification processing on the set of cross-modal attention weight values of the image towards early features by using the following formula to obtain the sparse set of cross-modal attention weight values of the image towards early features, where the formula is: ; where is each cross-modal attention weight value of the image towards early features in the set of cross-modal attention weight values of the image towards early features, is a preset threshold, is the attention sparsification processing, is each cross-modal attention weight value of the image towards early features in the sparse set of cross-modal attention weight values of the image towards early features.
6. The AI multi-modal dialogue system based on multi-modal recognition according to claim 5, characterized in that, The user intention residual enhancement unit includes: a sparse coupling symmetric optimization sub-unit for performing global-local sparse coupling symmetric optimization on the image-to-cross-modal early feature attention interaction encoding vector and the first user intention word-level semantic encoding vector to obtain an image-to-cross-modal early feature attention interaction optimized encoding vector and a first user intention word-level semantic optimized encoding vector; and a first user intention feature enhancement sub-unit for inputting the image-to-cross-modal early feature attention interaction optimized encoding vector and the first user intention word-level semantic optimized encoding vector into a residual module to obtain an enhanced user intention word-level semantic encoding vector corresponding to the first user intention word-level semantic encoding vector.
7. The AI multi-modal dialogue system based on multi-modal recognition according to claim 6, characterized in that, The sparse coupling symmetric optimization subunit is used for: performing global sparsity residual modeling on the image towards the cross-modal early feature attention interaction coding vector and the first user intention word granularity semantic coding vector to obtain the image towards the cross-modal early feature attention interaction coding global vector and the first user intention word granularity semantic coding global vector; using the image towards the cross-modal early feature attention interaction coding vector and the first user intention word granularity semantic coding vector as local space coordinate bases to calculate the smoothness mutual constraint factor of the global sparsity residual to obtain the image towards the cross-modal constraint factor and the first user intention constraint factor; Based on the image towards the cross-modal constraint factor and the first user intention constraint factor, performing symmetric transformation optimization on the image towards the cross-modal early feature attention interaction coding vector and the first user intention word granularity semantic coding vector to obtain the image towards the cross-modal early feature attention interaction optimized coding vector and the first user intention word granularity semantic optimized coding vector.
8. The AI multimodal dialogue system based on multimodal recognition according to claim 1, characterized in that, The early feature fusion module further includes: an image semantic vector extraction unit for extracting a first image local semantic coding feature vector from the set of the image local semantic coding feature vectors; a text to early cross-modal attention calculation unit for calculating the early feature cross-modal attention values of each user intention word granularity semantic coding vector in the set of the user intention word granularity semantic coding vectors relative to the first image local semantic coding feature vector to obtain a set of text to early feature cross-modal attention values; a text cross-modal feature fusion unit for fusing the set of the user intention word granularity semantic coding vectors based on the set of the text to early feature cross-modal attention values to obtain a text to cross-modal early feature attention interaction coding vector; an image semantic residual enhancement unit for inputting the text to cross-modal early feature attention interaction coding vector and the first image local semantic coding feature vector into a residual module to obtain an enhanced image local semantic coding feature vector corresponding to the first image local semantic coding feature vector.
Citation Information
Patent Citations
Multi-modal emotion recognition method based on attention enhancing mechanism
CN112489635A
Intestinal environment deep completion method based on multi-scale confidence coefficient and self-attention mechanism
CN116523986A
Semantic segmentation model and segmentation method for high-resolution remote sensing image
CN119206229A
AI dialogue system based on multi-modal input
CN119831058A
Multi-modal dialogue abstract method based on multi-level visual guidance
CN119918545A
Cited By
Semantic understanding and representation learning method for smart home dialogue system
CN120851042A
Multi-modal feature fused AI interactive voice intention recognition method and system
CN121963706A
Public opinion monitoring method and system based on multi-modal data fusion
CN121981714A
Interaction method based on multi-modal large model, storage medium and electronic device
CN122020573A