Multi-modal deep semantic intention recognition method, system and equipment and storage medium

Through multi-layer text and image semantic information extraction and cross-modal attention mechanism, the problem of insufficient use of semantic information in the pre-trained model is solved, and higher accuracy and robustness of intention recognition are achieved.

CN120277614APending Publication Date: 2025-07-08INSPUR SMART TECH (NANJING) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510548885.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing pretrained models fail to make full use of semantic information in the identification of graphic and text intentions, resulting in insufficient progressive relationship between semantic information in the cross-modal fusion module, affecting the performance and accuracy of the model.

Method used

Multi-layer text semantic information and image semantic information extraction methods are used to fusion of features using a cross-modal attention mechanism, and intent categories are identified through classifiers, and hierarchical feature extraction and dynamic weighted fusion are combined with RoBERTa, graph convolutional neural network and Transformer encoder.

Benefits of technology

It improves the accuracy of intention recognition, solves the semantic deviation problem caused by traditional models due to ignoring the language hierarchy structure, and enhances the robustness of cross-modal fusion and the accuracy of intention recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277614A_ABST
    Figure CN120277614A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intention recognition, and particularly provides a multi-modal deep semantic intention recognition method, system and device and a storage medium, and the method comprises the steps: obtaining text data, and extracting multi-layer text semantic information from the text data; acquiring image data, and extracting multilayer image semantic information from the image data; fusing the multi-layer text semantic information and the multi-layer image semantic information into a feature vector by using a cross-modal attention mechanism; and identifying the intention category corresponding to the feature vector by using a classifier. According to the method, by paying attention to hierarchical features, the problem of semantic deviation caused by neglecting a language hierarchical structure in a traditional model is solved, and the intention recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intention recognition, and particularly relates to a multi-modal deep semantic intention recognition method, system, device and storage medium. Background Art

[0002] Intention recognition has a wide range of applications in multiple fields, mainly including intelligent customer service, smart home, autonomous driving, medical diagnosis and sentiment analysis, etc. In a text and image intention recognition model, text information and image information need to be converted into vector form and input into a neural network, and the semantic features of the text and image are output.

[0003] However, since the pre-trained model is a black box to the users, the users only need to focus on the results of the model analysis without paying attention to the process of obtaining semantic features. This leads to the fact that the semantic information parsed by the model cannot be fully utilized, and the progressive relationship of semantic information from shallow to deep is not reflected in the subsequent cross-modal fusion module, affecting the performance and accuracy of the overall model. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the present invention provides a multi-modal deep semantic intention recognition method, system, device and storage medium to solve the above technical problems.

[0005] In a first aspect, the present invention provides a multi-modal deep semantic intention recognition method, including: Obtaining text data, and extracting multi-layer text semantic information from the text data; Obtaining image data, and extracting multi-layer image semantic information from the image data; Using a cross-modal attention mechanism to fuse the multi-layer text semantic information and multi-layer image semantic information into a feature vector; Using a classifier to identify the intention category corresponding to the feature vector.

[0006] In an optional embodiment, obtaining text data, and extracting multi-layer text semantic information from the text data, includes: Using a RoBERTa model to extract a word sequence from the text data; Performing dependency syntactic analysis and graph convolutional neural network processing on the word sequence to obtain word-level features; Inputting the word-level features into a bidirectional gated recurrent unit to obtain phrase-level features; Generating a sentence feature vector according to the word-level features and phrase-level features.

[0007] In an optional embodiment, performing dependency syntactic analysis and graph convolutional neural network processing on the word sequence to obtain word-level features, includes: Construct a syntactic dependency tree among words in the word sequence through dependency parsing; Use a graph convolutional network to perform semantic propagation on the syntactic dependency tree to capture the syntactic constraint relationships among words; Generate word-level features according to the syntactic constraint relationships and the word-level features.

[0008] In an optional implementation, input the phrase-level features into a bidirectional gated recurrent unit to obtain sentence-level features, including: Use the bidirectional gated recurrent unit to convert a phrase into a phrase feature vector containing positional encoding based on the temporal order of words in the phrase, and associate relative position information of the phrase in the sentence with the phrase feature vector; Generate a phrase-level semantic representation according to the phrase feature vector and the relative position information.

[0009] In an optional implementation, generate a sentence feature vector according to the word-level features and the phrase-level features, including: Construct a multi-head self-attention mechanism network, establish cross-phrase semantic associations globally according to the phrase-level semantic representation, and encode the cross-phrase semantic associations and the phrase-level semantic representation into phrase features; Adopt a residual connection method to dynamically weight and fuse the word-level features and the phrase features to obtain a sentence feature vector with hierarchical representation.

[0010] In an optional implementation, obtain image data and extract multi-layer image semantic information from the image data, including: Segment the image data into multiple image patches; Flatten the image patches along the RGB channels into image feature vectors of 16×16×3; Use the multi-head attention mechanism to calculate the image feature vectors of all image patches to obtain a set of image vectors; Inject positional encoding into the set of image vectors; Use a stack of Transformer encoders to process the set of image vectors after injecting positional encoding to obtain multi-layer image semantic information; The stack of Transformer encoders has a 12-layer network structure, which is divided into three groups with four layers in each group, and the three groups of network structures are arranged as the first group of networks, the second group of networks, and the third group of networks in the order from the upstream to the downstream of data processing. Output the semantic information extracted by the first group of networks as local block-to-block association information, output the semantic information extracted by the second group of networks as object part-level association information, and use the semantic information output by the third group of networks as global semantic association information.

[0011] In an alternative embodiment, a cross-modal attention mechanism is used to fuse the multi-layer text semantic information and multi-layer image semantic information into a feature vector, including: Concatenate the multi-layer text semantic information and multi-layer image semantic information along the channel dimension to obtain a concatenated feature; Calculate the similarity score matrix of the concatenated feature; Calculate the attention weight matrix of the similarity score matrix; Generate a weighted visual context according to the attention weight matrix, image features, and visual value transformation matrix; Use a gating mechanism to adjust the weighted visual context and text features; Perform multi-level fusion on the adjusted weighted visual context and text features to obtain a feature vector.

[0012] In a second aspect, the present invention provides a multi-modal deep semantic intent recognition system, including: A first feature extraction module, configured to obtain text data and extract multi-layer text semantic information from the text data; A second feature extraction module, configured to obtain image data and extract multi-layer image semantic information from the image data; A feature fusion module, configured to use a cross-modal attention mechanism to fuse the multi-layer text semantic information and multi-layer image semantic information into a feature vector; An intent classification module, configured to use a classifier to identify the intent category corresponding to the feature vector.

[0013] In a third aspect, a device is provided, including: A memory, configured to store a multi-modal deep semantic intent recognition program; A processor, configured to implement the steps of the multi-modal deep semantic intent recognition method provided in the first aspect when executing the multi-modal deep semantic intent recognition program.

[0014] In a fourth aspect, a computer-readable storage medium is provided, on which a multi-modal deep semantic intent recognition program is stored. When the multi-modal deep semantic intent recognition program is executed by a processor, the steps of the multi-modal deep semantic intent recognition method provided in the first aspect are implemented.

[0015] The beneficial effects of the present invention are as follows. The multi-modal deep semantic intention recognition method, system, device and storage medium provided by the present invention extract hierarchical features from semantic information, and combine the attention mechanism and adaptive gating to construct a cross-modal attention mechanism module for modal weight calculation and feature fusion. Finally, the fused cross-modal features are sent into a classifier composed of multiple linear layers to obtain the intention recognition result. By focusing on hierarchical features, the present invention solves the semantic deviation problem caused by traditional models ignoring the language hierarchical structure and improves the accuracy of intention recognition.

[0016] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention.

[0019] Figure 2 is a schematic principle diagram of the method according to an embodiment of the present invention.

[0020] Figure 3 is a schematic block diagram of the system according to an embodiment of the present invention.

[0021] Figure 4 is a schematic structural diagram of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0024] The multi-modal deep semantic intention recognition method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the multi-modal deep semantic intention recognition system runs in the computer device.

[0025] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a multi-modal deep semantic intention recognition system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0026] As Figure 1 shown, the method includes: S1. Obtain text data, and extract multi-layer text semantic information from the text data; S2. Obtain image data, and extract multi-layer image semantic information from the image data; S3. Use a cross-modal attention mechanism to fuse the multi-layer text semantic information and the multi-layer image semantic information into a feature vector; S4. Use a classifier to identify the intention category corresponding to the feature vector.

[0027] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.

[0028] S101. Use the RoBERTa model to extract a word sequence from the text data.

[0029] RoBERTa (Robustly Optimized BERT Pretraining Approach) is a pre-trained language model improved based on BERT. The following are the detailed steps: Text preprocessing: Tokenize the original text, remove special characters, stop words, etc., and convert the text to lowercase.

[0030] Input encoding: Convert the preprocessed text into an input format acceptable to the RoBERTa model, that is, a sequence of tokens. Usually, the tokenizer of RoBERTa is used to tokenize the text into tokens, and special start and end markers (such as [CLS] and [SEP]) are added.

[0031] Feature extraction: Input the token sequence after input encoding into the pre-trained RoBERTa model to obtain the hidden state representation of each token.

[0032] Assume the input text is T = [w1, w2, …, w n , after being processed by the tokenizer, the token sequence Ttokens =[t1, t2, …, t m , where m may be greater than n because a word may be split into multiple tokens. Input T tokens into the RoBERTa model, and the output of the model is H = [h1, h2, …, h m , where h i is the hidden state representation of the i-th token.

[0033] S102. Perform dependency syntactic analysis and graph convolutional neural network processing on the word sequence to obtain word-level features.

[0034] (1) Construct a syntactic dependency tree among the words in the word sequence through dependency syntactic analysis; Use a dependency syntactic analyzer (such as SpaCy, Stanford CoreNLP, etc.) to analyze the word sequence to obtain the syntactic dependency relationships among the words. The dependency relationships can be represented as a directed graph, where the nodes represent words and the edges represent the dependency relationships between words. Construct a syntactic dependency tree based on the dependency relationships, and the root node of the tree is usually the core verb of the sentence.

[0035] (2) Use the graph convolutional network to perform semantic propagation on the syntactic dependency tree to capture the syntactic constraint relationships among the words.

[0036] Perform graph convolutional operations on the syntactic dependency tree to capture the syntactic constraint relationships among the words. The basic formula of the graph convolutional network (GCN) is as follows:

[0037] where H (l) is the node feature matrix of the l-th layer, is the adjacency matrix A plus the self-loop (I is the identity matrix), is 's degree matrix, W(l) is the learnable weight matrix of the l-th layer, is the activation function (such as ReLU).

[0038] (3) Generate word-level features based on the syntactic constraint relationships and the word-level features.

[0039] Generate word-level features based on the syntactic constraint relationships and the word-level features. Assume the word-level feature is H word , and the feature after being processed by the graph convolutional network is H gcn , then the word-level feature H word_level can be calculated by the following formula: .

[0040] S103. Input the word-level features into a bidirectional gated recurrent unit to obtain phrase-level features.

[0041] Use the bidirectional gated recurrent unit to convert a phrase into a phrase feature vector containing positional encoding based on the temporal order of words in the phrase, and associate the relative position information of the phrase in the sentence with the phrase feature vector; generate a phrase-level semantic representation according to the phrase feature vector and the relative position information.

[0042] Specifically, use the bidirectional gated recurrent unit to convert a phrase into a phrase feature vector containing positional encoding based on the temporal order of words in the phrase, and associate the relative position information of the phrase in the sentence with the phrase feature vector. Assume the output of the bidirectional GRU is H bigru , the positional encoding is P, then the phrase feature vector V phrase can be calculated by the following formula:

[0043] Generate a phrase-level semantic representation S phrase :

[0044] where MLP is a multi-layer perceptron.

[0045] S104. Generate a sentence feature vector according to the word-level features and the phrase-level features.

[0046] (1) Construct a multi-head self-attention mechanism network to establish cross-phrase semantic associations globally according to the phrase-level semantic representation, and encode the cross-phrase semantic associations and the phrase-level semantic representation into phrase features.

[0047] Construct a multi-head self-attention mechanism network to establish cross-phrase semantic associations globally according to the phrase-level semantic representation. The calculation formula of multi-head self-attention is as follows:

[0048] where , , Q, K, and V are the query, key, and value matrices respectively, , , , are learnable weight matrices, h is the number of heads, and d k is the dimension of the key.

[0049] Encode the cross-phrase semantic associations and the phrase-level semantic representation into phrase features F phrase .

[0050] (2) The residual connection method is adopted to dynamically and weightedly fuse the word-level features and phrase-level features to obtain a sentence feature vector with hierarchical representation.

[0051] The residual connection method is adopted to dynamically and weightedly fuse the word-level features and phrase-level features to obtain a sentence feature vector V with hierarchical representation. sentence . The calculation formula is as follows:

[0052] Among them, α is the dynamic weight, which can be determined by a learnable parameter.

[0053] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0054] The image data is segmented into multiple image blocks; the image blocks are flattened into image feature vectors of 16×16×3 along the RGB channels; the multi-head attention mechanism is used to calculate the image feature vectors of all image blocks to obtain an image vector set; position encoding is injected into the image vector set.

[0055] The image vector set after injecting position encoding is processed by stacking Transformer encoders to obtain multi-layer image semantic information; the Transformer encoder stacks a 12-layer network structure, and the 12-layer network structure is divided into three groups with four layers in each group, and the three groups of network structures are arranged in the order from the upstream to the downstream of data processing as the first group of networks, the second group of networks, and the third group of networks. The semantic information extracted by the first group of networks is output as local block-to-block association information, the semantic information extracted by the second group of networks is output as object part-level association information, and the semantic information output by the third group of networks is used as global semantic association information.

[0056] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0057] S301. The multi-layer text semantic information and the multi-layer image semantic information are concatenated along the channel dimension to obtain a concatenated feature.

[0058]

[0059] Among them, T is the multi-layer text semantic information, that is, text features; V is the multi-layer image semantic information, that is, image features; is the set of real numbers, d t is the dimension of T, d v is the dimension of V.

[0060] Map the features of different modalities to the same space, providing a basis for subsequent cross-modal interactions.

[0061] S302. Calculate the similarity score matrix of the concatenated features.

[0062] Similarity score matrix:

[0063] Query: The text feature T is projected into the query space through the matrix W Q to represent the "question" of the text about the visual information.

[0064] Key: The image feature V is projected into the key space through the matrix W K to represent the "index identifier" of the image region.

[0065] Similarity calculation: Measure the association strength between text words and image regions through dot product, and the scaling factor d is used to prevent gradient explosion.

[0066] n t represents the length of the text sequence; n v represents the number of spatial regions of the visual feature (e.g., the number of blocks or regions into which the image is divided). If an image region is highly correlated with a text word (e.g., "dog" corresponding to the dog in the image), the score will be higher.

[0067] S303. Calculate the attention weight matrix of the similarity score matrix.

[0068]

[0069] Softmax: Normalize the similarity matrix row-wise into a probability distribution so that the sum of the attention weights of each text word to the image region is 1.

[0070] Highlight important associations (e.g., "running" focusing on the dynamic region) and suppress irrelevant backgrounds (e.g., sky or grass).

[0071] S304. Generate the weighted visual context according to the attention weight matrix, the image feature, and the visual value transformation matrix.

[0072]

[0073] Value: The image feature V is projected into the value space through the matrix W V to represent the semantic content of the image region.

[0074] Weighted summation: Aggregate the image region features weighted according to the attention weight A to generate a visual context related to the text.

[0075] The text word "red" will weight the red object region in the image to generate color-related visual features.

[0076] S305. Use a gating mechanism to adjust the weighted visual context and text features.

[0077]

[0078] Concatenate the input: Concatenate the text feature T and the visual context V ctx to capture the joint information after their interaction.

[0079] Sigmoid activation: Map the result of the linear transformation to the interval [0, 1] to represent the initial weight assignment of text and vision.

[0080] For example, gbase = 0.7 means the model defaults to relying more on text information.

[0081] The text confidence measures the intensity of the text feature through the L2 norm. The larger the value, the clearer the text information (such as rich keywords). The calculation formula is:

[0082] The visual confidence measures the overall response intensity of the visual context through the mean value. The larger the value, the more significant the visual information (such as clear objects). The calculation formula is: .

[0083] The dynamic gating adjustment formula is:

[0084] Modify the base gating value according to the modality confidence.

[0085] If the text confidence is high (c t ≫ c v ), then gdynamic → 1 and the model relies on text.

[0086] If the visual confidence is high (c v ≫ c t ), then gdynamic → 0 and the model relies on vision.

[0087] When the image is blurred, c v decreases and the model automatically turns to text dominance.

[0088] S306. Perform multi-level fusion on the adjusted weighted visual context and text features to obtain a feature vector.

[0089] The single-layer feature fusion calculation formula:

[0090] The text features and visual context are linearly combined through dynamic gating weights.

[0091] When gdynamic = 0.8, 80% of the fused features come from the text and 20% come from vision. When the confidence levels of both are close, the model balances the information of the two modalities.

[0092] Multi-level fusion:

[0093] For multi-level features (such as word, phrase, sentence levels), the fusion results of each layer are calculated separately.

[0094] Through the learnable weight λ k Integrate semantics at different levels (such as word-level details + sentence-level global information). Layer normalization stabilizes the training process and prevents gradient explosion.

[0095] Please refer to Figure 2 , and the following is a specific embodiment: 1. Text feature extraction.

[0096] The text is divided into three levels: words, phrases, and sentences, and encoded in these three levels in sequence. For the given text content, it is decomposed into a word sequence based on the RoBERTa pre-trained model . Among them represents the number of English words or Chinese characters in the text. In the present invention, the maximum value of is set to 100, and the excess part will be truncated, while the insufficient part will use the symbol <pad>Instead. The feature representation obtained after RoBERTa encoding is , where corresponds to the feature obtained after encoding.

[0097] Subsequently, a syntactic dependency graph is constructed through dependency parsing, and a graph convolutional network (GCN) is used to perform semantic propagation on the dependency tree to capture the syntactic constraint relationships between words, forming a word-level semantic representation . Among them, represents the set of dependency relationships indicating syntactic relationships such as subject-predicate and verb-object, represents the adjacency matrix. Calculation formula: , .

[0098] After obtaining the phrase-level semantic representation, a bidirectional gated recurrent unit (BIGRU) is used to perform temporal modeling on the words within the phrase. Each phrase unit outputs a feature vector containing positional encoding, retaining the relative position information of the phrase in the sentence. A phrase-level semantic representation is formed . Subsequently, a multi-head self-attention mechanism network is constructed to establish cross-phrase semantic associations globally. The word-level features and phrase-level features are weighted and fused in a residual connection manner to generate a sentence feature vector with hierarchical representation . Among them, is a learnable weight matrix. Calculation formula: , .

[0099] 2. Image feature extraction.

[0100] For the given picture content, first perform preprocessing to convert the image into a standard size of 3×224×22, then segment it into 16×16 image patches, and then flatten the image patches along the RGB channels into a feature vector of 16×16×3. After performing multi-head self-attention mechanism on all the image patches of the entire image, a vector set is obtained. After obtaining the feature vectors of the text and the picture respectively, a cross-modal attention mechanism module is used to calculate the dynamic weights and generate the visual context. And a dynamic interaction gated unit is used to dynamically adjust the cross-modal information flow and dynamically adjust the fusion ratio of the text and visual features.

[0101] 3. Feature fusion.

[0102] Specifically, first concatenate the feature vectors of the text and the picture, then send them into an attention module, calculate the similarity score and normalize it through softmax to obtain the cross-modal attention weights , and through and the visual feature matrix After weighted operation, the visually weighted context is obtained . Subsequently, the basic gating weight is generated through the Sigmoid activation function, and the basic gating value is dynamically adjusted according to the modality confidence evaluated in real time to obtain the dynamic fusion gating weight , where the text confidence is The L2 norm of, and the visual confidence is The mean value Indicates that the text features and visual context are concatenated, which is used to fuse the confidence information of each modality during the calculation process. The closer it is to 1, the higher the proportion of text weight. Similarly, The closer it is to 0, the higher the proportion of picture weight, which is weighted and fused with text features and visual features to obtain the fused features , and finally the normalized output obtains the final fusion result , 4. Intent classification.

[0103] The softmax calculation is performed on the fused results of each layer after splicing using the intent classifier. Among them, Indicates the text features, , Respectively represent the Q, K, and V projection matrices, Is the scaling factor, Is the dynamic gating parameter, Represents the Sigmoid function.

[0104] , .

[0105] This embodiment has the following beneficial effects: 1. Hierarchical semantic extraction improves intent recognition accuracy: Through RoBERTa, GCN, BiGRU, and multi-head attention, explicitly model syntactic constraints and context associations to solve the semantic deviation problem caused by traditional models ignoring the language hierarchical structure.

[0106] 2. Cross-modal dynamic fusion enhances robustness: Use confidence-aware gating to dynamically adjust the fusion weights, and hierarchical attention alignment realizes fine-grained intent understanding.

[0107] 3. Multi-scenario generalization ability: Only need to fine-tune the classifier layer to support intent recognition in different fields such as smart home, in-vehicle instructions, and medical consultations.

[0108] In some embodiments, the multimodal deep semantic intent recognition system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the multimodal deep semantic intent recognition system may be stored in the memory of a computer device and executed by at least one processor to perform (see Figure 1 description) the functions of multimodal deep semantic intent recognition.

[0109] In this embodiment, the multimodal deep semantic intent recognition system can be divided into multiple functional modules according to the functions it performs, as Figure 3 shown. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0110] The first feature extraction module is used to obtain text data and extract multi-layer text semantic information from the text data; The second feature extraction module is used to obtain image data and extract multi-layer image semantic information from the image data; The feature fusion module is used to fuse the multi-layer text semantic information and multi-layer image semantic information into a feature vector by using a cross-modal attention mechanism; The intent classification module is used to identify the intent category corresponding to the feature vector by using a classifier.

[0111] Figure 4 The multimodal deep semantic intent recognition method provided by the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.

[0112] Among them, the device 400 may include: a processor 410, a memory 420, and a communication unit 430. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0113] Among them, the memory 420 can be used to store the execution instructions of the processor 410. The memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 420 are executed by the processor 410, the device 400 can execute some or all of the steps in the above method embodiments.

[0114] The processor 410 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 420, and calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or composed of multiple packaged ICs with the same or different functions connected. For example, the processor 410 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single operation core or include multiple operation cores.

[0115] The communication unit 430 is used to establish a communication channel so that the storage device can communicate with other devices. Receive user data sent by other devices or send user data to other devices.

[0116] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0117] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., various media that can store program codes, including several instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0118] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the descriptions in the method embodiments.

[0119] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the systems or modules can be in electrical, mechanical, or other forms.

[0120] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0121] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0122] Although the present invention has been described in detail by referring to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all fall within the scope of the present invention / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.< / pad>

Claims

1. A multi-modal deep semantic intention recognition method, characterized in that Including: Obtain text data, and extract multi-layer text semantic information from the text data; Obtain image data, and extract multi-layer image semantic information from the image data; Use a cross-modal attention mechanism to fuse the multi-layer text semantic information and multi-layer image semantic information into a feature vector; Use a classifier to identify the intent category corresponding to the feature vector.

2. The method according to claim 1, characterized in that Obtain text data, and extract multi-layer text semantic information from the text data, including: Use the RoBERTa model to extract a word sequence from the text data; Perform dependency syntactic analysis and graph convolutional neural network processing on the word sequence to obtain word-level features; Input the word-level features into a bidirectional gated recurrent unit to obtain phrase-level features; Generate a sentence feature vector according to the word-level features and phrase-level features.

3. The method according to claim 2, wherein Perform dependency syntactic analysis and graph convolutional neural network processing on the word sequence to obtain word-level features, including: Construct a syntactic dependency tree between words in the word sequence through dependency syntactic analysis; Use a graph convolutional network to perform semantic propagation on the syntactic dependency tree to capture the syntactic constraint relationships between words; Generate word-level features according to the syntactic constraint relationships and the word-level features.

4. The method according to claim 2, wherein Input the word-level features into a bidirectional gated recurrent unit to obtain phrase-level features, including: Use a bidirectional gated recurrent unit to convert a phrase into a phrase feature vector containing position encoding based on the time sequence of words in the phrase, and associate relative position information of the phrase in the sentence with the phrase feature vector; Generate a phrase-level semantic representation according to the phrase feature vector and the relative position information.

5. The method according to claim 2, wherein Generate a sentence feature vector according to the word-level features and phrase-level features, including: Construct a multi-head self-attention mechanism network, establish cross-phrase semantic associations globally according to the phrase-level semantic representation, and encode the cross-phrase semantic associations and phrase-level semantic representation into phrase features; Adopt a residual connection method to dynamically weight and fuse the word-level features and phrase features to obtain a sentence feature vector with hierarchical representation.

6. The method according to claim 1, characterized in that, Obtain image data, and extract multi-layer image semantic information from the image data, including: Segment the image data into multiple image patches; Flatten the image patches along the RGB channels into an image feature vector of 16×16×3; Use a multi-head attention mechanism to calculate the image feature vectors of all image patches to obtain a set of image vectors; Inject position encoding into the set of image vectors; Use a stack of Transformer encoders to process the set of image vectors with injected position encoding to obtain multi-layer image semantic information; The stack of Transformer encoders has a 12-layer network structure. The 12-layer network structure is divided into three groups with four layers in each group. The three groups of network structures are arranged as the first group of networks, the second group of networks, and the third group of networks in the order from the upstream to the downstream of data processing. The semantic information extracted by the first group of networks is output as local inter-block association information, the semantic information extracted by the second group of networks is output as object part-level association information, and the semantic information output by the third group of networks is used as global semantic association information.

7. The method according to claim 1, wherein Fusing the multi-layer text semantic information and multi-layer image semantic information into a feature vector by using a cross-modal attention mechanism, including: Concatenating the multi-layer text semantic information and multi-layer image semantic information along the channel dimension to obtain a concatenated feature; Calculating a similarity score matrix of the concatenated feature; Calculating an attention weight matrix of the similarity score matrix; Generating a weighted visual context according to the attention weight matrix, image features and a visual value transformation matrix; Adjusting the weighted visual context and text features by using a gating mechanism; Performing multi-level fusion on the adjusted weighted visual context and text features to obtain a feature vector.

8. A multi-modal deep semantic intention recognition system, characterized in that, Including: A first feature extraction module for obtaining text data and extracting multi-layer text semantic information from the text data; A second feature extraction module for obtaining image data and extracting multi-layer image semantic information from the image data; A feature fusion module for fusing the multi-layer text semantic information and multi-layer image semantic information into a feature vector by using a cross-modal attention mechanism; An intent classification module for identifying the intent category corresponding to the feature vector by using a classifier.

9. A device, characterized in that, Including: A memory for storing a multi-modal deep semantic intent recognition program; A processor for implementing the steps of the multi-modal deep semantic intent recognition method as described in any one of claims 1-7 when executing the multi-modal deep semantic intent recognition program.

10. A computer-readable storage medium storing a computer program, characterized in that, The multi-modal deep semantic intent recognition program is stored on the readable storage medium, and when the multi-modal deep semantic intent recognition program is executed by the processor, the steps of the multi-modal deep semantic intent recognition method as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Crop disease and pest segmentation detection method based on TVFNet network model

    CN120543554A

  • Intelligent agent training data set construction method and system in combination with cross-modal learning

    CN121212191A

  • Agent training data set construction method and system combined with cross-modal learning

    CN121212191B

  • Text generation method and system based on AI

    CN121388196A

  • Cabin scene semantic graph construction method and device and program product

    CN121480661A