Multimodal information processing method, device, electronic device and storage medium

By extracting the features in multimodal information and using attention mechanism and multimodal decomposition bilinear pooling processing, the problem of waste of information by singlemodal method is solved, and more accurate intention classification results are achieved.

CN113762319BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110239408.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-04
Publication Date
2025-05-23
Estimated Expiration
2041-03-04

AI Technical Summary

Technical Problem

When the prior art adopts a single-modal method in multimodal information processing, information from other modalities is wasted, resulting in low accuracy of intention classification results.

Method used

By extracting the first and second modal features in the multimodal information, the first modal features are weighted by using the attention mechanism, and combined with the multimodal decomposition bilinear pooling process, more accurate intention classification results are generated.

Benefits of technology

The full interactive fusion of the first modal feature and the second modal feature is achieved, and the accuracy of the intention classification results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113762319B_ABST
    Figure CN113762319B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a multimodal information processing method, device, electronic device and storage medium, the method comprising: extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities; based on an attention mechanism, using the second modal feature, performing attention weighted processing on the first modal feature to obtain a first weighted modal feature; performing multimodal decomposition bilinear pooling processing on the first weighted modal feature and the second modal feature to obtain a bimodal vector; and generating an intent classification result corresponding to the multimodal information according to the bimodal vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and relates to but is not limited to a multimodal information processing method, device, electronic device and storage medium. Background Art

[0002] In the related technologies, the solutions for multimodal information processing mainly include: processing methods based on classification or matching models of unimodal languages, such as TextCNN (Text Convolutional Neural Networks, using convolutional neural networks for text classification); processing methods based on unimodal image classification models, such as ResNet (Residual Network), etc.

[0003] If a single-modal method (i.e., only using images or text) is used to solve multimodal information processing, information from other modalities will be wasted, and the accuracy of the intent classification results will be low. Summary of the invention

[0004] In view of this, embodiments of the present application provide a multimodal information processing method, device, electronic device, and storage medium.

[0005] In a first aspect, an embodiment of the present application provides a multimodal information processing method, the method comprising: extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities; based on an attention mechanism, using the second modal feature, performing attention weighted processing on the first modal feature to obtain a first weighted modal feature; performing multimodal decomposition bilinear pooling processing on the first weighted modal feature and the second modal feature to obtain a bimodal vector; and generating an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0006] In the second aspect, an embodiment of the present application provides a multimodal information processing device, including: an extraction module, used to extract a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities; a weighting module, used to perform attention weighted processing on the first modal feature based on an attention mechanism and using the second modal feature to obtain a first weighted modal feature; a pooling module, used to perform multimodal decomposition bilinear pooling processing on the first weighted modal feature and the second modal feature to obtain a bimodal vector; a first generation module, used to generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps of any multimodal information processing method described in the embodiment of the present application are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any multimodal information processing method described in the embodiments of the present application.

[0009] In an embodiment of the present application, based on the attention mechanism, the second modal feature is utilized to perform attention-weighted processing on the first modal feature to obtain the first weighted modal feature, and then the first weighted modal feature and the second modal feature are subjected to multimodal decomposition and bilinear pooling processing to obtain a bimodal vector, and then the intent classification result is generated, so that the first modal feature and the second modal feature can be fully interactively integrated to obtain a more accurate intent classification result. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 A schematic diagram of a multimodal information processing method according to an embodiment of the present application;

[0011] Figure 2 A schematic diagram of the structure of a neural network model based on a top-down attention mechanism according to an embodiment of the present application;

[0012] Figure 3 This is a schematic diagram of the structure of an MFB according to an embodiment of the present application;

[0013] Figure 4 A schematic diagram of multimodal information according to an embodiment of the present application;

[0014] Figure 5 This is a schematic diagram of the matching relationship between the intent classification result and the knowledge points in an embodiment of the present application;

[0015] Figure 6 This is a schematic diagram of an intent card according to an embodiment of the present application;

[0016] Figure 7 A schematic diagram of a method for generating intent classification results according to an embodiment of the present application;

[0017] Figure 8 A schematic diagram of the structure of a multimodal information processing device according to an embodiment of the present application;

[0018] Fig. 9 A hardware entity schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0019] The technical solution of the present application is further described in detail below with reference to the accompanying drawings and embodiments.

[0020] Figure 1 The implementation flow diagram of the multimodal information processing method provided in the embodiment of the present application is applied to electronic devices, such as Figure 1 As shown, the method includes:

[0021] Step 102: extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities;

[0022] Among them, each source or form of information can be called a modality; multimodal information can be information composed of at least two different sources or different forms of information; the multimodal information can include at least two of the following information: image information, audio information, and text information; correspondingly, the first modal feature and the second modal feature can be image features, audio features, text features, etc., and the first modal feature and the second modal feature are features of different modalities.

[0023] Step 104: Based on the attention mechanism, using the second modal feature, perform attention weighting processing on the first modal feature to obtain a first weighted modal feature;

[0024] Among them, in cognitive science, due to the bottleneck of information processing, part of all information will be selectively focused on, while other visible information will be ignored. The above mechanism is usually called an attention mechanism. The attention mechanism can be a top-down conscious attention (Top Down Attention), called focused attention; focused attention refers to attention that is intentional, task-dependent, and actively and consciously focused on an object. The attention mechanism can also be a bottom-up unconscious attention (Bottom Up Attention), called salience-based attention. Salience-based attention is attention driven by external stimuli, does not require active intervention, and is not related to the task.

[0025] The number of the first modal features may be at least one; the attention weighted processing may be to use the second modal features to assign a corresponding attention weight to each of the first modal features, and to perform weighted summation processing on each first modal feature according to the attention weight of each first modal feature to obtain a first weighted modal feature.

[0026] Step 106: performing multimodal decomposition bilinear pooling processing on the first weighted modal feature and the second modal feature to obtain a bimodal vector;

[0027] Among them, multi-modal factorized bilinear pooling (MFB) processing can be used to fuse the features of the two modalities; the bimodal vector is a vector obtained by fusing the first weighted modal feature and the second modal feature.

[0028] Step 108: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0029] Among them, the bimodal vector can be input into a recurrent neural network, and the bimodal vector can be trained using the recurrent neural network to generate a context vector; and the context vector can be used to generate an intent classification result; assuming that the multimodal information is text information and image information sent by the user, when the content of the text information and image information is a consultation question about price and insurance, the intent classification result can be "order and logistics screenshot"; when the content of the text information is "help me order fried chicken to be delivered to XX hotel" and the content of the image information is a picture of fried chicken, the intent classification result can be "ordering food".

[0030] In the embodiment of the present application, not only is attention weighted processing performed on the first modal feature based on the attention mechanism and the second modal feature to obtain the first weighted modal feature, but multimodal decomposition and bilinear pooling processing are also performed on the first weighted modal feature and the second modal feature to obtain a bimodal vector, thereby generating an intent classification result, so that each point of the first modal feature and the second modal feature can interact with each other, resulting in a better fusion effect.

[0031] The present application also provides a multimodal information processing method, the method comprising:

[0032] Step S202: extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities;

[0033] Wherein, the first modal feature and the second modal feature are each at least one;

[0034] Step S204: mapping at least one first modal feature and at least one second modal feature to a first spatial dimension;

[0035] Among them, assuming that the first modal feature is an M-dimensional feature, the second modal feature is an N-dimensional feature, and the first space dimension is a P-dimensional space, the M-dimensional first modal feature and the N-dimensional second modal feature can be mapped to the P-dimensional space; that is, the first modal feature is changed from an M-dimensional feature to a P-dimensional feature, and the second modal feature is changed from an N-dimensional feature to a P-dimensional feature.

[0036] For each of the second modal features, the following steps S206 to S210 are performed:

[0037] Step S206: determining, in the first spatial dimension, a correlation between each of the first modal features and the second modal features based on an attention mechanism;

[0038] Step S208: determining an attention weight corresponding to the first modal feature according to each of the correlations;

[0039] Among them, the correlation and the attention weight may be proportional, that is, the higher the correlation between the first modal feature and the second modal feature, the greater the attention weight of the first modal feature.

[0040] Step S210: using each of the attention weights, performing weighted summation on each of the first modal features to obtain a first weighted modal feature corresponding to the second modal feature.

[0041] The method described in steps S206 to S210 may be used to determine the first weighted modal feature corresponding to each second modal feature.

[0042] Step S212: performing multimodal decomposition bilinear pooling processing on each of the first weighted modal features and the corresponding second modal features to obtain a bimodal vector;

[0043] Step S214: Generate an intent classification result corresponding to the multimodal information based on each of the bimodal vectors.

[0044] Among them, each bimodal vector can be processed to obtain a processed bimodal vector, and based on the processed bimodal vector, an intention classification result corresponding to the multimodal information is generated.

[0045] In an embodiment of the present application, the attention weight of the first modal feature is determined by the correlation between the first modal feature and the second modal feature, so that the determination of the attention weight can be more accurate, and the first modal feature and the second modality can be integrated to a certain extent.

[0046] The present application also provides a multimodal information processing method, the method comprising:

[0047] Step S302: extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities;

[0048] Wherein, the first modal feature and the second modal feature are each at least one;

[0049] Step S304: mapping at least one first modal feature and at least one second modal feature to a first spatial dimension;

[0050] It is assumed that the first modal feature is an image feature, the second modal feature is a text feature, and the attention mechanism is a top-down attention mechanism. Figure 2 This is a schematic diagram of the structure of a neural network model based on a top-down attention mechanism in an embodiment of the present application, see Figure 2 , W can be represented as a fully connected layer, softmax represents a softmax layer, and the image feature 201 and the text feature 202 can be first mapped to a 512-dimensional spatial dimension, and then the image feature and the text feature can be mapped to a 2048-dimensional spatial dimension.

[0051] For each of the second modal features, the following steps S306 to S314 are performed:

[0052] Step S306: in the first spatial dimension, based on the attention mechanism, determining a first feature vector corresponding to each of the first modal features and a second feature vector corresponding to the second modal features;

[0053] Step S308: determining the dot product between each of the first eigenvectors and the second eigenvectors;

[0054] Step S310: According to each of the dot products, determine the correlation between the corresponding first modal feature and the second modal feature.

[0055] The dot product and the correlation can be proportional, that is, the larger the dot product, the greater the correlation between the first eigenvector and the second eigenvector, and correspondingly, the more similar the first modal feature and the second modal feature are; Figure 2 , the dot product can be normalized using the softmax function of the softmax layer to obtain the correlation between the first modal feature and the second modal feature.

[0056] Step S312: determining an attention weight corresponding to the first modal feature according to each of the correlations;

[0057] Among them, the relevance and the attention weight may be positively proportional, that is, the greater the relevance, the greater the attention weight of the first modal feature may be considered.

[0058] Step S314: using each of the attention weights, performing weighted summation on each of the first modal features to obtain a first weighted modal feature corresponding to the second modal feature.

[0059] Among them, see Figure 2, ∑ can represent a summation operation, k can represent the number of first modal features, then the determined attention weights of the first modal features can be used to perform a weighted summation operation on the k first modal features to obtain the first weighted modal feature corresponding to the second modal feature; similarly, the first weighted modal feature corresponding to each second modal feature in at least one second modal feature can be determined.

[0060] Step S316: performing multimodal decomposition bilinear pooling processing on each of the first weighted modal features and the corresponding second modal features to obtain a bimodal vector;

[0061] Step S318: Generate an intent classification result corresponding to the multimodal information based on each of the bimodal vectors.

[0062] In the embodiment of the present application, the correlation between the first modal feature and the second modal feature is determined based on the dot product between the first eigenvector and the second eigenvector, so that the correlation can be determined more accurately.

[0063] The present application also provides a multimodal information processing method, the method comprising:

[0064] Step S402: extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities;

[0065] Wherein, the first modal feature and the second modal feature are each at least one;

[0066] Step S404: mapping at least one first modal feature and at least one second modal feature to a first spatial dimension;

[0067] For each of the second modal features, the following steps S406 to S422 are performed:

[0068] Step S406: in the first spatial dimension, based on the attention mechanism, determining a first feature vector corresponding to each of the first modal features and a second feature vector corresponding to the second modal features;

[0069] Step S408: determining the dot product between each of the first eigenvectors and the second eigenvectors;

[0070] Step S410: Determine the correlation between the corresponding first modal feature and the second modal feature according to each of the dot products.

[0071] Step S412: determining an attention weight corresponding to the first modal feature according to each of the correlations;

[0072] Step S414: using each of the attention weights, performing weighted summation on each of the first modal features to obtain a first weighted modal feature corresponding to the second modal feature.

[0073] Step S416: mapping the second modal feature and the corresponding first weighted modal feature to a second spatial dimension;

[0074] Wherein, assuming that the first weighted modal feature is an m-dimensional feature and the second modal feature is an n-dimensional feature, the second spatial dimension may be an o-dimensional spatial dimension; Figure 3 This is a schematic diagram of the structure of an MFB in the embodiment of the present application, see Figure 3 , assuming that the first weighted modal feature is a weighted image feature 301, the weighted image feature 301 can be expressed as x∈R m ; The second modal feature is the text feature 302, which can be expressed as y∈R n , the output of the multimodal decomposition bilinear pooling model can be expressed as Z i ∈R, then Z i The dimension can be o; after mapping the first weighted modal feature to the second spatial dimension, the first weighted modal feature 303 in the second spatial dimension can be obtained, and after mapping the second modal feature to the second spatial dimension, the second modal feature 304 in the second spatial dimension can be obtained.

[0075] Step S418: determining, in the second spatial dimension, a third eigenvector corresponding to the first weighted modal feature and a fourth eigenvector corresponding to the second modal feature;

[0076] Step S420: determining the outer product between the third eigenvector and the fourth eigenvector;

[0077] Among them, the multi-modal decomposition bilinear pooling model can be defined as shown in the following formula (1):

[0078]

[0079] Among them, w i ∈R mxn is a mapping matrix, Z i ∈R is the output of the multimodal decomposition bilinear pooling model. In order to obtain an output of dimension o, all mapping matrices w = [w 1 ,...w o ]∈R mxnxo , the multi-modal decomposition bilinear pooling model can be converted into the following formula (2):

[0080]

[0081] where k is the potential dimension of the factor or factorized matrix, I represents the identity matrix, and u i =[u 1 ,...u k ]∈R mxk , v i =[v 1 ,...v k ]∈R nxk , It can be the outer product 305 between the third eigenvector and the fourth eigenvector, which can also be called the Hadmard product or the element-wise multiplication of the third eigenvector and the fourth eigenvector. The attention weight to be learned is u=[u 1 ,...u k ]∈R mxkxo , v=[v 1 ,...v k ]∈R nxkxo .

[0082] Step S422: performing summing and pooling on the outer products to obtain a bimodal vector corresponding to the first weighted modal feature;

[0083] The outer product 305 can be summed and pooled using the SumPooling function; U and V can be rewritten as Then the bimodal vector 306 can be expressed using the following formula (3):

[0084]

[0085] Among them, the function SumPooling Represents the use of a one-dimensional non-overlapping window of size k Perform SumPooling.

[0086] Assume x 1 ∈R ko ,y 1 ∈R ko , Then formula (3) can be converted to:

[0087] Z = SumPooling(x 1 oy 1 ,k) Formula (4);

[0088] The method described in steps S406 to S422 may be used to determine the first weighted modal feature corresponding to each second modal feature, and then determine the corresponding bi-modal vector.

[0089] Step S424: Generate an intent classification result corresponding to the multimodal information based on each of the bimodal vectors.

[0090] In an embodiment of the present application, the second modal feature is used to perform attention weighted processing on the first modal feature to obtain the first weighted modal feature, and then the first weighted modal feature and the second modal feature are subjected to multimodal decomposition and bilinear pooling processing, so that each point of the first modal feature and the second modal feature can interact with each other, resulting in a better fusion effect.

[0091] The present application also provides a multimodal information processing method, the method comprising:

[0092] Step S502: extracting image features and text features from the acquired multimodal information;

[0093] Step S504: Based on the attention mechanism, using the text features, performing attention weighting processing on the image features to obtain weighted image features;

[0094] Step S506: performing multimodal decomposition and bilinear pooling processing on the weighted image features and the text features to obtain a bimodal vector;

[0095] Step S508: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0096] In the embodiment of the present application, not only is attention weighted processing performed on image features based on the attention mechanism and text features to obtain weighted image features, but multimodal decomposition and bilinear pooling processing are also performed on the weighted image modal features and text features to obtain a bimodal vector, thereby generating an intent classification result, thereby enabling each point of the image features and the text features to interact and achieve a better fusion effect.

[0097] The present application also provides a multimodal information processing method, the method comprising:

[0098] Step S602: using a convolutional neural network to extract image features from the acquired multimodal information;

[0099] Among them, the convolutional neural network (CNN) can be a type of feedforward neural network that includes convolution calculations and has a deep structure. The convolutional neural network can be ResNet (Residual Network) or Fast R-CNN (Fast Region-based Convolution Neural Networks), etc.

[0100] Step S604: extracting text features from the multimodal information using a recurrent neural network.

[0101] Among them, a recurrent neural network (RNN) is a type of recurrent neural network that takes sequence data as input, recursively in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain; the recurrent neural network can be a GRU (Gate Recurrent Unit) or an LSTM (Long Short-Term Memory), etc.

[0102] Wherein, each of the image feature and the text feature is at least one;

[0103] Step S606: mapping at least one image feature and at least one text feature to a first spatial dimension;

[0104] For each of the text features, perform the following steps:

[0105] Step S608: Under the first spatial dimension, based on the attention mechanism, determine a first feature vector corresponding to each of the image features and a second feature vector corresponding to the text feature;

[0106] Among them, when the convolutional neural network is ResNet, the attention mechanism can be a Top DownAttention mechanism; when the convolutional neural network is Fast R-CNN, the attention mechanism can be a Bottom Up Attention mechanism.

[0107] Step S610: determining the dot product between each of the first eigenvectors and the second eigenvectors;

[0108] Step S612: According to each of the dot products, determine the correlation between the corresponding image feature and the text feature.

[0109] Step S614: determining the attention weight of the corresponding image feature according to each of the correlations;

[0110] Step S616: using each of the attention weights, perform weighted summation on each image feature to obtain a weighted image feature corresponding to the text feature.

[0111] Step S618: Mapping each of the weighted image features and the corresponding text features to a second spatial dimension;

[0112] Step S620: determining a third feature vector corresponding to the weighted image feature and a fourth feature vector corresponding to the text feature in the second spatial dimension;

[0113] Step S622: determining the outer product between the third eigenvector and the fourth eigenvector;

[0114] Step S624: performing summing and pooling on the outer product to obtain a bimodal vector corresponding to the weighted image feature;

[0115] Step S626: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0116] In the embodiment of the present application, not only the image features are weighted by attention based on the attention mechanism and text features to obtain weighted image features, but also the weighted image features and text features are multimodally decomposed and bilinearly pooled to obtain bimodal vectors, thereby generating intent classification results, so that each point of the image features and text features can interact with each other, and the fusion effect is better. In addition, convolutional neural networks are used to extract image features from the acquired multimodal information; recurrent neural networks are used to extract text features from the multimodal information, so that the extraction of image features and text features can be more accurate.

[0117] The present application also provides a multimodal information processing method, the method comprising:

[0118] Step S702: using a convolutional neural network to extract image features from the acquired multimodal information;

[0119] Figure 4 A schematic diagram of multimodal information provided in an embodiment of the present application, see Figure 4, assuming that the multimodal information processing method of the embodiment of the present application is applied to a dialogue scenario between a customer 401 (or user) and a merchant (or business), and the customer 401 sends multimodal information on the interactive interface, and the multimodal information includes a plain text message 402 "Why is there no movement", and a screenshot 403 of an order and logistics; a convolutional neural network can be used to extract image features from the plain text message 402 and the screenshot 403 of the order and logistics, such as the color features, texture features, shape features, and spatial relationship features contained in the picture of shoes in the screenshot 403 of the order and logistics.

[0120] Step S704: Utilize a recurrent neural network to extract text features from the multimodal information.

[0121] Among them, a recurrent neural network can be used to extract text features such as "why", "what", "what's not", "no movement", "movement", etc. from the plain text information 402; text features such as "in progress", "out of warehouse", "entering", "third party", "seller", "warehouse", "preparation", "out of warehouse", "reminder", "order", "logistics", etc. can also be extracted from the screenshot 403 of orders and logistics.

[0122] Step S706: Based on the attention mechanism, using the text features, performing attention weighting processing on the image features to obtain weighted image features;

[0123] Step S708: performing multimodal decomposition and bilinear pooling processing on the weighted image features and the text features to obtain a bimodal vector;

[0124] Step S710: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0125] Step S712: Determine the knowledge points that match the intention classification result;

[0126] Step S714: Generate a multimodal interactive text based on the knowledge points.

[0127] In an embodiment of the present application, by generating a multimodal interactive text based on knowledge points that match the intent classification results, the generation efficiency and accuracy of the multimodal interactive text can be improved.

[0128] The present application also provides a multimodal information processing method, the method comprising:

[0129] Step S802: using a convolutional neural network to extract image features from the acquired multimodal information;

[0130] Step S804: Utilize a recurrent neural network to extract text features from the multimodal information.

[0131] Step S806: Based on the attention mechanism, using the text features, performing attention weighting processing on the image features to obtain weighted image features;

[0132] Step S808: performing multimodal decomposition and bilinear pooling processing on the weighted image features and the text features to obtain a bimodal vector;

[0133] Step S810: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0134] Step S812: determining at least one candidate knowledge point that matches the intention classification result;

[0135] Figure 5 A schematic diagram of the matching relationship between the intent classification result and the knowledge point provided in the embodiment of the present application, see Figure 5 , the intent classification result "Order and Logistics Screenshot" 501 can be generated according to the multimodal information; the Order and Logistics Screenshot 501 can be regarded as a label of the multimodal information; the Order and Logistics Screenshot 501 can be first mapped to the selected knowledge points, which are the most likely reply knowledge points corresponding to the intent classification results obtained through massive data information statistics and manual annotation. The selected knowledge points include: Query logistics 5021, no logistics update 5022, when to ship out 5023, return logistics 5024, whether to deliver 5025 and no logistics record 5026, etc.

[0136] Step S814: displaying the at least one knowledge point to be selected;

[0137] Figure 6 A schematic diagram of an intent card provided in an embodiment of the present application, see Figure 6 , an intention card 603 may be popped up on the interactive interface according to the usage frequency of the candidate knowledge points, and the intention card 603 is used to display the possible inquiry intention of the customer 601 , and part or all of the candidate knowledge points may be displayed on the intention card 603 .

[0138] Step S816: determining a target knowledge point from the at least one candidate knowledge point according to the received instruction;

[0139] Among them, the instruction can be a selection instruction for the selected knowledge point generated according to the click operation or input operation of the user 601; assuming that the customer 601 clicks or inputs the "When will the goods be shipped out" in the intention card 603, the electronic device 602 operated by the merchant can determine the selected knowledge point "When will the goods be shipped out" as the target knowledge point.

[0140] Step S818: Determine the target knowledge point as a knowledge point that matches the intention classification result.

[0141] Step S820: Generate a multimodal interactive text based on the knowledge points.

[0142] Among them, the electronic device can determine the reply content 605 that matches the corresponding selected knowledge point 604 "When will it be shipped out" according to the click operation of the customer 601, and display the reply content as a multimodal interactive text on the interactive interface for the user to view. The reply content matching the selected knowledge point can be configured by the merchant according to its actual situation, or the electronic device can automatically configure it according to the selected knowledge point; when the selected knowledge point is "When will it be shipped out", the reply content can be configured as "Hello, I'm very sorry~ Due to the serious disaster in place A, the express delivery in place A has been seriously affected, and the specific delivery time cannot be guaranteed."

[0143] In the embodiment of the present application, the target knowledge point can be determined from the candidate knowledge points according to the received instruction, so that the determination of the knowledge point can be made more flexible.

[0144] The present application also provides a multimodal information processing method, the method comprising:

[0145] Step S902: using a convolutional neural network to extract image features from the acquired multimodal information;

[0146] Step S904: Utilize a recurrent neural network to extract text features from the multimodal information.

[0147] Step S906: Based on the attention mechanism, using the text features, performing attention weighting processing on the image features to obtain weighted image features;

[0148] Step S908: performing multimodal decomposition and bilinear pooling processing on the weighted image features and the text features to obtain a bimodal vector;

[0149] Step S910: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0150] Step S912: determining at least one candidate knowledge point that matches the intention classification result;

[0151] Step S914: determining the frequency of each candidate knowledge point being determined as a target knowledge point at a historical moment;

[0152] The frequency is also called the usage frequency of the candidate knowledge point.

[0153] Step S916: sorting the at least one candidate knowledge point according to the descending order of frequency;

[0154] Step S918: Displaying the at least one knowledge point to be selected in order of arrangement.

[0155] Among them, the top three candidate knowledge points with high frequency responses can be displayed on the intention card 603: when to ship out of the warehouse, no logistics record, and no logistics update.

[0156] Step S920: determining a target knowledge point from the at least one candidate knowledge point according to the received instruction;

[0157] The instruction may be a selection instruction of a candidate knowledge point generated according to a click operation or an input operation of the user 601 .

[0158] Step S922: Determine the target knowledge point as a knowledge point that matches the intention classification result.

[0159] Step S924: Generate a multimodal interactive text based on the knowledge points.

[0160] In the embodiment of the present application, the display order of the candidate knowledge points can be determined according to the usage frequency of the candidate knowledge points, thereby improving the user's interactive experience.

[0161] The present application also provides a multimodal information processing method, the method comprising:

[0162] Step S1002: using a convolutional neural network to extract image features from the acquired multimodal information;

[0163] Step S1004: extracting text features from the multimodal information using a recurrent neural network.

[0164] Step S1006: Based on the attention mechanism, using the text features, performing attention weighting processing on the image features to obtain weighted image features;

[0165] Step S1008: performing multimodal decomposition and bilinear pooling processing on the weighted image features and the text features to obtain a bimodal vector;

[0166] Step S1010: Generate an intent classification result corresponding to the multimodal information based on the bimodal vector.

[0167] Step S1012: determining at least one candidate knowledge point that matches the intention classification result;

[0168] Step S1014: determining the frequency of each candidate knowledge point being determined as a target knowledge point at a historical moment;

[0169] Step S1016: determining the candidate knowledge point with the highest frequency as the knowledge point matching the intention classification result;

[0170] Among them, the candidate knowledge point with the highest frequency may also be directly determined as the knowledge point that matches the intention classification result.

[0171] Step S1018: Generate a multimodal interactive text based on the knowledge points.

[0172] In an embodiment of the present application, the candidate knowledge point with the highest frequency can also be directly determined as the knowledge point that matches the intention classification result, thereby improving the intelligence of the knowledge point determination.

[0173] Multimodal response means that the computer needs to make an intelligent response to a given piece of information containing text and images. Multimodal response can be called multimodal image response, where text is a modality and image is another modality. Information that contains both images and text is called multimodal. Unlike general response methods based on image or text classification, multimodality involves two or more information streams, so it is more difficult and challenging to solve. Multimodal information response has extremely high application value in intelligent customer service such as e-commerce, because in the customer service conversation, users will not only send plain text information, but also may contain image information, see Figure 4 , the user sent a plain text message "Why is there no movement", as well as image information, which is a screenshot of the order and logistics. The traditional answering method based on image or text classification cannot solve this kind of multimodal problem well. Therefore, multimodal image answering can not only save labor costs, but also answer users' questions more quickly and accurately.

[0174] In the related technology, the solutions for multimodal image response mainly include: response method based on classification or matching model of unimodal language, such as TextCNN (Text Convolutional Neural Networks, using convolutional neural network for text classification); response method based on unimodal image classification model, such as ResNet (Residual Network); classification response method based on attention mechanism for multimodal fusion, which usually uses attention to fuse two kinds of information, such as SAN (Stacked Attention Networks) and other models. SAN is a VQA (Visual Question Answering) network that uses a double-layer attention mechanism.

[0175] If a single-modal approach (i.e., only using images or text) is used to solve multimodal responses, information from other modalities will be wasted, and the accuracy of the intent classification results will be low. Methods based on the attention mechanism use the interaction of modal information and perform better than single-modal methods, but the accuracy is still poor.

[0176] At the same time, the reply based on the intent classification results requires a complete set of post-processing methods. Simply replying with the intent classification results will make the answer seem not detailed enough and cannot solve the user needs well.

[0177] The reason why unimodal classification methods perform poorly is that the multimodal information has a complex information flow, and using only one type of information cannot fully understand the context. Figure 4 If only text information is used without understanding the image information, the customer service will not be able to understand that the user means "there is no movement in the logistics" based on the text information "Why is there no movement", and will lead to incorrect responses due to insufficient understanding of the information.

[0178] When using the Attention mechanism to interact with multimodal information, only the image features are weighted, and the text features in the text information only play a role in generating the attention weights of the image features. This has a limited degree of fusion and interaction, which may cause some important information in the text to be ignored, resulting in poor classification results. At the same time, the reply based on the intent classification result requires a complete set of processing methods. Simply replying with the intent classification result will make the answer seem not detailed enough. See Figure 4 If you only reply with "order and logistics screenshots", users may not understand what you are saying, and user problems cannot be solved well, and the interactivity is also relatively poor.

[0179] The embodiment of the present application proposes a two-stage interactive multimodal response method, which can utilize multiple modal information and enable the information to be fully interactively integrated to obtain accurate multimodal intent classification. At the same time, in order to make the reply more detailed, the embodiment of the present application adopts a "multimodal intent classification-intent mapping" reply method based on the two-stage interactive method, so as to better respond.

[0180] The embodiment of the present application provides a multimodal response method, which is applied to an electronic device, and the method comprises the following steps:

[0181] Step S1102: Generate intent classification results through multimodal information.

[0182] In an application scenario, a user asks about price insurance and sends a text message "Why is there no movement" and image information on the interactive interface: screenshots of order information and logistics information; through the two-stage multimodal intent classification method provided in the embodiment of the present application, the intent classification result can be obtained as "order and logistics screenshots".

[0183] In the entire multimodal response process, multimodal intent recognition is a very important part, because the accuracy of the user's intent classification result is directly related to the quality of the subsequent response. The embodiment of the present application proposes a two-stage interactive multimodal intent classification model to improve the quality of intent classification results.

[0184] Figure 7 A schematic diagram of a method for generating intent classification results provided in an embodiment of the present application is shown in FIG7 . First, image features 705 and text features 706 can be extracted. A plurality of image features 705 can be extracted from image information 701 through a convolutional neural network ResNet 703, and text features 706 can be extracted from text information 702 through a double-layer recurrent neural network GRU (Gate Recurrent Unit) 704.

[0185] Secondly, in the first stage of multimodal interaction, the attention mechanism (Attention) 707 can be used to fuse multiple image features and text features to obtain attention-weighted weighted image features 709.

[0186] Among them, the first stage of information interaction, i.e., fusion of the attention stage, is performed on the two types of information (i.e., image features 705 and text features 706) through Top down Attention (top-down attention model). The information of the two modes is first mapped to a common spatial dimension, and then dot product fusion is performed, i.e., the dot product between each image feature and the text feature is calculated, and each of the dot products is normalized through the softmax function to obtain the attention weight 708 of each image feature; according to the attention weight 708 of each image, the image features are weighted and summed to obtain the attention-weighted weighted image feature 709.

[0187] However, the attention mechanism at this time only performs weighted summation on the image features, while text information is also a key information in multimodal dialogue. Therefore, in the second stage of multimodal interaction, the attention-weighted weighted image features 709 and text features 706 are subjected to MFB (Multi-modal Factorized Bilinear Pooling) processing 710 to obtain a multimodal vector, and based on the multimodal vector, the intent classification result 711 corresponding to the multimodal information is generated.

[0188] Among them, the MFB structure interacts with each point of the text feature and the weighted image feature after attention weighting, and the fusion effect is better. The MFB structure is as follows Figure 3 As shown in the figure, the two features (weighted image features and text features after attention weighting) are mapped to a larger spatial dimension respectively, and then the outer product operation is performed on the attention weighted image features and text features, and finally the sum and pooling are performed. This allows the two features to be fully integrated and retains the text semantic information that is "ignored" in the first stage to the greatest extent, which is more conducive to multimodal classification.

[0189] Step S1104: Map the intention classification result to the knowledge point to be selected.

[0190] Among them, the candidate knowledge points are the most likely reply knowledge points corresponding to the classification obtained through massive data information statistics and manual annotation. Figure 5, the selected knowledge points that match the "order and logistics screenshot" include: query logistics, no logistics update, when to ship out, return logistics, whether to deliver, and no logistics record. According to the frequency of use of the selected knowledge points, an intention card can be popped up on the interactive interface. The intention card is used to display the user's possible inquiry intention, and the intention card can display the top three high-frequency reply selected knowledge points. The user clicks on the corresponding selected knowledge point in the intention card, and the electronic device determines the reply content that matches the corresponding selected knowledge point according to the user's click operation, and displays the reply content as a multimodal interactive text on the interactive interface for the user to view. Among them, the reply content matched by the selected knowledge point can be configured by the merchant according to its actual situation. When the selected knowledge point is "when to ship out", the reply content can be configured as "Hello, I'm very sorry~ Due to the serious disaster in A, the express delivery in A has been seriously affected, and the specific delivery time cannot be guaranteed."

[0191] Table 1 shows the multimodal response method of the two-stage interactive mode provided in the embodiment of the present application, as well as the classification accuracy Acc, precision Precision, recall Recall and the harmonic mean F1 of the precision and recall of the single-modal image classification model ResNet, the stacked attention network model SAN, and the top-down attention model Top down Attention;

[0192] Table 1

[0193] Model Acc Precision Recall F1 ResNet 0.901 0.769 0.730 0.749 SAN 0.889 0.747 0.700 0.716 Top down Attention 0.903 0.762 0.736 0.750 Two-stage interaction 0.927 0.766 0.782 0.766

[0194] Referring to Table 1, the two-stage interactive multimodal response method provided in the embodiment of the present application achieves the best results in the two most important indicators Acc (accuracy) and F1 (harmonic mean of precision and recall), thereby improving the accuracy of determining the intent classification results.

[0195] Based on the foregoing embodiments, the embodiments of the present application provide a multimodal information processing device, which includes the modules included and can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU, Central Processing Unit), a microprocessor (MPU, Microprocessor Unit), a digital signal processor (DSP, Digital Signal Processing) or a field programmable gate array (FPGA, Field Programmable Gate Array), etc.

[0196] Figure 8Schematic diagram of the composition structure of the multimodal information processing device according to the embodiment of the present application. Figure 8 As shown, the apparatus 800 includes an extraction module 801, a weighting module 802, a pooling module 803 and a first generation module 804, wherein:

[0197] An extraction module 801 is used to extract a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities;

[0198] A weighting module 802 is used to perform attention weighting processing on the first modal feature by using the second modal feature based on an attention mechanism to obtain a first weighted modal feature;

[0199] A pooling module 803 is used to perform multimodal decomposition bilinear pooling processing on the first weighted modal feature and the second modal feature to obtain a bimodal vector;

[0200] The first generating module 804 is used to generate an intention classification result corresponding to the multimodal information according to the bimodal vector.

[0201] In one embodiment, when there is at least one of the first modal feature and the second modal feature, the weighting module 802 includes: a first mapping submodule, used to map at least one first modal feature and at least one second modal feature to a first spatial dimension; a first determination submodule, used to determine, for each of the second modal features, the correlation between each of the first modal feature and the second modal feature in the first spatial dimension based on an attention mechanism; a second determination submodule, used to determine, for each of the second modal features, the attention weight corresponding to the first modal feature according to each of the correlations; and a weighting submodule, used to perform weighted summation of the first modal features for each of the second modal features using each of the attention weights to obtain a first weighted modal feature corresponding to the second modal feature.

[0202] In one embodiment, the first determination submodule includes: a first determination unit, used to determine, for each of the second modal features, a first eigenvector corresponding to each of the first modal features and a second eigenvector corresponding to the second modal feature; a second determination unit, used to determine, for each of the second modal features, a dot product between each of the first eigenvectors and the second eigenvectors; and a third determination unit, used to determine, for each of the second modal features, a correlation between the corresponding first modal feature and the second modal feature based on each of the dot products.

[0203] In one embodiment, when both the first weighted modality feature and the second modality feature are at least one, the pooling module 803 includes: a second mapping sub-module, configured to map each of the first weighted modality features and the corresponding second modality features to a second spatial dimension; a third determination sub-module, configured to determine, in the second spatial dimension, a third feature vector corresponding to the first weighted modality feature and a fourth feature vector corresponding to the second modality feature; a fourth determination sub-module, configured to determine an outer product between the third feature vector and the fourth feature vector; and a pooling sub-module, configured to perform sum pooling on the outer product to obtain a bimodal vector corresponding to the first weighted modality feature.

[0204] In one embodiment, the first modality feature includes an image feature, and the second modality information includes a text feature. The extraction module includes: a first extraction sub-module, configured to extract an image feature from the acquired multimodal information by using a convolutional neural network; and a second extraction sub-module, configured to extract a text feature from the multimodal information by using a recurrent neural network.

[0205] In one embodiment, the apparatus further includes: a determination module, configured to determine knowledge points that match the intent classification result; and a second generation module, configured to generate multimodal interaction text according to the knowledge points.

[0206] In one embodiment, the determination module includes: a fifth determination sub-module, configured to determine at least one candidate knowledge point that matches the intent classification result; a display sub-module, configured to display the at least one candidate knowledge point; a sixth determination sub-module, configured to determine a target knowledge point from the at least one candidate knowledge point according to a received instruction; and a seventh determination sub-module, configured to determine the target knowledge point as the knowledge point that matches the intent classification result.

[0207] In one embodiment, the display sub-module includes: a fourth determination unit, configured to determine the frequency at which each of the candidate knowledge points was determined as the target knowledge point at a historical moment; a sorting unit, configured to sort the at least one candidate knowledge point in a descending order according to the frequency; and a display unit, configured to display the at least one candidate knowledge point in the sorted order.

[0208] In one embodiment, the determination module includes: an eighth determination sub-module, configured to determine at least one candidate knowledge point that matches the intent classification result; a ninth determination sub-module, configured to determine the frequency at which each of the candidate knowledge points was determined as the target knowledge point at a historical moment; and a tenth determination sub-module, configured to determine the candidate knowledge point with the highest frequency as the knowledge point that matches the intent classification result.

[0209] It should be noted that in the embodiment of the present application, if the above-mentioned multimodal information processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a mobile phone, a tablet computer, a desktop computer, a personal digital assistant, a digital phone, a video phone, a television, a sensor device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), magnetic disk or optical disk and other media that can store program codes. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0210] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0211] Correspondingly, an embodiment of the present application provides an electronic device, Fig. 9 A hardware entity diagram of an electronic device according to an embodiment of the present application is shown in FIG. Fig. 9 As shown, the hardware entity of the electronic device 900 includes: a memory 901 and a processor 902, wherein the memory 901 stores a computer program that can be run on the processor 902, and the processor 902 implements the steps in the multimodal information processing method of the above embodiment when executing the program.

[0212] The memory 901 is configured to store instructions and applications executable by the processor 902, and can also cache data to be processed or processed by the processor 902 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0213] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the multimodal information processing method provided in the above embodiment are implemented.

[0214] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the device embodiments. For technical details not disclosed in the storage medium and method embodiments of this application, please refer to the description of the device embodiments of this application for understanding.

[0215] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0216] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0217] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0218] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the embodiments of the present application may be all integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0219] It can be understood by a person of ordinary skill in the art that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiment are executed; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, read-only memory (ROM), magnetic disks or optical disks. Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a mobile phone, a tablet computer, a desktop computer, a personal digital assistant, a digital phone, a video phone, a television, a sensor device, etc.) to execute all or part of the methods described in each embodiment of the present application. And the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, magnetic disks or optical disks.

[0220] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain a new method embodiment. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain a new product embodiment. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain a new method embodiment or device embodiment.

[0221] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A multimodal information processing method, It is characterized in that The method comprises: Extracting a first modal feature and a second modal feature from the acquired multimodal information; the first modal feature and the second modal feature are features of two different modalities; wherein the multimodal information includes at least two of image information, audio information, and text information; Based on the attention mechanism, the second modal feature is used to perform attention weighting processing on the first modal feature to obtain a first weighted modal feature; wherein the number of the first modal features is at least one, and the attention weighting processing is used to use the second modal feature to assign a corresponding attention weight to each of the first modal features, and according to the attention weight of each of the first modal features, weighted summing processing is performed on each of the first modal features to obtain the first weighted modal feature; Performing multimodal decomposition and bilinear pooling processing on the first weighted modal features and the second modal features to obtain a bimodal vector; Based on the bimodal vector, an intent classification result corresponding to the multimodal information is generated.

2. The method according to claim 1, It is characterized in that The second modal feature is at least one, and the attention-based mechanism uses the second modal feature to perform attention-weighted processing on the first modal feature to obtain a first weighted modal feature, including: mapping at least one first modality feature and at least one second modality feature to a first spatial dimension; For each of the second modal features, perform the following steps: In the first spatial dimension, based on an attention mechanism, determining a correlation between each of the first modal features and the second modal features; According to each of the correlations, determining an attention weight corresponding to the first modality feature; Using each of the attention weights, weighted summation is performed on each of the first modal features to obtain a first weighted modal feature corresponding to the second modal feature.

3. The method according to claim 2, It is characterized in that The determining, in the first spatial dimension, based on the attention mechanism, the correlation between each of the first modal features and the second modal features includes: Determine a first eigenvector corresponding to each of the first modal features and a second eigenvector corresponding to each of the second modal features; determining a dot product between each of the first eigenvectors and the second eigenvector; According to each of the dot products, a correlation between the corresponding first modal feature and the second modal feature is determined.

4. The method according to claim 1, It is characterized in that When both the first weighted modal feature and the second modal feature are at least one, performing multimodal decomposition bilinear pooling processing on the first weighted modal feature and the second modal feature to obtain a bimodal vector includes: Mapping each of the first weighted modal features and the corresponding second modal features to a second spatial dimension; In the second spatial dimension, determining a third eigenvector corresponding to the first weighted modal feature and a fourth eigenvector corresponding to the second modal feature; determining an outer product between the third eigenvector and the fourth eigenvector; Performing summation pooling on the outer product to obtain the bimodal vector corresponding to the first weighted modality feature.

5. The method according to any one of claims 1 to 4, wherein, the first modality feature includes an image feature, the second modality information includes a text feature, and extracting the first modality feature and the second modality feature from the obtained multimodal information includes: using a convolutional neural network to extract image features from the obtained multimodal information; using a recurrent neural network to extract text features from the multimodal information.

6. The method according to any one of claims 1 to 4, wherein, the method further includes: determining knowledge points matching the intent classification result; generating multimodal interaction text according to the knowledge points.

7. The method according to claim 6, wherein, determining the knowledge points matching the intent classification result includes: determining at least one candidate knowledge point matching the intent classification result; displaying the at least one candidate knowledge point; determining a target knowledge point from the at least one candidate knowledge point according to a received instruction; determining the target knowledge point as the knowledge point matching the intent classification result.

8. The method according to claim 7, wherein, displaying the candidate knowledge points includes: determining the frequency of each candidate knowledge point being determined as the target knowledge point at a historical moment; sorting the at least one candidate knowledge point in descending order according to the frequency; displaying the at least one candidate knowledge point in the sorted order.

9. The method according to claim 6, wherein, determining the knowledge points matching the intent classification result includes: determining at least one candidate knowledge point matching the intent classification result; determining the frequency of each candidate knowledge point being determined as the target knowledge point at a historical moment; determining the candidate knowledge point with the highest frequency as the knowledge point matching the intent classification result.

10. A multimodal information processing device, wherein, the device includes: an extraction module for extracting a first modality feature and a second modality feature from the obtained multimodal information; the first modality feature and the second modality feature are features of two different modalities; wherein, the multimodal information includes at least two of image information, audio information, and text information; a weighting module for performing attention weighting processing on the first modality feature by using the second modality feature based on an attention mechanism to obtain a first weighted modality feature; wherein, the number of the first modality features is at least one, and the attention weighting processing is used to assign corresponding attention weights to each of the first modality features by using the second modality feature, and perform weighted summation processing on each of the first modality features according to the attention weights of each of the first modality features to obtain the first weighted modality feature; a pooling module for performing multimodal decomposition bilinear pooling processing on the first weighted modality feature and the second modality feature to obtain a bimodal vector; The first generating module is used to generate the intent classification result corresponding to the multimodal information according to the bimodal vector.

11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, It is characterized in that When the processor executes the program, the steps in the multimodal information processing method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the multimodal information processing method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Image content question and answer method based on multi-modality low-rank dual-linear pooling

    CN107480206A

  • Multi-modal dialogue system and method guided by user attention

    CN110209789A