Question Answering Method, Device, Equipment and Storage Medium Based on Bimodal Feature Fusion
Through the bimodal feature fusion method, the target area and word vector are extracted using the image and text detection model, which solves the problem of image noise and text-independent words, and generates more accurate answers.
Patent Information
- Application Number
- CN202210633637.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-06
AI Technical Summary
In the prior art, when processing images and text data sets, image features are susceptible to noise interference, and text features are susceptible to unrelated words, resulting in inaccurate feature extraction.
The dual-modal feature fusion method is adopted to extract the target area features through the image detection model and input it into the convolutional neural network to obtain the average image features; the part-of-speech annotation is performed through the text detection model and word vectors are extracted to obtain the average text features; the two are fused and the long and short-term memory neural networks are input to generate the answer.
Effectively removes the influence of image noise and text-independent words, improves the accuracy of feature extraction, and the generated answer text is more accurate.
Smart Images

Figure CN114972792B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data fusion technologies, for example, to a question-answering method, apparatus, device, and storage medium based on bimodal feature fusion. Background Art
[0002] With the development of deep learning, the Internet has generated a large amount of data, especially text and image data. For example, the CLIP model open-sourced by OpenAI is trained on 400 million pairs of image-text datasets; video data is naturally a combination of text and image data. Therefore, it is easy to obtain cross-modal data of images and text. Compared with previous text search tasks, novel tasks such as searching for text based on images and collecting text answers based on text and images have emerged. Among them, Visual Question Answering (VQA) has been widely studied and applied in the past two years. The VQA task refers to a VQA system that takes a series of images and free, open-ended natural language questions related to the series of images as inputs and generates a natural language answer as an output. This task involves multi-modal learning tasks of two datasets, images and text.
[0003] Existing technologies usually need to input the entire image in an image dataset when processing it. The size of the entire image is large, and the background and noise of the image will make the extracted image features inaccurate. Existing technologies usually need to input the entire sentence when processing a text dataset, which will bring many irrelevant words and make the extracted text features inaccurate. Summary of the Invention
[0004] This application provides a question-answering method, apparatus, device, and storage medium based on bimodal feature fusion, aiming to solve the problems of noisy image features and irrelevant words in text features in the question-answering method.
[0005] To solve the above problems, this application adopts the following technical solutions:
[0006] This article provides a question-answering method based on bimodal feature fusion, which is characterized by including:
[0007] Obtain an input image and an input question; extract the target region features of the input image through a trained image detection model, and input the target region features into a trained convolutional neural network to obtain average image features;
[0008] Perform part-of-speech tagging on the input question to obtain a tagged question; extract the word vectors of the tagged question through a trained text detection model to obtain average text features;
[0009] Fuse the average image features and the average text features to obtain fused features; input the fused features into a long short-term memory neural network for decoding to obtain an answer text.
[0010] The target region features include the target region coordinates and the target region image;
[0011] Extracting the target region features of the input image by the trained image detection model includes:
[0012] Obtaining the detection region of the input image;
[0013] Extracting the image category of the detection region and extracting the confidence of the detection region;
[0014] Filtering out the target detection regions where the image category is the target category and the confidence is greater than or equal to the confidence threshold;
[0015] Extracting the target region coordinates of the target detection region;
[0016] Extracting the target region image according to the target region coordinates.
[0017] Inputting the target region features into the trained convolutional neural network to obtain the average image features includes:
[0018] Inputting the target region features into the trained convolutional neural network to obtain image features;
[0019] Calculating the average value of the image features to obtain the average image features.
[0020] Part-of-speech tagging the input question sentence to obtain the tagged question sentence includes:
[0021] Using a part-of-speech tagger to split the input question sentence to obtain split words;
[0022] Performing part-of-speech tagging on all the split words to obtain the tagged question sentence.
[0023] Extracting the word vectors of the tagged question sentence by the trained text detection model to obtain the average text features includes:
[0024] Extracting the question words and related nouns of the tagged question sentence by the trained text detection model;
[0025] Converting the question words and the related nouns into word vectors;
[0026] Calculating the average value of the word vectors to obtain the average text features.
[0027] Fusing the average image features and the average text features to obtain the fused features includes:
[0028] Feature fusion is performed on the average image feature and the average text feature through multi-modal Tucker fusion, multi-modal bilinear fusion, or linear fusion to obtain the fused feature.
[0029] The trained convolutional neural network includes:
[0030] An input layer for preprocessing the target region feature to obtain a preprocessed feature;
[0031] A hidden layer for performing convolution, activation, and pooling on the preprocessed feature to obtain a hidden layer output;
[0032] A fully connected layer for integrating the hidden layer output to obtain an integrated feature;
[0033] An output layer for classifying the integrated feature to obtain the image feature.
[0034] This application also provides a question and answer device based on bimodal feature fusion, including:
[0035] An image and question acquisition module, a target region feature extraction module, an average image feature extraction module, a part-of-speech tagging module, a text detection module, a feature fusion module, and a feature decoding module;
[0036] The image and question acquisition module is used to acquire an input image and an input question;
[0037] The target region feature extraction module is used to extract the target region feature of the input image through a trained image detection model;
[0038] The average image feature extraction module is used to input the target region feature into a trained convolutional neural network to obtain an average image feature;
[0039] The part-of-speech tagging module is used to perform part-of-speech tagging on the input question to obtain a tagged question;
[0040] The text detection module is used to extract the word vector of the tagged question through a trained text detection model to obtain an average text feature;
[0041] The feature fusion module is used to perform feature fusion on the average image feature and the average text feature to obtain a fused feature;
[0042] The feature decoding module is used to input the fused feature into a long short-term memory neural network for decoding to obtain an answer text.
[0043] The present application also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above-mentioned question-answering method based on dual-modal feature fusion are implemented.
[0044] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned question-answering method based on dual-modal feature fusion are implemented.
[0045] For the question-answering method based on dual-modal feature fusion of the present application, the target region features of the input image are extracted by a trained image detection model, and the target region features are input into a trained convolutional neural network to obtain average image features. The target region features are the target parts in the input image, without including the background and having a relatively low noise content. The convolutional neural network can remove the noise in the target feature region, and the obtained average image features can better reflect the characteristics of the target in the image. The input question sentence is subjected to part-of-speech tagging to obtain a tagged question sentence; the word vectors of the tagged question sentence are extracted by a trained text detection model to obtain average text features. Part-of-speech tagging can distinguish the nature of different words in the text, only retaining the interrogative words and relevant nouns, and the average text features obtained based on only the interrogative words and relevant nouns are not affected by other words. When the average image features and the average text features are relatively accurate, the answer text obtained by passing the fusion features of the two through a long short-term memory neural network is also relatively accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic flowchart of the question-answering method based on dual-modal feature fusion according to an embodiment;
[0047] Figure 2 It is a schematic flowchart of the method for extracting the target region features of the input image according to an embodiment;
[0048] Figure 3 It is a schematic flowchart of obtaining average image features by a trained convolutional neural network according to an embodiment;
[0049] Figure 4 It is a schematic flowchart of the method for obtaining average text features by using a trained text detection model according to an embodiment;
[0050] Figure 5 It is a schematic block diagram of the structure of the question-answering device based on dual-modal feature fusion according to an embodiment;
[0051] Figure 6 It is a schematic block diagram of the structure of a computer device according to an embodiment.
[0052] The realization, functional features, and advantages of the present application will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners
[0053] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the above", and "the" used herein may also include the plural forms. It should be further understood that the term "including" used in the description of the present application means the presence of features, integers, steps, operations, elements, units, units, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, units, units, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0055] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0056] Refer to Figure 1 , which is a schematic flowchart of the question-and-answer method based on dual-modal feature fusion of the present application, including:
[0057] S1: Obtain an input image and an input question sentence.
[0058] The input image contains target information, and the input question sentence is a question sentence related to the target information of the input image.
[0059] S2: Extract the target region features of the input image through a trained image detection model.
[0060] The target region features include target region coordinates and a target region image;
[0061] The extracting the target region features of the input image through a trained image detection model includes:
[0062] Obtain the detection region of the input image;
[0063] The size of the input image is large and may contain multiple objects, among which there is a target. Therefore, the input image is detected to obtain multiple detection regions.
[0064] Extract the image category of the detection region and extract the confidence of the detection region;
[0065] According to the pre-prepared labels of image categories, all detection regions of the input image are classified to obtain the image categories of the detection regions. Extract the confidence of the detection region. The range of the confidence is 0-1. The higher the confidence, the more reliable the obtained image category of the detection region.
[0066] Filter out the target detection regions whose image category is the target category and whose confidence is greater than or equal to the confidence threshold;
[0067] Remove the background and noise by combining the image category and the confidence, and filter out the target detection regions.
[0068] Extract the target region coordinates of the target detection region;
[0069] Extract the target region image according to the target region coordinates.
[0070] Extracting the target region coordinates and the target region image obtains the features of the target in the image. The target region coordinates are used to locate the position of the target, and the target region image contains the specific information of the target, such as color, contour and texture features.
[0071] S3: Input the target region features into the trained convolutional neural network to obtain the average image features.
[0072] Input the target region features into the trained convolutional neural network to obtain the image features;
[0073] Calculate the average value of the image features to obtain the average image features.
[0074] The convolutional neural network can extract image features. By calculating the average value of the image features, it can better reflect the essential features of the target in the image.
[0075] S4: Perform part-of-speech tagging on the input question sentence to obtain the tagged question sentence.
[0076] Use a part-of-speech tagger to split the input question sentence to obtain the split words;
[0077] Perform part-of-speech tagging on all the split words to obtain the tagged question sentence.
[0078] The input question sentence is split, and the split words are part-of-speech tagged. Words unrelated to obtaining the answer text, such as adverbs, verbs, and adjectives, can be removed according to the part-of-speech of the split words, improving the operation efficiency while reducing interference.
[0079] S5: Extract the word vectors of the labeled question sentence through the trained text detection model to obtain the average text features.
[0080] Extract the question words and related nouns of the labeled question sentence through the trained text detection model;
[0081] The question words include "who, how, and where". To extract text features, the key is to extract the question of the labeled question sentence, which is composed of question words and related nouns.
[0082] Convert the question words and the related nouns into word vectors;
[0083] Word vectors have the advantage of being convenient for calculation.
[0084] Calculate the average value of the word vectors to obtain the average text features.
[0085] By calculating the average value of the word vectors, the essential features of the input question sentence can be better reflected.
[0086] S6: Perform feature fusion on the average image features and the average text features to obtain fusion features.
[0087] Perform feature fusion on the average image features and the average text features through multi-modal Tucker fusion, multi-modal bilinear fusion, or linear fusion to obtain the fusion features.
[0088] By fusing the average image features and the average text features to obtain fusion features, the image features and text features of the target can be described by one feature simultaneously.
[0089] S7: Input the fusion features into a long short-term memory neural network for decoding to obtain the answer text.
[0090] The long short-term memory network can combine the previously input fusion features and the currently input fusion features to obtain results, and has a good decoding effect on the same type of fusion features.
[0091] In the specific implementation process, on the premise that step S1 is completed, steps S2 - S7 can be executed sequentially, or steps S2 and S4 can be executed in parallel. After S2 is completed, S3 is executed. After S4 is completed, S5 is executed. S6 is executed according to the results of S3 and S5, and finally S7 is executed. Executing steps S2 and S4 in parallel can calculate the target region features of the input image and perform part-of-speech tagging on the input question sentence simultaneously, calculate the average image features based on the target region features, and obtain the average text features according to the tagged question sentence. In the case of better computing resources, by executing steps S2 - S3 and steps S4 - S5 in parallel, the operation efficiency can be improved, and the final answer text can be obtained quickly.
[0092] In addition, on the premise that step S1 is completed, steps S4 - S5 can also be executed first. After step S5 is completed, steps S2, S3, S6, and S7 are executed sequentially. The specific execution order depends on the actual situation and is not limited here.
[0093] The question - answering method based on bimodal feature fusion in the embodiment of this application obtains the input image and the input question sentence, extracts the target region features of the input image through a trained image detection model, and inputs the target region features into a trained convolutional neural network to obtain the average image features. The target region features are the target parts in the input image, do not include the background, and have a relatively low noise content. Through the convolutional neural network, the noise in the target feature region can be removed, and the obtained average image features can better reflect the characteristics of the target in the image. Perform part-of-speech tagging on the input question sentence to obtain the tagged question sentence; extract the word vectors of the tagged question sentence through a trained text detection model to obtain the average text features. Part-of-speech tagging can distinguish the nature of different words in the text, and only retain the interrogative words and relevant nouns. The average text features obtained based on only the interrogative words and relevant nouns are not affected by other words. When the average image features and the average text features are relatively accurate, the answer text obtained by passing the fusion features of the two through a long short-term memory neural network is also relatively accurate.
[0094] In one embodiment, the image category, confidence level, and target region coordinates of the input image are extracted through a trained image detection model. Refer to Figure 2 , which is a schematic flow diagram of the method for extracting the target region features of the input image in this application, including:
[0095] S21: Extract the image category of the detection region.
[0096] The trained image detection model is a trained Yolov5 model, and the Yolov5 model is used to extract the image category of the detection region.
[0097] The Yolov5 model is trained using a classification dataset and can obtain the classification labels corresponding to the input images.
[0098] The Yolov5 model includes a Focus layer, a first convolutional layer, a first Bottleneck layer, a second convolutional layer, a second Bottleneck layer, a third convolutional layer, a third Bottleneck layer, a fourth convolutional layer, a Spatial Pyramid Pooling layer, and a first Bottleneck layer connected in sequence.
[0099] Exemplarily, the parameters of the Focus layer, convolutional layers, Bottleneck layers, and Spatial Pyramid Pooling layer of the Yolov5 model are as follows:
[0100] The number of convolutional kernels in the Focus layer is 1, the number of convolutional kernels in the first convolutional layer is 1, the number of convolutional kernels in the first Bottleneck layer is 3, the number of convolutional kernels in the second convolutional layer is 1, the number of convolutional kernels in the second Bottleneck layer is 9, the number of convolutional kernels in the third convolutional layer is 1, the number of convolutional kernels in the third Bottleneck layer is 9, the number of convolutional kernels in the fourth convolutional layer is 1, and the number of convolutional kernels in the Spatial Pyramid Pooling layer is 1.
[0101] The number of convolutional kernels in each layer of the above Yolov5 model is only an example and is specifically determined according to the actual situation, and is not limited here.
[0102] S22: Extract the confidence of the detection region.
[0103] The Yolov5 model can output the confidence of all detection regions. The confidence of the detection region represents the probability that the region belongs to the image category obtained in step S21.
[0104] S23: Screen out the target detection regions where the image category is the target category and the confidence is greater than or equal to the confidence threshold.
[0105] Set the confidence threshold, compare the confidence of all detection regions obtained in step S22 with the confidence threshold, retain the detection regions with confidence greater than or equal to the confidence threshold, and delete the detection regions with confidence less than the confidence threshold. The retained detection regions are used as the target detection regions.
[0106] Exemplarily, set the confidence threshold to 0.95. The Yolov5 model detects a total of 4 detection regions. The confidences of the 1st to 4th detection regions are 0.85, 0.9, 0.95, and 0.78 respectively. Retain the 3rd detection region and use the 3rd detection region as the target detection region.
[0107] S24: Extract the target region coordinates of the target detection region.
[0108] The target detection region is a rectangle, and the target detection region has a total of four positioning points: the upper left corner positioning point, the lower left corner positioning point, the upper right corner positioning point, and the lower right corner positioning point.
[0109] Extract the coordinates of the upper left corner positioning point and the lower right corner positioning point of the target detection area.
[0110] Exemplarily, the coordinates of the upper left corner positioning point of the target detection area are (10, 26), the coordinates of the lower left corner positioning point are (30, 26), the coordinates of the upper right corner positioning point are (10, 76), and the coordinates of the lower right corner positioning point are (30, 76). Extract the upper left corner positioning point coordinates (10, 26) and the lower right corner positioning point coordinates (30, 76).
[0111] S25: Extract the target area image according to the target area coordinates.
[0112] Extract the image composed of the upper left corner coordinates of the target detection area and the lower right corner coordinates of the target detection area, and use this image as the target area image. The length of the target area image is the difference between the abscissas of the lower right corner positioning point and the upper left corner positioning point, and the width of the target area image is the difference between the ordinates of the lower right corner positioning point and the upper left corner positioning point.
[0113] Exemplarily, the coordinates of the upper left corner positioning point are (10, 26), and the coordinates of the lower right corner positioning point are (30, 76). Use the image within the abscissa range of 10 - 30 and the ordinate range of 26 - 76 as the target area image, and the length of the target area image is 20 and the width is 50.
[0114] The method for extracting the target area features of the input image in the embodiments of the present application filters out the target detection areas whose image category is the target category and whose confidence level is greater than or equal to the confidence level threshold by extracting the image category and confidence level of the detection area. By filtering the image category and confidence level, it can be ensured that the selected target detection areas can be used for subsequent fusion. Extract the target area coordinates of the target detection area, and extract the target area image according to the target area coordinates. The obtained image features include the target area coordinates and the target area image. The accuracy of the image features after screening is relatively high and can correspond to the text features, so as to ensure that the answer text obtained by fusing the image features and the text features has a relatively high accuracy.
[0115] In one embodiment, input the target area features into a trained convolutional neural network to obtain average image features. Refer to Figure 3 , Figure 3 is the flow schematic diagram of obtaining the average image features through a trained convolutional neural network in this application, including:
[0116] S31: Input the target area features into a trained convolutional neural network to obtain image features.
[0117] Input the target region coordinates and the target region image included in the target region features into the trained convolutional neural network. The trained convolutional neural network can use the trained fast-rcnn convolutional neural network, or other convolutional neural networks. In this embodiment, the fast-rcnn convolutional neural network is taken as an example.
[0118] Fast-rcnn first uses the hidden layer to extract the feature map of the target region image. The hidden layer includes: a convolutional layer, a Relu activation layer, and a pooling layer. This feature map is shared by the subsequent candidate region network layer and the fully connected layer.
[0119] The candidate region network layer is used to generate candidate regions. The candidate region network layer determines whether the anchor is positive or negative through the softmax function, and then uses bounding box regression to correct the anchor to obtain the accurate candidate target position.
[0120] The pooling layer collects the input feature map and the candidate target position, extracts the feature map of the candidate target position according to the feature map and the candidate target position, and inputs the feature map of the candidate target position into the fully connected layer to judge the target category.
[0121] Obtain the final target region coordinates through bounding box regression according to the target region coordinates, and obtain the image features according to the final target region coordinates.
[0122] Calculate the image features according to the following formula:
[0123]
[0124] Where, D1 and D2 are linear mappings, LN represents layer normalization, m is the serial number of the image feature, 1 ≤ m ≤ M, and M is the total number of image features. is the m-th target region image. is the m-th target region coordinate. is the m-th image feature.
[0125] S32: Calculate the average value of the image features to obtain the average image feature.
[0126] Calculate the average image feature through the following formula:
[0127]
[0128] Where, mean represents the averaging operation, and x obj is the average image feature.
[0129] In the embodiment of the present application, an average image feature is obtained through a trained convolutional neural network. The target region feature is input into the trained convolutional neural network to obtain an image feature. The average value of the image feature is calculated to obtain the average image feature. Since the convolutional neural network used has been trained, the average image feature can be obtained quickly, and the operation efficiency is relatively high. By calculating the average value of the image feature, the error can be reduced, and the obtained average image feature corresponds to the average text feature.
[0130] In one embodiment, the word vectors of the labeled question sentence are extracted through a trained text detection model to obtain an average text feature. Refer to Figure 4 , Figure 4 FIG. is a schematic flowchart of a method for obtaining an average text feature by using a trained text detection model according to the present application, including:
[0131] S51: Extract the question words and related nouns of the labeled question sentence through the trained text detection model.
[0132] Use the trained text detection model to extract the part-of-speech of all words in the labeled question sentence, retain the question words and related nouns, and remove other words such as adverbs, verbs, and adjectives.
[0133] Exemplarily, the labeled question sentence is "What does the tall giraffe like to eat?", and the trained text detection model is used to extract the question words and related nouns of the labeled question sentence, obtaining "What does the giraffe eat?".
[0134] S52: Convert the question words and the related nouns into word vectors.
[0135] Obtain the glove word vectors of the question words and related nouns. The dimension of the word vectors can be 300 dimensions, or other values.
[0136] S53: Calculate the average value of the word vectors to obtain the average text feature.
[0137] The formula for calculating the average text feature is as follows:
[0138] x text = mean(v i );
[0139] where mean represents taking the average value, v i is the i-th word vector, and x text is the average text feature.
[0140] The method for obtaining the average text feature using a trained text detection model provided by the embodiment of the present application extracts the query words and related nouns of the annotated question sentence through the trained text detection model, converts the query words and related nouns into word vectors, calculates the average value of the word vectors, and obtains the average text feature. Calculating the average text feature only using the word vectors of the query words and related nouns can eliminate the influence of adverbs, verbs, and adjectives, and the obtained text feature is relatively accurate. In addition, since only the query words and related nouns need to be converted into word vectors, this method has high operation efficiency.
[0141] Referring to Figure 5 , which is a structural schematic block diagram of a question-answering device based on dual-modal feature fusion according to the present application. The device includes:
[0142] An image and question sentence acquisition module 10, a target region feature extraction module 20, an average image feature extraction module 30, a part-of-speech tagging module 40, a text detection module 50, a feature fusion module 60, and a feature decoding module 70.
[0143] The image and question sentence acquisition module 10 is configured to acquire an input image and an input question sentence;
[0144] The target region feature extraction module 20 is configured to extract the target region feature of the input image through a trained image detection model;
[0145] The average image feature extraction module 30 is configured to input the target region feature into a trained convolutional neural network to obtain an average image feature;
[0146] The part-of-speech tagging module 40 is configured to perform part-of-speech tagging on the input question sentence to obtain an annotated question sentence;
[0147] The text detection module 50 is configured to extract the word vectors of the annotated question sentence through a trained text detection model to obtain an average text feature;
[0148] The feature fusion module 60 is configured to perform feature fusion on the average image feature and the average text feature to obtain a fusion feature;
[0149] The feature decoding module 70 is configured to input the fusion feature into a long short-term memory neural network for decoding to obtain an answer text.
[0150] The question-answering device based on dual-modal feature fusion provided by the embodiment of the present application includes an image and question acquisition module, a target region feature extraction module, an average image feature extraction module, a part-of-speech tagging module, a text detection module, a feature fusion module, and a feature decoding module. The image and question acquisition module is used to acquire an input image and an input question. The target region feature extraction module is used to extract the target region features of the input image through a trained image detection model. The average image feature extraction module is used to input the target region features into a trained convolutional neural network to obtain average image features. The part-of-speech tagging module is used to perform part-of-speech tagging on the input question to obtain a tagged question. The text detection module is used to extract the word vectors of the tagged question through a trained text detection model to obtain average text features. The feature fusion module is used to perform feature fusion on the average image features and the average text features to obtain fusion features. The feature decoding module is used to input the fusion features into a long short-term memory neural network for decoding to obtain an answer text. The question-answering device based on dual-modal feature fusion is used to implement the question-answering method based on dual-modal feature fusion.
[0151] Referring to Figure 6 , an embodiment of the present application further provides a computer device, which may be a server, and its internal structure may be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. It is characterized in that the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store average image features, average text features, etc. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the question-answering method based on dual-modal feature fusion.
[0152] Those skilled in the art can understand that Figure 6 the structure shown in
[0153] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
[0154] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. It is characterized in that any reference to a memory, storage, database or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0155] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article or method including that element.
[0156] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A question-answering method based on dual-modal feature fusion, characterized in that Including: Obtain an input image and an input question; Extract the target region features of the input image through a trained image detection model, and input the target region features into a trained convolutional neural network to obtain average image features; Perform part-of-speech tagging on the input question to obtain a tagged question; Extract the word vectors of the tagged question through a trained text detection model to obtain average text features; Perform feature fusion on the average image features and the average text features to obtain fusion features; input the fusion features into a long short-term memory neural network for decoding to obtain an answer text; The target region features include target region coordinates and a target region image; The extracting the target region features of the input image through a trained image detection model includes: Obtain the detection region of the input image; Extract the image category of the detection region and extract the confidence of the detection region; Filter out target detection regions whose image category is the target category and whose confidence is greater than or equal to a confidence threshold; Extract the target region coordinates of the target detection region; Extract the target region image according to the target region coordinates, and the target region image includes the color, contour and texture features of the target.
2. The question-answering method based on dual-modal feature fusion according to claim 1, wherein The inputting the target region features into a trained convolutional neural network to obtain average image features includes: Input the target region features into a trained convolutional neural network to obtain image features; Calculate the average value of the image features to obtain average image features.
3. The question answering method based on dual-modal feature fusion according to claim 1, wherein The performing part-of-speech tagging on the input question to obtain a tagged question includes: Use a part-of-speech tagger to split the input question to obtain split words; Perform part-of-speech tagging on all the split words to obtain a tagged question.
4. The question-answering method based on dual-modal feature fusion according to claim 1, wherein The extracting the word vectors of the tagged question through a trained text detection model to obtain average text features includes: Extract the query words and related nouns of the tagged question through the trained text detection model; Convert the query words and the related nouns into word vectors; Calculate the average value of the word vectors to obtain the average text features.
5. The question-answering method based on dual-modal feature fusion according to claim 1, wherein The performing feature fusion on the average image features and the average text features to obtain fusion features includes: Perform feature fusion on the average image features and the average text features through multi-modal Tucker fusion, multi-modal bilinear fusion or linear fusion to obtain the fusion features.
6. The question-answering method based on bimodal feature fusion according to claim 3, wherein The trained convolutional neural network includes: An input layer for preprocessing the target region features to obtain preprocessed features; A hidden layer for performing convolution, activation and pooling on the preprocessed features to obtain a hidden layer output; A fully connected layer for integrating the hidden layer output to obtain integrated features; An output layer for classifying the integrated features to obtain the image features.
7. A question-answering device based on dual-modal feature fusion, characterized in that, Including: An image and question acquisition module, a target region feature extraction module, an average image feature extraction module, a part-of-speech tagging module, a text detection module, a feature fusion module and a feature decoding module; The image and question acquisition module is used to acquire an input image and an input question; The target region feature extraction module is used to extract the target region features of the input image through a trained image detection model; The average image feature extraction module is used to input the target region features into a trained convolutional neural network to obtain average image features; The part-of-speech tagging module is used to perform part-of-speech tagging on the input question to obtain a tagged question; The text detection module is used to extract the word vectors of the tagged question through a trained text detection model to obtain average text features; The feature fusion module is used to fuse the average image features and the average text features to obtain fused features; The feature decoding module is used to input the fused features into a long short-term memory neural network for decoding to obtain an answer text; The target region features include target region coordinates and target region images; The extracting the target region features of the input image through a trained image detection model includes: Obtaining the detection region of the input image; Extracting the image category of the detection region and extracting the confidence of the detection region; Filtering out the target detection regions whose image category is the target category and whose confidence is greater than or equal to the confidence threshold; Extracting the target region coordinates of the target detection region; Extracting the target region image according to the target region coordinates, and the target region image includes the color, contour and texture features of the target.
8. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the question-answering method based on dual-modal feature fusion according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the question-answering method based on dual-modal feature fusion according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image region feature extraction method, device and equipment and readable medium
CN109635821A
Visual question answering method of dynamic memory network model based on multi-attention mechanism
CN113886626A
Intelligent question answering method based on machine reading understanding and common question answering model
CN114357127A