Question and Answer Method and Device
By extracting and analyzing the multi-layer features of the target image and combining text features, the problem that existing question-and-answer models cannot effectively process image input is solved, achieving higher answer accuracy and matching.
Patent Information
- Application Number
- CN202210412205.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing question and answer models can only support input questions in text or pronunciation, and the answers determined through similarity calculations are less accurate and cannot effectively match the questions and image content.
By extracting the first image features and the second image features of the target image and performing attention analysis on the first image features, the attention coefficient is obtained, the third image features are determined by combining the second image features and the attention coefficient, and finally the answer to the target problem is determined based on the text features and the third image features.
It improves the degree of matching between the target answer and the target question and the target image, enhances the accuracy of the answer, and can more accurately understand and match the image content.
Smart Images

Figure CN114722178B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and in particular, to a question-answering method. This application also relates to a question-answering device, a computing device, and a computer-readable storage medium. Background Art
[0002] Artificial Intelligence (AI) refers to the ability of an engineered (i.e., designed and manufactured) system to perceive the environment, as well as the ability to acquire, process, apply, and represent knowledge. Natural Language Processing (NLP) is an important direction in the field of computer science and artificial intelligence, which refers to the processing of information such as the form, sound, and meaning of natural language by a computer, that is, the operations and processing of the input, output, recognition, analysis, understanding, generation, etc. of words, phrases, sentences, and texts. It studies various theories and methods that can realize effective communication between humans and computers in natural language. The specific manifestations of natural language processing include machine translation, text summarization, text classification, text proofreading, information extraction, speech synthesis, speech recognition, intelligent question answering, etc.
[0003] With the development of Internet technology, it is very common to use a question-answering model to determine the answer to a question. However, in the existing question-answering models, the input questions are usually in the form of pure text. After the question is input into the question-answering model, the text features of the input question are extracted, and then the similarity between the text features and the sentences in the knowledge base is calculated, and the sentence with a higher similarity is determined as the answer to the question. Sometimes, the input question can also be in the form of speech, but usually, the speech is first converted into text, and then the answer to the question is determined in the same way.
[0004] It can be seen that the existing question-answering methods only support inputting questions into the question-answering model in the form of text or speech, with relatively large limitations. Moreover, the sentences determined by similarity calculation only indicate that they are relatively similar to the question, which may be the source of the question or a similar question, and are not necessarily the answer to the question. Therefore, the accuracy of the answers determined by the above methods is relatively low. Summary of the Invention
[0005] In view of this, embodiments of this application provide a question-answering method to solve the technical defects existing in the prior art. Embodiments of this application also provide a question-answering device, a computing device, and a computer-readable storage medium.
[0006] According to the first aspect of the embodiments of this application, a question-answering method is provided, including:
[0007] Extract the first image feature, the second image feature of the target image, and the text feature of the target question, where the dimension of the first image feature is lower than that of the second image feature;
[0008] Perform attention analysis on the first image feature to obtain an attention coefficient, and determine a third image feature according to the second image feature and the attention coefficient;
[0009] Determine the target answer of the target question based on the text feature and the third image feature.
[0010] According to the second aspect of the embodiments of the present application, there is provided a question-answering device, including:
[0011] An extraction module configured to extract the first image feature, the second image feature of the target image, and the text feature of the target question, where the dimension of the first image feature is lower than that of the second image feature;
[0012] A first determination module configured to perform attention analysis on the first image feature to obtain an attention coefficient, and determine a third image feature according to the second image feature and the attention coefficient;
[0013] A second determination module configured to determine the target answer of the target question based on the text feature and the third image feature.
[0014] According to the third aspect of the embodiments of the present application, there is provided a computing device, including:
[0015] A memory and a processor;
[0016] The memory is used to store computer-executable instructions, and when the processor executes the computer-executable instructions, the steps of the question-answering method are implemented.
[0017] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing computer-executable instructions, and when the instructions are executed by a processor, the steps of the question-answering method are implemented.
[0018] According to the fifth aspect of the embodiments of the present application, there is provided a chip storing a computer program, and when the computer program is executed by the chip, the steps of the question-answering method are implemented.
[0019] The question-answering method provided by this application extracts the first image feature, the second image feature of the target image, and the text feature of the target question, and then performs attention analysis on the first image feature to obtain attention coefficients. Since the dimension of the first image feature is lower than that of the second image feature, the first image feature can include more detailed information of the target image. Therefore, the attention coefficients determined based on the first image feature will be more accurate. Then, according to the second image feature and the attention coefficients, a third image feature is determined. The third image feature carries the attention coefficients, that is, the third image feature carries the information about the degree of attention to each content in the target image. Then, based on the text feature and the third image feature, the target answer to the target question is determined, that is, the question and the image are combined, and the target answer is determined according to the degree of attention to different contents in the image, improving the matching degree of the target answer with the target question and the target image, and further improving the accuracy of the target answer. Description of the Drawings
[0020] Figure 1 Shows a schematic structural diagram of a question-answering model provided by an embodiment of the present application;
[0021] Figure 2 Shows a flowchart of a question-answering method provided by an embodiment of the present application;
[0022] Figure 3 Shows a flowchart of a method for extracting image features in a question-answering method provided by an embodiment of the present application;
[0023] Figure 4 Shows a flowchart of a method for determining attention coefficients in a question-answering method provided by an embodiment of the present application;
[0024] Figure 5 Shows a flowchart of a method for determining a third image feature in a question-answering method provided by an embodiment of the present application;
[0025] Figure 6 Shows a flowchart of a method for determining a target answer in a question-answering method provided by an embodiment of the present application;
[0026] Figure 7 Shows a flowchart of a method for determining a fusion feature in a question-answering method provided by an embodiment of the present application;
[0027] Figure 8 Shows a flowchart of another method for determining a fusion feature in a question-answering method provided by an embodiment of the present application;
[0028] Figure 9 Shows a flowchart of another method for determining a target answer in a question-answering method provided by an embodiment of the present application;
[0029] Figure 10 shows a flowchart of another question - answering method provided according to an embodiment of the present application;
[0030] Figure 11 shows a flowchart of a training method of a question - answering model used in a question - answering method provided according to an embodiment of the present application;
[0031] Figure 12 shows a flowchart of yet another question - answering method provided according to an embodiment of the present application;
[0032] Figure 13 shows a processing flowchart of a question - answering method provided according to an embodiment of the present application;
[0033] Figure 14 shows a schematic structural diagram of a question - answering device provided according to an embodiment of the present application;
[0034] Figure 15 shows a structural block diagram of a computing device provided according to an embodiment of the present application. Detailed implementation manners
[0035] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0036] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.
[0037] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first can also be referred to as the second, and similarly, the second can also be referred to as the first.
[0038] First, the noun terms related to one or more embodiments of the present application are explained.
[0039] Multi-modal learning: This solution is implemented based on the idea of multi-modal learning. A modality refers to the characteristic distribution of a research problem or dataset that contains various different manifestations such as vision, touch, hearing, etc., and is described as multi-modal.
[0040] TransFormer: A self-attention mechanism model. This solution uses this model for the interaction between visual features (image features) and text features.
[0041] Bert: A natural language processing model that can capture the context information of the input statement. In this solution, this model is used for feature extraction of the target question.
[0042] LSTM (Long Short-Term Memory): A natural language processing model. In this solution, this model is used to extract the third image feature with attention of the target image.
[0043] ResNet: A computer vision model, which is a residual network. This solution uses this model to extract the image features of the target image.
[0044] Image-text description: Input an image and obtain the corresponding Chinese description statement related to the image.
[0045] First image feature: The vectorized representation of the image with lower dimension obtained by feature extraction of the target image.
[0046] Second image feature: The vectorized representation of the image with higher dimension obtained by feature extraction of the target image.
[0047] Third image feature: The image feature carrying attention coefficients, which integrates the information of the degree of attention to each content in the target image.
[0048] Text feature: The vectorized representation of the target question obtained by feature extraction of the target question.
[0049] Image feature sequence: The image features after serialization processing.
[0050] Self-attention mechanism: The self-attention mechanism actually enables the model to notice the correlations between different parts of the entire input.
[0051] Next, the application scenarios of the question-and-answer method provided in the embodiments of this application will be described.
[0052] In the current question-and-answer model, the user inputs through voice / text. If it is voice input, the voice is recognized and converted into text form for input. The similarity between the input text and the statements in the knowledge base is calculated. The calculation of similarity can be divided into two forms. One is based on the physical form of the input text and statements, such as the edit distance. The other is based on the vector form. The input text is encoded to obtain a text vector, and the statements in the knowledge base are encoded in the same way to obtain a sentence vector. The cosine similarity between the text vector and the sentence vector is calculated. This cosine similarity value can be the cosine distance. The closer the cosine similarity is to 0, the more similar they are. There are two ways to generate sentence vectors. One is the discrete one-hot form, which can be encoded through TF-IDF (term frequency–inverse document frequency). The other is the distributed representation, which uses continuous dense vectors to replace sparse vectors and maps the semantic features of the text to a high-dimensional space.
[0053] However, in the current question-and-answer scenario, the input form of the question-and-answer model can only be text, which has relatively large limitations. Moreover, the method of encoding the input text is relatively single. If faced with the situation of determining the answer to a question from an image, the output target answer may only be determined based on the text, with a low correlation with the content in the image. Additionally, the focus on the content in the image is not focused, which may lead to inaccurate determination of the target answer.
[0054] Therefore, Figure 1 The structural schematic diagram of a question-and-answer model provided by an embodiment of the present application is shown. Next, in combination with Figure 1 a brief description of the question-and-answer method provided by the embodiments of the present application is given.
[0055] The user inputs a target image and a target question through the terminal 104. The terminal 104 inputs the target image and the target question into the feature extraction unit 102-1. The target image is subjected to feature extraction to obtain a second image feature and a first image feature with a dimension lower than the second image feature. Moreover, the target question is subjected to feature extraction to obtain a text feature. Then, the first image feature is input into the attention processing unit 102-2 to perform attention analysis on the first image feature to obtain an attention coefficient. The attention coefficient and the second image feature are input into the attention feature extraction unit 102-3, and a third image feature can be obtained. Then, the third image feature and the text feature are input into the determination unit 102-4 to output the target answer to the target question.
[0056] This method determines the attention coefficient using the first image feature that includes more detailed information of the target image, which can improve the accuracy of the determined attention coefficient. Then, based on the second image feature and the attention coefficient, a third image feature carrying the degree of attention to each content in the target image is determined. Subsequently, based on the text feature and this third image feature, the target answer to the target question is determined. That is, the question and the image are combined, and the target answer is determined according to the degree of attention to different contents in the image, improving the matching degree of the target answer with the target question and the target image, and thus improving the accuracy of the target answer.
[0057] In this application, a question-answering method is provided. This application also relates to a question-answering device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0058] Figure 2 The flowchart of a question-answering method according to an embodiment of this application is shown, which specifically may include the following steps:
[0059] Step 202: Extract the first image feature, the second image feature of the target image, and the text feature of the target question, where the dimension of the first image feature is lower than that of the second image feature.
[0060] As an example, both the first image feature and the second image feature are vectorized representations of the target image. The difference is that the dimension of the first image feature is lower than that of the second image feature. Then, the first image feature can include more detailed information of the target image and is used to represent the target image more accurately. The text feature is a vectorized representation of the target question and is used to represent the target question. The target question can be a question based on the target image or any question, and the embodiments of this application do not make limitations in this regard.
[0061] The core of the embodiments of this application lies in determining the answer to the question based on multimodal information, such as text and image. And it is difficult to determine the answer solely based on the image and text. However, by extracting the image feature and the text feature, the target image and the target question can be represented more accurately. Then, based on the image feature and the text feature to determine the answer, a more accurate target answer can be obtained. Therefore, this solution needs to first extract the image feature of the target image and the text feature of the target question. And since features with different dimensions carry different amounts of information, it is more appropriate to use the feature with more information when determining the attention coefficient. Therefore, this solution can extract the first image feature and the second image feature of the target image, and the dimension of the first image feature is lower than that of the second image feature.
[0062] In some embodiments, the first image feature and the second image feature of the target image can be extracted through an image feature extraction model for processing images, and the text feature of the target question can be extracted through a text feature extraction model for processing text.
[0063] As an example, if the image feature extraction model includes at least two feature extraction layers, and the process of feature extraction by the feature extraction layer is to perform dimensionality increase processing on the features, then the features output by the last feature extraction layer are used as the second image features, or the features output by the output layer of the image extraction model are used as the second image features, and the features output by the first feature extraction layer are used as the first image features. In this way, it can be ensured that the dimension of the obtained first image features is lower than that of the second image features.
[0064] As an example, taking the text feature extraction model as the Bert model, the target problem can be feature-extracted through the Bert model, and the [CLS] feature of Bert is extracted as the text feature. This text feature can be a feature with a dimension of [1, 1, 21128].
[0065] Furthermore, since the dimension of the text features extracted by the text feature extraction model may be relatively high and the information of the target problem included is relatively small, therefore, the extracted text features can be processed by 1x1 convolution for dimensionality reduction to obtain features with a dimension of [1, 1, 20992], and then the features with a dimension of [1, 1, 20992] are reshaped to obtain features with a dimension of [16, 16, 82], and the text features with a dimension of [16, 16, 82] are used subsequently.
[0066] Furthermore, the features with a dimension of [16, 16, 82] can also be input into 1x1 convolution for channel transformation processing to obtain text features with a dimension of [16, 16, 500], and the features with a dimension of [16, 16, 500] are used when using the text features subsequently, so as to obtain more information of the target problem.
[0067] In this step, by extracting the first image features and the second image features of the target image, and extracting the text features of the target problem, it prepares for the features required to determine the target answer of the target problem subsequently, so as to facilitate subsequent use.
[0068] Step 204: Perform attention analysis on the first image features to obtain attention coefficients, and determine the third image features according to the second image features and the attention coefficients.
[0069] As an example, the attention coefficient can be understood as the degree of attention to the image content in the target image. The higher the attention coefficient, the higher the importance of the content, and the lower the attention coefficient, the lower the importance of the content. The third image feature is an image feature carrying the attention coefficient, that is, the third image feature includes the information of the degree of attention to each content in the target image.
[0070] In some embodiments of the present application, since the target image may include multiple contents that need attention, some are more important and some may be less important. In order to make the target answer focus on key points of the target image, attention analysis can be performed on the second image feature to determine the attention coefficient, so that the attention degrees to different contents in the target image can be distinguished. After determining the attention coefficient, since it is necessary to combine the image feature and the text feature when determining the target answer subsequently, the attention coefficient can be incorporated into the second image feature, that is, the third image feature is determined based on the second image feature and the attention coefficient.
[0071] In some embodiments, the second image feature can be input into a model including an attention mechanism. By performing attention calculation on the feature values of pixel points in the second image feature, the attention coefficient can be obtained. Then, the attention coefficient is multiplied point by point with the second image feature points, and the third image feature can be obtained.
[0072] In this step, performing attention analysis on the first image feature including more image detail information can obtain a more accurate attention coefficient. The third image feature is determined based on the attention coefficient and the second image feature. Then, the third image feature carries the information of the attention degrees to different contents in the target image. When determining the target answer subsequently, different contents in the target image will not be treated equally, and different contents can be focused on targeted. Then, the correlation degree between the determined target answer and the target image will be relatively high, that is, the accuracy of the target answer is higher.
[0073] Step 206: Determine the target answer of the target question based on the text feature and the third image feature.
[0074] In some embodiments, the text feature and the third image feature can be added to obtain a fusion feature, and then the target answer is determined based on the fusion feature.
[0075] Exemplarily, pooling processing is performed on the fusion feature to obtain a pooled feature. Normalization processing is performed on the pooled feature, and the probability that each predicted answer is the target answer can be obtained. The predicted answer with the maximum probability is determined as the target answer of the target question.
[0076] The question-answering method provided by the embodiment of the present application extracts the first image feature, the second image feature of the target image, and the text feature of the target question, and then performs attention analysis on the first image feature to obtain an attention coefficient. Since the dimension of the first image feature is lower than that of the second image feature, the first image feature can include more detailed information of the target image. Therefore, the attention coefficient determined based on the first image feature will be more accurate. Then, according to the second image feature and the attention coefficient, a third image feature is determined. The third image feature carries the attention coefficient, that is, the third image feature carries the information of the degree of attention to each content in the target image. Then, based on the text feature and the third image feature, the target answer to the target question is determined, that is, the question and the image are combined, and the target answer is determined according to the degree of attention to different contents in the target image, improving the matching degree of the target answer with the target question and the target image, and further improving the accuracy of the target answer.
[0077] Figure 3 The flowchart of the method for extracting image features in a question-answering method provided by an embodiment of the present application is shown, which may specifically include the following steps:
[0078] Step 302: Input the target image into the residual network for feature extraction, where the residual network includes at least two feature extraction layers.
[0079] Among them, the residual network can be any network with image feature extraction function. For example, the residual network can be ResNet. Specifically, the residual network can be ResNet101. Since the characteristic of the ResNet network is that it can ensure the accuracy while ensuring the performance, using the ResNet network to extract image features can improve the accuracy of the extracted image features.
[0080] In some embodiments, the residual network may include at least two feature extraction layers, and the feature extraction in the feature extraction layer is actually a dimensionality increase process for the features input to the feature extraction layer, that is, every time passing through a feature extraction layer, the dimension of the output feature increases once, that is, the dimension of the image feature output by the first feature extraction layer is the lowest, and the dimension of the image feature output by the last feature extraction layer is the highest.
[0081] Taking the residual network as the ResNet101 model as an example, assume that the ResNet101 model includes 4 feature extraction layers. The dimension of the target image input to the ResNet101 model can be [256, 256, 3]. Feature extraction is performed in the first feature extraction layer (stage1) of the ResNet101 model, and the dimension of the extracted image features can be [128, 128, 256]. Feature extraction is performed in the fourth feature extraction layer (stage4) of the ResNet101 model, and the dimension of the extracted image features can be [16, 16, 1024].
[0082] Step 304: Determine the features output by the last feature extraction layer as the second image features.
[0083] In some embodiments, the last feature extraction layer may be the output layer of the residual network. The dimension of the features output by the output layer is not too low and is more suitable for representing the target image. Therefore, the features output by the last feature layer can be determined as the second image features.
[0084] Continuing with the above example, the image features with the dimension of [16, 16, 1024] output by the fourth feature extraction layer (stage4) of the ResNet101 model can be determined as the second image features.
[0085] Step 306: Determine the features output by any feature extraction layer other than the last feature extraction layer as the first image features.
[0086] In some embodiments, since the dimensions of the image features output by other feature extraction layers are all lower than the dimensions of the image features output by the last feature extraction layer, the image features output by other feature extraction layers include more image detail information. Therefore, any one of the image features output by other feature extraction layers can be determined as the first image features.
[0087] As an example, if the image features output by the first feature extraction layer are determined as the first image features, more image detail information about the target image can be obtained, which is convenient for subsequent attention analysis and can obtain more accurate attention coefficients.
[0088] In the embodiments of the present application, by using the residual network to extract features from the target image, taking the image features output by the last feature extraction layer as the second image features, and taking the image features output by any other feature extraction layer as the first image features, more accurate second image features representing the target image and first image features including more image detail information can be obtained, which is convenient for determining more accurate attention coefficients, thereby improving the relevance between the target image and the determined target answer and improving the accuracy of the target answer.
[0089] It should be noted that the above steps 302 - 306 are a specific implementation process for extracting the first image feature and the second image feature of the target image in step 201.
[0090] Figure 4 The figure shows a flowchart of a method for determining an attention coefficient in a question - answering method provided by an embodiment of the present application, which specifically includes the following steps:
[0091] Step 402: Serialize the first image feature to obtain an image feature sequence.
[0092] In some embodiments, the serialization process may be to slice the first image feature according to the depth dimension to obtain an image feature sequence.
[0093] For example, assume that the dimension of the first image feature is [128, 128, 36], where 36 is the number of features in the depth dimension. After serializing the first image feature, the obtained image feature sequence includes 36 image features with a dimension of 128x138.
[0094] Step 404: Perform attention analysis on the image feature sequence to obtain attention coefficients.
[0095] In some embodiments, an attention analysis can be performed on the image feature sequence through a model including an attention mechanism to obtain attention coefficients.
[0096] As an example, taking the model including the attention mechanism as an LSTM network, the image feature sequence can be input into this LSTM network. In the LSTM network, calculate the attention values between each image feature in the image feature sequence and other image features, then multiple attention values corresponding to each image feature can be determined. These multiple attention values form an attention matrix. By multiplying the attention matrix with the attention conversion matrix in the LSTM network, the attention coefficient corresponding to each image feature in the image feature sequence can be obtained.
[0097] For example, taking the image feature sequence including 36 image features with a dimension of 128x138 as an example, input this image feature sequence into the LSTM model, calculate the attention values between each 128x138 image feature and other image features. Then, 36 corresponding attention coefficients can be obtained for each image feature. Multiply the 36 attention coefficients corresponding to image feature A by the attention conversion matrix to obtain the attention coefficient corresponding to image feature A. Applying the above method to each image feature in the image feature sequence, 36 attention coefficients can be obtained, that is, attention coefficients with a dimension of [1, 36] are obtained.
[0098] In the embodiments of the present application, first, the first image feature is serially processed into an image feature sequence, and then attention analysis is performed on the image feature sequence to determine the degree of attention to different image contents in the target image, which is represented by an attention coefficient. Then, different image contents in the target image can be distinguished, and when determining the target answer subsequently, not all image contents will be treated equally. Instead, the primary and secondary aspects can be clearly distinguished, so that the matching degree between the obtained target answer and the target image will be higher, and the determined target answer will be more accurate.
[0099] Further, for the convenience of calculation in this step, after extracting the first image feature of the target image, the first image feature can be dimensionally reduced to obtain a first image feature with a lower dimension, and then the operations of step 402 and step 404 above are performed on the first image feature with a lower dimension.
[0100] For example, the first image feature with dimensions [128, 128, 256] can be subjected to channel dimensional reduction processing through a convolution kernel with dimensions [1, 1, 36] to obtain a first image feature with dimensions [128, 128, 36].
[0101] Through the above channel dimensional reduction processing, the first image feature for serial processing can include more image detail information. Based on such a first image feature for attention analysis, the accuracy of the determined attention coefficient can be further improved.
[0102] It should be noted that the above step 402 - step 404 is a specific implementation manner for determining the attention coefficient in step 202.
[0103] Figure 5 The flowchart shows a method for determining a third image feature in a question - answering method provided by an embodiment of the present application, which may specifically include the following steps:
[0104] Step 502: Determine a visual feature matrix carrying attention according to the second image feature and the attention coefficient.
[0105] In some embodiments, the attention coefficient is used to represent the degree of attention to different image contents in the target image. Since the image feature and the text need to be combined subsequently to determine the target answer, and not all contents in the target image are important. Some contents may be more relevant to the question, while some contents may be basically useless. Therefore, if it is necessary to distinguish different image contents in the target image, the attention coefficient can be incorporated into the image feature. Thus, a visual feature matrix carrying attention can be determined according to the second image feature and the attention coefficient.
[0106] In some embodiments, the attention coefficient can be represented in the form of a matrix, and the second image feature is also represented in the form of a matrix. Then, the second image feature can be multiplied pointwise with the attention coefficient, that is, the two matrices are multiplied pointwise, and a visual feature matrix carrying the attention coefficient can be obtained.
[0107] For example, assume that the dimension of the second image feature is [16, 16, 36], and the dimension of the attention coefficient is [1, 36]. Multiplying the second image feature and the attention coefficient pointwise gives a visual feature matrix with a dimension of [16, 16, 36].
[0108] Step 504: Perform channel transformation on the visual feature matrix to obtain a third image feature.
[0109] In some embodiments, since the visual feature matrix needs to be calculated together with other matrices or features later, but the determined dimension of the visual feature matrix may not be convenient for subsequent calculations. Therefore, it is necessary to perform channel transformation on the visual feature matrix to obtain a third image feature that meets the calculation requirements.
[0110] As an example, the visual feature matrix can be processed by a 1x1 convolutional kernel for channel transformation to obtain a matrix with a dimension of [16, 16, x]. Then, this matrix is the third image feature. Here, x represents the number of predicted answers, which is related to the training of the model and can be considered as a parameter determined during the model training process. Then, the dimension of the features in the depth direction of the third image feature is the same as the number of predicted answers.
[0111] In the embodiments of the present application, the attention coefficient is fused into the second image feature to obtain a visual feature matrix carrying attention, and the visual feature matrix is converted into a third image feature that meets the calculation requirements through channel transformation. The third image feature can not only represent the target image but also represent the degree of attention to different image contents in the target image. Then, the target answer determined based on the third image feature has a higher matching degree with the target image, and the importance of different image contents in the target image is considered when determining the target answer, and a more accurate target answer can be obtained.
[0112] It should be noted that the above Step 502 - Step 504 is a specific implementation process for determining the third image feature in Step 204.
[0113] Figure 6 The flowchart shows a method for determining a target answer in a question - answering method provided by an embodiment of the present application, which may specifically include the following steps:
[0114] Step 602: Fuse the text feature and the third image feature to obtain a fused feature.
[0115] In a possible implementation manner of the present application, the target answer is determined based on the target question and the target image. Then, the target answer must be related to both the target question and the target image. Therefore, it is necessary to combine the target image and the target question. Then, the text feature and the third image feature can be added to obtain a fusion feature. As an example, when adding the third image feature and the text feature, it is necessary to ensure that the dimensions of the third image feature and the text feature are the same so that the addition operation can be performed. Otherwise, addition cannot be performed due to different dimensions.
[0116] Exemplarily, assume that the third image feature is a feature with a dimension of [16, 16, 500], and the text feature is also a feature with a dimension of [16, 16, 500]. The third image feature and the text feature can be added to obtain a fusion feature.
[0117] In another possible implementation manner of the present application, the text feature and the third image feature are fused to obtain a fusion feature, including: using the self-attention mechanism to fuse the text feature and the third image feature to obtain a fusion feature.
[0118] As an example, the self-attention mechanism is a variant of the attention mechanism, which reduces the dependence on external information and is more proficient in capturing the internal correlation of data or features. In this solution, the self-attention mechanism is used to capture the internal correlation of the text feature and the internal correlation of the third image feature, and then capture the correlation between the third image feature and the text feature.
[0119] In some embodiments, the text feature and the third image feature can be input into a model including the self-attention mechanism for self-attention calculation to obtain a fusion feature. As an example, the model including the self-attention mechanism can be a transformer model.
[0120] By obtaining the fusion feature of the text feature and the third image feature through the self-attention mechanism, the correlation between the text feature and the third image feature can be captured. Then, the obtained fusion feature combines the association relationship between the target image and the target question. Based on this, determining the target answer can obtain a more accurate answer.
[0121] Step 604: Determine the target answer to the target question based on the fusion feature.
[0122] In some embodiments, the fused feature can be input into a normalization processing model to obtain the target answer. Since the dimension of the feature in the depth direction in the third image feature is the same as the number of predicted answers, and the dimension of the feature in the depth direction in the text feature is the same as the number of predicted answers, the dimension of the feature in the depth direction in the obtained fused feature is also the same as the number of predicted answers. After inputting the fused feature into the normalization processing model, the probability that each predicted answer is the target answer can be obtained, and then the predicted answer with the highest probability is determined as the target answer to the target question.
[0123] In the embodiments of the present application, the text feature and the third image feature are first fused to obtain a fused feature, and then the target answer is determined according to the fused feature. The target answer is determined by combining the fused feature of the target question and the target image, and has a high matching degree with both the target question and the target image, and higher accuracy.
[0124] It should be noted that steps 602-step 604 are a specific implementation process for step 206 to determine the target answer.
[0125] Figure 7 The flowchart shows a method for determining a fused feature in a question-answering method provided by an embodiment of the present application, which may specifically include the following steps:
[0126] Step 702: Based on the third image feature and the query matrix, determine the image query feature matrix.
[0127] Among them, the query matrix can be a Q (Query) matrix, which is a preset matrix in the self-attention mechanism.
[0128] As an example, the third image matrix can be multiplied by the query matrix, and the obtained result is used as the image query feature matrix. The internal correlation relationship of the pixel points in the third image feature is incorporated into the image query feature matrix.
[0129] Step 704: Based on the third image feature and the keyword matrix, determine the image keyword feature matrix.
[0130] Among them, the keyword matrix can be a K (Key) matrix, which is a preset matrix in the self-attention mechanism.
[0131] As an example, the third image feature and the keyword matrix can be multiplied, and the obtained result is used as the image keyword feature matrix. The internal correlation relationship of the pixel points in the third image feature is incorporated into the image keyword feature matrix.
[0132] Step 706: Based on the text feature and the key-value matrix, determine the text key-value feature matrix.
[0133] Among them, the key-value matrix can be a V (Value) matrix, which is a preset matrix in the self-attention mechanism.
[0134] As an example, the text features can be multiplied by the key-value matrix, and the result obtained is used as the text key-value feature matrix. Then, the internal relationship between the characters in the text features is incorporated into the text key-value feature matrix.
[0135] Step 708: Determine the fusion feature based on the image query feature matrix, the image keyword feature matrix, and the text key-value feature matrix.
[0136] In some embodiments, the image query feature matrix, the image keyword feature matrix, and the text key-value feature matrix can be added together to obtain the fusion feature.
[0137] In the embodiments of the present application, the relationship between the features of the internal pixels of the third image feature is extracted through the Q and K matrices to obtain the image query feature matrix and the image keyword feature matrix. The relationship between the features of the internal characters of the text features is extracted through the V matrix to obtain the text key-value feature matrix. Then, these three features fuse the internal relationships of the target image and the target question. The fusion feature determined based on this can not only carry the relationship between the target image and the target question, but also carry the relationship between the pixels in the target image, the degree of attention to different image contents in the target image, and the relationship between the characters in the target question. All of these can improve the matching degree between the obtained target answer and the target question and the target image, and thus improve the accuracy of the determined target answer.
[0138] It should be noted that steps 702 - 708 are a specific implementation process for determining the fusion feature in the above step 602.
[0139] Figure 8 The flowchart shows another method for determining the fusion feature in a question-answering method provided by an embodiment of the present application, which may specifically include the following steps:
[0140] Step 802: Determine the autocorrelation coefficient matrix of the target image based on the image query feature matrix and the image keyword feature matrix.
[0141] Among them, the autocorrelation coefficient matrix can be used to represent the correlation relationship within the features of the third image feature of the target image.
[0142] In some embodiments, the image query feature matrix can be multiplied by the image keyword feature matrix, and the product is used as the autocorrelation coefficient matrix of the target image.
[0143] As an example, the transpose of the image keyword feature matrix can be multiplied by the image query feature matrix to obtain the autocorrelation coefficient matrix of the target image. Each value in the autocorrelation coefficient matrix represents the magnitude of the correlation between the corresponding two pixel points in the target image.
[0144] Exemplarily, the autocorrelation coefficient matrix can be determined by Equation (1).
[0145] A = K T * Q (1)
[0146] Where, A represents the autocorrelation coefficient matrix, K represents the image keyword feature matrix, Q represents the image query feature matrix. Taking the target image including n pixel points as an example, the obtained autocorrelation coefficient matrix A can be a ij represents the correlation between the i-th pixel point and the j-th pixel point.
[0147] Step 804: Determine the fusion feature based on the autocorrelation coefficient matrix and the text key value feature matrix.
[0148] In some embodiments, the autocorrelation coefficient matrix can be multiplied by the text key value feature matrix to obtain the fusion feature.
[0149] As an example, the fusion feature can be determined by the following Equation (2).
[0150] O = V * A (2)
[0151] Where, O represents the fusion feature, V represents the text key value feature matrix, and A represents the autocorrelation coefficient matrix.
[0152] In the embodiments of the present application, based on the image query feature matrix and the image keyword feature matrix, the autocorrelation coefficient matrix of the target image is determined, and the internal association relationship of the features in the target image can be obtained. Then, based on the autocorrelation coefficient matrix and the text key value feature matrix, the fusion feature is determined, so that the internal association relationship of the features can also be fused into the fusion feature, making the final target answer determined based on the fusion feature consider the association relationship between the features in the target image, and the accuracy is higher.
[0153] Further, before determining the fusion feature based on the autocorrelation coefficient matrix and the text key value feature matrix, it further includes: normalizing the autocorrelation coefficient matrix.
[0154] In some embodiments, in order to prevent overflow caused by too large matrix values, the autocorrelation coefficient matrix can be normalized, and then the fusion feature is determined based on the normalized autocorrelation coefficient matrix and the text key value feature matrix.
[0155] Exemplarily, the self-correlation coefficient matrix can be normalized through the Scale layer and Softmax, and the normalized self-correlation coefficient matrix can be obtained.
[0156] Exemplarily, a softmax operation or a relu operation can be performed on the self-correlation coefficient matrix to obtain the normalized self-correlation coefficient matrix.
[0157] By normalizing the self-correlation coefficient matrix, it is possible to prevent the matrix values from overflowing due to being too large, which is convenient for subsequent fusion processing.
[0158] It should be noted that steps 802 - 804 are another specific implementation manner for determining the fusion feature in the above step 708.
[0159] Figure 9 The flowchart shows another method for determining the target answer in a question-and-answer method provided by an embodiment of the present application, which may specifically include the following steps:
[0160] Step 902: Perform pooling processing on the fusion feature to obtain the pooled feature.
[0161] Among them, the pooling operation (Pooling) is a common operation in convolutional neural networks. It mimics the human visual system to perform dimensionality reduction on data, and can reduce overfitting while reducing network parameters and computational costs.
[0162] As an example, the fusion feature is a three-dimensional matrix, which can be understood as a three-dimensional image including multiple feature points. When performing pooling processing on the fusion feature, the three-dimensional image (fusion feature) can be divided into several rectangular regions.
[0163] In some embodiments, global average pooling operation can be performed on the fusion feature to obtain the pooled feature. Among them, average pooling outputs the average value of all elements in each sub-region. Exemplarily, assuming the fusion feature is a feature with a dimension of [1, 16, 16, 500], the pooled feature obtained after pooling processing is a feature with a dimension of [1, 1, 500].
[0164] For example, assuming the fusion feature is Use a 2*2 filter, with a stride of 2 for scanning, select the maximum value and output it to the next layer to obtain the pooled feature
[0165] In other embodiments, max pooling operation can be performed on the fusion feature to obtain the pooled feature. Among them, max pooling outputs the maximum value of all elements in each sub-region.
[0166] For example, assuming the fusion feature is Scan using a 2*2 filter with a stride of 2, select the maximum value and output it to the next layer to obtain the pooled features
[0167] Average pooling takes the average value in each matrix region, and can extract the information of all features in the feature map (fused features) and input it into the next layer, which can retain more information of the fused features.
[0168] Step 904: Based on the pooled features, determine the probability that each predicted answer is the target answer to the target question through normalization processing to obtain multiple probabilities.
[0169] In some embodiments, the dimension of the features in the depth direction in the pooled features is the same as the number of predicted answers. Therefore, by performing normalization processing on the pooled features, the probability corresponding to each predicted answer can be obtained, that is, the probability that each predicted answer may be the target answer.
[0170] As an example, the pooled features can be input into the softmax layer for normalization processing to obtain multiple probabilities. For example, assuming that the pooled features are features with a dimension of [1,1,500], 500 probabilities can be obtained.
[0171] Step 906: Determine the target answer to the target question based on multiple probabilities.
[0172] In some embodiments, since the probability refers to the probability that the predicted answer is the target answer, the predicted answer with the highest probability can be determined as the target answer.
[0173] In the embodiments of the present application, the fused features are first pooled to obtain the pooled features. The dimension in the depth direction of the pooled features is the same as the number of predicted answers. Then, the pooled features are normalized to obtain the probability that each predicted answer is the target answer, and the predicted answer with the highest probability is determined as the target answer, which can obtain a more accurate target answer.
[0174] It should be noted that steps 902-step 906 are a specific implementation manner for determining the target answer in the above step 604.
[0175] Figure 10 The flowchart shows another question-and-answer method provided according to an embodiment of the present application, which may specifically include the following steps:
[0176] Step 1002: Obtain the target image and the target question, and input the target image and the target question into the question-and-answer model.
[0177] Among them, the question-and-answer model includes an image feature extraction layer, a text feature extraction layer, an attention feature extraction layer, a feature fusion layer, and an output layer.
[0178] In some embodiments, the target problem may be a problem posed for the target image, or the target image may be an image associated with the target problem.
[0179] It should be noted that step 1002 may be a step executed before step 202.
[0180] Step 1004: Input the target image into the image feature extraction layer, and output the first image feature and the second image feature of the target image.
[0181] In some embodiments, the image feature extraction layer may be a ResNet101 network. The second image feature of the target image is determined through the first feature extraction layer of the ResNet101 network, and the first image feature of the target image is determined through the last feature extraction layer of the ResNet101 network.
[0182] As an example, dimensionality reduction processing may be performed on the first image feature and the second image feature, and it is ensured that the dimensions of the first image feature and the second image feature are the same for subsequent processing.
[0183] It should be noted that step 1004 is a specific implementation manner for extracting the first image feature and the second image feature in step 202.
[0184] Step 1006: Input the target problem into the text feature extraction layer, and output the text feature of the target problem.
[0185] In some embodiments, the text feature extraction layer may be a Bert model. Inputting the target problem into this Bert model can obtain the text feature of the target problem. Since the Bert model can better capture the context information of the sentence, extracting text features based on the Bert model can improve the semantic accuracy of the extracted text features.
[0186] It should be noted that step 1006 is a specific implementation manner for extracting text features in step 202.
[0187] Step 1008: Input the first image feature and the second image feature into the attention feature extraction layer, determine the attention coefficient based on the first image feature, and output the third image feature according to the attention coefficient and the second image feature.
[0188] In some embodiments, the attention feature extraction layer may include an attention processing sub-layer and an attention feature extraction sub-layer. The attention processing sub-layer may include an LSTM model. The LSTM model is used to perform attention feature extraction on the first image feature to determine the attention coefficient, that is, the output of the LSTM model is the attention coefficient. Then, the attention coefficient and the second image feature are input into the attention feature extraction sub-layer for dot product processing, and the third image feature can be obtained.
[0189] It should be noted that step 1008 is a specific implementation manner of step 204.
[0190] Step 1010: Input the text feature and the third image feature into the feature fusion layer, and output the fused feature.
[0191] In some embodiments, the feature fusion layer is used to fuse the text feature and the third image feature. The feature extraction layer has a self-attention mechanism. The text feature and the third image feature are input into the feature extraction layer, and the text feature and the third image feature are fused through the self-attention mechanism to obtain the fused feature.
[0192] Step 1012: Input the fused feature into the output layer, and output the target answer to the target question.
[0193] In some embodiments, the output layer may be a normalization processing layer. Inputting the fused feature into the normalization processing layer can obtain the target answer to the target question.
[0194] In the embodiments of the present application, after obtaining the target image and the target question, the target image is input into the image feature extraction layer to obtain the first image feature and the second image feature. The target question is input into the text image feature layer to output the text feature. Then, through the attention feature extraction layer, the third image feature carrying the attention coefficient is determined based on the first image feature and the second image feature. Then, the text feature and the third image feature are fused through the feature fusion layer to obtain the fused feature. Finally, the target answer is output according to the fused feature through the output layer. The above method combines the target image and the target question, and takes into account the differences in different image contents in the target image and the association between the features in the target image and the target question, so that a more accurate target answer can be obtained.
[0195] It should be noted that steps 1010 - 1012 are a specific implementation manner of step 206. For the specific implementation manners of the above steps 102 - 1012, reference can be made to the relevant descriptions of the above various embodiments, and details are not described herein again.
[0196] Figure 11 The flowchart of a training method of a question-answering model used in a question-answering method according to an embodiment of the present application is shown, which may specifically include the following steps:
[0197] Step 1102: Obtain sample pairs, where the sample pairs include sample questions, sample images, and answer labels.
[0198] As an example, the sample pairs can be obtained from an existing training sample set, and the existing training sample set can be obtained by collecting the Q&A tasks of the Q&A model.
[0199] Step 1104: Input the sample pairs into the Q&A model. Through the image feature extraction layer, determine the first image feature and the second image feature of the sample image. Through the text feature extraction layer, determine the text feature of the sample question.
[0200] In some embodiments, the image feature extraction layer can be a ResNet101 network. Determine the second image feature of the sample image through the first feature extraction layer of the ResNet101 network, and determine the first image feature of the sample image through the last feature extraction layer of the ResNet101 network.
[0201] As an example, dimensionality reduction processing can be performed on the first image feature and the second image feature, and ensure that the dimensions of the first image feature and the second image feature are the same for subsequent processing.
[0202] In some embodiments, the text feature extraction layer can be a Bert model. Input the sample question into the Bert model, and the text feature of the sample question can be obtained. Since the Bert model can better capture the context information of the sentence, extracting text features based on the Bert model can improve the semantic accuracy of the extracted text features.
[0203] Step 1106: Input the first image feature and the second image feature into the attention feature extraction layer. Determine the attention coefficient based on the first image feature, and output the third image feature according to the attention coefficient and the second image feature.
[0204] In some embodiments, the attention feature extraction layer can include an attention processing sublayer and an attention feature extraction sublayer. The attention processing sublayer can include an LSTM model. The attention coefficient is the output of the attention processing sublayer, and the third image feature is the output of the attention feature extraction sublayer.
[0205] As an example, before inputting the first image feature into the attention processing sublayer, the first image feature can be serialized to obtain an image feature sequence, and then the image feature sequence is input into the attention processing sublayer, that is, the input of the attention processing sublayer is the image feature sequence.
[0206] As another example, the first image feature can be input into the attention processing sub-layer. Through this attention sub-layer, the first image feature is first serialized to obtain an image feature sequence, and then attention coefficients are obtained based on the image feature sequence. That is, the input of the attention processing sub-layer is the first image feature.
[0207] Exemplarily, the specific implementation of determining the attention coefficient through the image feature sequence may include: calculating the attention values between each image feature in the image feature sequence and other image features, then multiple attention values corresponding to each image feature can be determined. These multiple attention values form an attention matrix. Multiplying the attention matrix by the attention conversion matrix in the LSTM network can obtain the attention coefficients corresponding to each image feature in the image feature sequence.
[0208] In some embodiments, the attention coefficients can be represented in the form of a matrix, and the second image feature is also represented in the form of a matrix. Then, in the attention feature extraction sub-layer, the second image feature can be multiplied by the attention coefficients, that is, multiplying the two matrices, to obtain a visual feature matrix carrying the attention coefficients. Then, channel transformation processing is performed on the visual feature matrix to obtain the third image feature.
[0209] Step 1108: Input the text feature and the third image feature into the feature fusion layer and output a fusion feature.
[0210] In some embodiments, the feature fusion layer is used to fuse the text feature and the third image feature. This feature extraction layer has a self-attention mechanism. Input the text feature and the third image feature into the feature extraction layer, and fuse the text feature and the third image feature through the self-attention mechanism to obtain a fusion feature.
[0211] Step 1110: Input the fusion feature into the output layer and output the predicted answer to the sample question.
[0212] In some embodiments, the output layer may include a pooling layer and a normalization processing layer. Input the fusion feature into the pooling layer for pooling processing to obtain a pooled feature, and then input the pooled feature into the normalization processing layer to obtain the probability corresponding to each predicted answer. Determine the predicted answer corresponding to the maximum probability as the predicted answer to the sample question.
[0213] Step 1112: Determine the loss value based on the predicted answer and the answer label, and adjust the parameters of the question-and-answer model based on the loss value. Return to the step of inputting the sample pair into the question-and-answer model until the training stop condition is met.
[0214] In some embodiments, the predicted answer and the answer label can be input into a loss function to determine a loss value. If the loss value is greater than a loss threshold, it indicates that the accuracy of the model is not sufficient. Therefore, the parameters of the question-answering model are adjusted based on the loss value, and then the question-answering model with the adjusted parameters is trained based on the sample pairs until the loss value is less than or equal to the loss threshold, at which point the training can be stopped and the current model is determined as the question-answering model.
[0215] In this way, when it is determined that the loss value is less than or equal to the loss threshold, it can be considered that the result determined based on the current parameters of the model is relatively accurate. Therefore, the training of the question-answering model can be stopped.
[0216] In other embodiments, the predicted answer and the answer label can be input into a loss function to determine a loss value. If the loss value is greater than the loss threshold, the parameters of the question-answering model can be adjusted based on the loss value, and each time the model parameters are adjusted, the training count is incremented by 1. When the training count is less than a count threshold, the question-answering model with the adjusted parameters is continuously trained based on the sample pairs until the training count is greater than or equal to the count threshold. At this point, it can be considered that continuing the training is basically not helpful for improving the performance of the model. Therefore, the training can be stopped and the current model is determined as the question-answering model.
[0217] In this way, when the training count reaches a certain number, it can be considered that the parameters of the model will basically no longer change, and continuing the training is not very helpful for the performance of the model. Therefore, the training of the question-answering model can be stopped.
[0218] In the embodiments of the present application, the question-answering model is trained in the form of sample pairs of sample questions and sample images, so that the question-answering model can learn the features of the pixel points in the image and the associations between them, and can also learn the associations between the words in the question and the degree of attention to different image contents in the image. Then, the question-answering model can learn more features for determining the target answer, improving the performance of the question-answering model and the accuracy of the question-answering model in determining the answer.
[0219] The following Figure 12 and 13 further illustrates the question-answering method provided by the embodiments of the present application. Figure 12 shows a flowchart of another question-answering method provided by an embodiment of the present application. Figure 13 shows a processing flowchart of a question-answering method provided by an embodiment of the present application, which specifically may include the following steps:
[0220] Step 1202: Obtain a target question and a target image.
[0221] Step 1204: Input the target image into the ResNet101 model, extract the first image feature from the first feature extraction layer, and extract the second image feature from the last feature extraction layer.
[0222] For example, the target image is an image with dimensions [256, 256, 3]. Feature extraction is performed in stage1 of the ResNet101 model to extract the first image feature with dimensions [128, 128, 256], and feature extraction is performed in stage4 of the ResNet101 model to extract the second image feature with dimensions [16, 16, 1024].
[0223] See Figure 13 , and extract the first image feature and the second image feature of the target image.
[0224] Step 1206: Perform dimensionality reduction on the first image feature.
[0225] For example, perform channel dimensionality reduction on the first image feature with dimensions [128, 128, 256] through a convolution with dimensions [1, 1, 36] to obtain the first image feature with dimensions [128, 128, 36] after dimensionality reduction.
[0226] Step 1208: Serialize the first image feature after dimensionality reduction to obtain an image feature sequence.
[0227] For example, serialize the first image feature with dimensions [128, 128, 36] to obtain 36 image features of 128x138, which form an image feature sequence.
[0228] Step 1210: Input the image feature sequence into the LSTM model to obtain attention coefficients.
[0229] For example, input 36 image features of 128x138 into the LSTM, and after processing, obtain [1, 36] as the attention coefficients.
[0230] See Figure 13 , input the first image feature into the LSTM model to obtain attention coefficients.
[0231] Step 1212: Perform dimensionality reduction on the second image feature.
[0232] For example, perform channel dimensionality reduction on the second image feature with dimensions [16, 16, 1024] through a convolution with dimensions [1, 1, 36] to obtain the second image feature with dimensions [16, 16, 36] after dimensionality reduction.
[0233] Step 1214: Multiply the second image feature after dimensionality reduction by the attention coefficients to obtain a visual feature matrix carrying attention.
[0234] For example, multiplying the second image feature after dimensionality reduction processing by the attention coefficient can obtain a visual feature matrix with attention, and the dimensionality size is [16, 16, 36].
[0235] Step 1216: Perform channel transformation on the visual feature matrix to obtain a third image feature.
[0236] For example, perform channel transformation on the visual feature matrix through a 1x1 convolution to obtain a third image feature with a dimension of [16, 16, 500].
[0237] See Figure 13 , multiply the attention coefficient by the second image feature to obtain a third image feature.
[0238] Step 1218: Input the target problem into the Bert model to obtain the text feature.
[0239] For example, perform feature extraction on the target problem based on the Bert model, extract the feature of the Bert with a dimension of [1, 1, 21128], then perform a dimensionality reduction operation on the extracted feature through a 1x1 convolution to obtain a feature with a dimension of [1, 1, 20992], then perform a reshaping operation on the feature with a dimension of [1, 1, 20992] to obtain a feature with a dimension of [16, 16, 82], and finally input the feature with a dimension of [16, 16, 82] into a 1x1 convolution for channel transformation to obtain a feature with a dimension of [16, 16, 500] as the text feature.
[0240] See Figure 13 , input the target problem into the Bert model to extract the text feature of the target problem.
[0241] Step 1220: Multiply the third image feature by the Q matrix to obtain an image query feature matrix.
[0242] See Figure 13 , obtain an image query feature matrix based on the third image feature.
[0243] Step 1222: Multiply the third image feature by the K matrix to obtain an image keyword feature matrix.
[0244] See Figure 13 , obtain an image keyword feature matrix based on the third image feature.
[0245] Step 1224: Multiply the text feature by the V matrix to obtain a text key-value feature matrix.
[0246] See Figure 13 , obtain a text key-value feature matrix based on the text feature.
[0247] Step 1226: Multiply the image query feature matrix by the image keyword feature matrix to obtain the autocorrelation coefficient matrix of the target image.
[0248] See Figure 13 , and perform a fusion process on the image query feature matrix and the image keyword feature matrix to obtain the autocorrelation coefficient matrix.
[0249] Step 1228: Normalize the autocorrelation coefficient matrix.
[0250] For example, see Figure 13 , and normalize the autocorrelation coefficient matrix through Scale and Softmax operations to prevent overflow caused by excessive matrix values.
[0251] Step 1230: Multiply the text key-value feature matrix by the normalized autocorrelation coefficient matrix to obtain the fused feature.
[0252] Step 1232: Perform global average pooling on the fused feature to obtain the pooled feature.
[0253] For example, assume the fused feature is a feature with dimensions [1, 16, 16, 500]. Perform global average pooling on this fused feature to obtain a pooled feature with dimensions [1, 1, 500].
[0254] Step 1234: Input the pooled feature into the softmax layer for normalization to obtain the target answer to the target question.
[0255] See Figure 13 , fuse the text key-value feature matrix with the normalized autocorrelation coefficient matrix to obtain the fused feature, and determine the target answer to the target question based on the fused feature.
[0256] The question-answering method provided by the embodiments of the present application extracts the first image feature and the second image feature of the target image through the ResNet101 model, and obtains the attention coefficient according to the first image feature through the LSTM model. Since the dimension of the first image feature is lower than that of the second image feature, the first image feature can include more detailed information of the target image. Therefore, the attention coefficient determined based on the first image feature will be more accurate. Then, the attention coefficient is multiplied by the second image feature to obtain the third image feature carrying attention. The third image feature carries the information about the degree of attention to each content in the target image. The text feature is obtained by performing feature extraction on the third image feature through the Bert model, and then the self-attention mechanism is used to fuse the text feature and the third image feature, which can capture the internal correlation of the text feature and the internal correlation of the third image feature, and then capture the correlation between the third image feature and the text feature. The obtained fusion feature combines the association relationship between the target image and the target question. Furthermore, the target answer determined based on the fusion feature has a high matching degree with both the target question and the target image, and the accuracy is also higher.
[0257] Corresponding to the above method embodiments, the present application also provides embodiments of a question-answering device. Figure 14 The structural schematic diagram of a question-answering device provided by an embodiment of the present application is shown. As Figure 14 shown, the device may include:
[0258] An extraction module 1402, configured to extract the first image feature, the second image feature of the target image, and the text feature of the target question, where the dimension of the first image feature is lower than that of the second image feature;
[0259] A first determination module 1404, configured to perform attention analysis on the first image feature to obtain an attention coefficient, and determine a third image feature according to the second image feature and the attention coefficient;
[0260] A second determination module 1406, configured to determine the target answer of the target question based on the text feature and the third image feature.
[0261] In one or more embodiments of the present application, the first determination module 1404 is further configured to:
[0262] Perform serialization processing on the first image feature to obtain an image feature sequence;
[0263] Perform attention analysis on the image feature sequence to obtain an attention coefficient.
[0264] In one or more embodiments of the present application, the first determination module 1404 is further configured to:
[0265] Determine a visual feature matrix with attention according to the second image feature and the attention coefficient;
[0266] Perform channel transformation on the visual feature matrix to obtain a third image feature.
[0267] In one or more embodiments of the present application, the second determination module 1406 is further configured to:
[0268] Fuse the text feature and the third image feature to obtain a fused feature;
[0269] Determine the target answer to the target question based on the fused feature.
[0270] In one or more embodiments of the present application, the second determination module 1406 is further configured to:
[0271] Use the self-attention mechanism to fuse the text feature and the third image feature to obtain a fused feature.
[0272] In one or more embodiments of the present application, the second determination module 1406 is further configured to:
[0273] Determine an image query feature matrix based on the third image feature and the query matrix;
[0274] Determine an image keyword feature matrix based on the third image feature and the keyword matrix;
[0275] Determine a text key-value feature matrix based on the text feature and the key-value matrix;
[0276] Determine a fused feature based on the image query feature matrix, the image keyword feature matrix, and the text key-value feature matrix.
[0277] In one or more embodiments of the present application, the second determination module 1406 is configured to:
[0278] Determine the self-correlation coefficient matrix of the target image based on the image query feature matrix and the image keyword feature matrix;
[0279] Determine a fused feature based on the self-correlation coefficient matrix and the text key-value feature matrix.
[0280] In one or more embodiments of the present application, the second determination module 1406 is further configured to:
[0281] Perform normalization processing on the self-correlation coefficient matrix.
[0282] In one or more embodiments of the present application, the second determination module 1406 is further configured to:
[0283] Perform pooling processing on the fused feature to obtain a pooled feature;
[0284] Based on the pooled features, determine the probability that each predicted answer is the target answer to the target question through normalization processing, and obtain multiple probabilities;
[0285] Determine the target answer to the target question based on the multiple probabilities.
[0286] In one or more embodiments of the present application, the extraction module 1402 is further configured to:
[0287] Input the target image into the residual network for feature extraction, where the residual network includes at least two feature extraction layers;
[0288] Determine the features output by the last feature extraction layer as the second image features;
[0289] Determine the features output by any feature extraction layer other than the last feature extraction layer as the first image features.
[0290] In one or more embodiments of the present application, the question and answer device further includes:
[0291] An acquisition module, configured to acquire a target image and a target question, and input the target image and the target question into the question and answer model, where the question and answer model includes an image feature extraction layer, a text feature extraction layer, an attention feature extraction layer, a feature fusion layer, and an output layer;
[0292] The extraction module 1402 is further configured to:
[0293] Input the target image into the image feature extraction layer, and output the first image features and the second image features of the target image;
[0294] Input the target question into the text feature extraction layer, and output the text features of the target question;
[0295] The first determination module 1404 is further configured to:
[0296] Input the first image features and the second image features into the attention feature extraction layer, determine the attention coefficients based on the first image features, and output the third image features according to the attention coefficients and the second image features;
[0297] The second determination module 1406 is further configured to:
[0298] Input the text features and the third image features into the feature fusion layer, and output the fusion features;
[0299] Input the fusion features into the output layer, and output the target answer to the target question.
[0300] In one or more embodiments of the present application, the question and answer device further includes a training module, configured to:
[0301] Obtain sample pairs, where the sample pairs include sample questions, sample images, and answer labels;
[0302] Input the sample pairs into the question-answering model. Through the image feature extraction layer, determine the first image feature and the second image feature of the sample image. Through the text feature extraction layer, determine the text feature of the sample question;
[0303] Input the first image feature and the second image feature into the attention feature extraction layer. Based on the first image feature, determine the attention coefficient, and according to the attention coefficient and the second image feature, output the third image feature;
[0304] Input the text feature and the third image feature into the feature fusion layer, and output the fusion feature;
[0305] Input the fusion feature into the output layer, and output the predicted answer to the sample question;
[0306] Based on the predicted answer and the answer label, determine the loss value, and based on the loss value, adjust the parameters of the question-answering model. Return to execute the step of inputting the sample pairs into the question-answering model until the training stop condition is met.
[0307] The question-answering device provided by the embodiment of the present application extracts the first image feature, the second image feature of the target image, and the text feature of the target question through the extraction module, and then performs attention analysis on the first image feature through the first determination module to obtain the attention coefficient. Since the dimension of the first image feature is lower than that of the second image feature, the first image feature can include more detailed information of the target image. Therefore, the attention coefficient determined based on the first image feature will be more accurate. Then, according to the second image feature and the attention coefficient, determine the third image feature. The third image feature carries the attention coefficient, that is, the third image feature carries the information of the degree of attention to each content in the target image. Then, through the second determination module, based on the text feature and the third image feature, determine the target answer to the target question, that is, combine the question and the image, and determine the target answer according to the degree of attention to different contents in the image, improving the matching degree of the target answer with the target question and the target image, and further improving the accuracy of the target answer.
[0308] The above is a schematic solution of a question-and-answer device according to this embodiment. It should be noted that the technical solution of this question-and-answer device and the technical solution of the above question-and-answer method belong to the same concept. For the details not described in detail in the technical solution of the question-and-answer device, reference can be made to the description of the technical solution of the above question-and-answer method. In addition, each component in the device embodiment should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional division or separation limitation. The device claim defined by such a set of functional modules should be understood as mainly implementing the functional module framework of the solution through the computer program recorded in the specification, rather than understanding as an entity device mainly implementing the solution through hardware means.
[0309] Figure 15 FIG. shows a structural block diagram of a computing device 1500 according to an embodiment of the present application. The components of the computing device 1500 include, but are not limited to, a memory 1510 and a processor 1520. The processor 1520 is connected to the memory 1510 through a bus 1530, and the database 1550 is used to store data.
[0310] The computing device 1500 further includes an access device 1540, which enables the computing device 1500 to communicate via one or more networks 1560. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1540 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Controller (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.
[0311] In an embodiment of the present application, the above components of the computing device 1500 and Figure 15Other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 15 The block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0312] The computing device 1500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 1500 can also be a mobile or stationary server.
[0313] Among them, the processor 1520 is used to execute the computer-executable instructions of the question-and-answer method.
[0314] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above question-and-answer method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above question-and-answer method.
[0315] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions that, when executed by a processor, are used for the question-and-answer method.
[0316] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the above question-and-answer method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above question-and-answer method.
[0317] The computer instructions include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0318] An embodiment of the present application also provides a chip, which stores a computer program that, when executed by the chip, implements the steps of the question-and-answer method.
[0319] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0320] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0321] The preferred embodiments of this application disclosed above are only used to help explain this application. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this application. This application selects and specifically describes these embodiments in order to better explain the principle and practical application of this application, so that those skilled in the art can well understand and utilize this application. This application is only limited by the claims and their full scope and equivalents.
Claims
1. A question-and-answer method, characterized in that, Including: Obtain a target image and a target question; Extract a first image feature, a second image feature of the target image, and a text feature of the target question, where the dimension of the first image feature is lower than that of the second image feature; Perform attention analysis on the first image feature to obtain an attention coefficient, and determine a third image feature according to the second image feature and the attention coefficient; Determine a target answer to the target question based on the text feature and the third image feature.
2. The method according to claim 1, wherein The performing attention analysis on the first image feature to obtain an attention coefficient includes: Perform serialization processing on the first image feature to obtain an image feature sequence; Perform attention analysis on the image feature sequence to obtain an attention coefficient.
3. The method according to claim 1, characterized in that, The determining a third image feature according to the second image feature and the attention coefficient includes: Determine a visual feature matrix carrying attention according to the second image feature and the attention coefficient; Perform channel transformation on the visual feature matrix to obtain a third image feature.
4. The method according to claim 1, wherein The determining a target answer to the target question based on the text feature and the third image feature includes: Fuse the text feature and the third image feature to obtain a fused feature; Determine a target answer to the target question based on the fused feature.
5. The method according to claim 4, characterized in that, The fusing the text feature and the third image feature to obtain a fused feature includes: Use a self-attention mechanism to fuse the text feature and the third image feature to obtain a fused feature.
6. The method according to claim 5, wherein The using a self-attention mechanism to fuse the text feature and the third image feature to obtain a fused feature includes: Determine an image query feature matrix based on the third image feature and a query matrix; Determine an image keyword feature matrix based on the third image feature and a keyword matrix; Determine a text key-value feature matrix based on the text feature and a key-value matrix; Determine a fused feature based on the image query feature matrix, the image keyword feature matrix, and the text key-value feature matrix.
7. The method according to claim 6, wherein The determining a fused feature based on the image query feature matrix, the image keyword feature matrix, and the text key-value feature matrix includes: Determine a self-correlation coefficient matrix of the target image based on the image query feature matrix and the image keyword feature matrix; Determine a fused feature based on the self-correlation coefficient matrix and the text key-value feature matrix.
8. The method according to claim 7, wherein Before the determining a fused feature based on the self-correlation coefficient matrix and the text key-value feature matrix, it further includes: Perform normalization processing on the self-correlation coefficient matrix.
9. The method according to any one of claims 4-8, characterized in that, The determining a target answer to the target question based on the fused feature includes: Perform pooling processing on the fused feature to obtain a pooled feature; Based on the pooled feature, determine the probability that each predicted answer is the target answer to the target question through normalization processing to obtain multiple probabilities; Determine the target answer to the target question based on multiple probabilities.
10. The method according to claim 1, characterized in that, The extracting a first image feature and a second image feature of the target image includes: Input the target image into a residual network for feature extraction, where the residual network includes at least two feature extraction layers; Determine the features output by the last feature extraction layer as the second image features; Determine the features output by any feature extraction layer other than the last feature extraction layer as the first image features.
11. The method according to claim 1, wherein Before extracting the first image features, second image features of the target image and the text features of the target question, it further includes: Input the target image and the target question into a question-answering model, where the question-answering model includes an image feature extraction layer, a text feature extraction layer, an attention feature extraction layer, a feature fusion layer, and an output layer; The extraction of the first image features, second image features of the target image and the text features of the target question includes: Input the target image into the image feature extraction layer to output the first image features and second image features of the target image; Input the target question into the text feature extraction layer to output the text features of the target question; The attention analysis of the first image features to obtain attention coefficients, and determining the third image features according to the second image features and the attention coefficients includes: Input the first image features and the second image features into the attention feature extraction layer, determine the attention coefficients based on the first image features, and output the third image features according to the attention coefficients and the second image features; Based on the text features and the third image features, determining the target answer to the target question includes: Input the text features and the third image features into the feature fusion layer to output the fused features; Input the fused features into the output layer to output the target answer to the target question.
12. The method according to claim 11, wherein The question-answering model is trained by the following method: Obtain sample pairs, where the sample pairs include sample questions, sample images, and answer labels; Input the sample pairs into the question-answering model, determine the first image features and second image features of the sample image through the image feature extraction layer, and determine the text features of the sample question through the text feature extraction layer; Input the first image features and the second image features into the attention feature extraction layer, determine the attention coefficients based on the first image features, and output the third image features according to the attention coefficients and the second image features; Input the text features and the third image features into the feature fusion layer to output the fused features; Input the fused features into the output layer to output the predicted answers to the sample questions; Determine the loss value based on the predicted answers and the answer labels, and adjust the parameters of the question-answering model based on the loss value, and return to execute the step of inputting the sample pairs into the question-answering model until the training stop condition is met.
13. A question-and-answer device, characterized in that, It includes: An acquisition module configured to acquire a target image and a target question; An extraction module configured to extract the first image features, second image features of the target image and the text features of the target question, where the dimension of the first image features is lower than that of the second image features; A first determination module, configured to perform attention analysis on the first image feature to obtain an attention coefficient, and determine a third image feature according to the second image feature and the attention coefficient; A second determination module, configured to determine a target answer to the target question based on the text feature and the third image feature.
14. A computing device, characterized in that, Comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the question-answering method according to any one of claims 1 to 12.
15. A computer-readable storage medium storing computer instructions, characterized in that, When the instruction is executed by the processor, the steps of the question-answering method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Super-resolution reconstruction method based on attention mechanism
CN113706386A
Picture-based question and answer processing method and device, readable medium and electronic equipment
CN113761153A