An image description method based on multi-interaction information fusion
Through the multi-visual semantic information interaction module and the multi-modal interactive information network, the problem of incomplete understanding of semantic information in the image description model is solved, and sentences that are closer to the real description are generated, achieving a comprehensive understanding and accurate expression of image semantic information.
Patent Information
- Application Number
- CN202211194469.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-28
AI Technical Summary
In the prior art, the image description model has a relatively one-sided understanding of image semantic information and has failed to fully capture the interactive information between image visual information and the various complementary information between image visual information and text semantic information.
The multi-visual semantic information interaction module and multi-modal interactive information network are used to extract the prominent regional features of the image through the object detection model, and the multi-headed attention mechanism and the AoA mechanism are used to model the relationship between visual semantic information, and the relationship between global image fusion features and text semantic information is mined through the long-term and short-term memory network to generate the probability distribution of the output word sequence.
A more comprehensive understanding of image semantic information is achieved, the generated sentences are closer to the real description, and can accurately express image semantic structure information.
Smart Images

Figure CN115512195B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to two major fields of computer vision and natural language processing, and particularly relates to an image description method based on multi-interaction information fusion. Background Art
[0002] Image description is a task of describing the semantics such as objects, relationships, and attributes contained in an image through natural language. Image description has broad application prospects and has high practical value in assisting the life of visually impaired people, children's education, medical image analysis, etc. In the prior art, the encoder-decoder is the mainstream framework adopted by the image description task model. In this framework, the encoder uses a convolutional neural network (CNN) to encode the input image, and then the decoder using a recurrent neural network (RNN) decodes it to obtain natural language matching the input image.
[0003] In terms of capturing the interaction information between image visual information, the prior art mines the visual semantic information between objects through the attention mechanism, such as the authorized patent: CN113378919B. Although this method refines the representation of feature vectors by modeling the relationship between feature vectors, the mining of the relationship between feature vectors is not sufficient.
[0004] In terms of capturing the interaction information between image visual information and text semantic information, recent research usually adopts a long short-term memory network (LSTM) with an attention mechanism, performing selective forgetting and memory at each time step of decoding. For example, the authorized patent: CN110991515B. However, this only represents a specific perspective on decoding image visual information and text semantic information, so the model's understanding of the semantic information of the image is relatively one-sided. Summary of the Invention
[0005] Object of the Invention: Aiming at the problem that the model's understanding of the semantic information of the image is relatively one-sided pointed out in the background art, the present invention proposes an image description method based on multi-interaction information fusion, which can fully capture the interaction information between image visual information, as well as various complementary information of the interaction information between image visual information and text semantic information, and achieve a more comprehensive understanding of image semantics.
[0006] Technical Solution: The present invention discloses an image description method based on multi-interaction information fusion, including the following steps:
[0007] Step 1: Preprocess the data set and the real text description of the image;
[0008] Step 2: Extract the global image fusion features of the images in the data set;
[0009] Step 3: Use the multi-modal interaction information network to mine the relationship between the global image fusion features and the text semantic information, and obtain the context information at this time step;
[0010] Step 4: Use the linear unit of semantic decoding to decode the context information to generate the probability distribution of the output word sequence.
[0011] Further, the preprocessing in Step 1 specifically includes the following steps:
[0012] Step 1.1: Divide the data set in sequence, where 92% is divided into the training set, 4% is divided into the validation set, and the remaining 4% is divided into the test set;
[0013] Step 1.2: Convert the text of the 5 true descriptions corresponding to each picture in the data set to lowercase;
[0014] Step 1.3: Statistically count each word of the true description converted to lowercase to obtain a corpus, and the corpus is <unk>As the end flag, remove the words that appear less than 5 times in the corpus;
[0015] Step 1.4: Statistically analyze the length L of the real text description of each image, L = {L1, L2,..., L i}, and set the length of the real text description of each image to argmax(L) + 2. For those with a real text description length less than argmax(L) + 2, pad them with tokens.
[0016] Furthermore, the specific steps for the above-mentioned Step 2 to extract the global image fusion features of the images in the dataset are as follows:
[0017] Step 2.1: Use an object detection model to extract all the significant region features of the training set images, denoted as v = [v1, v2,..., v a}, where v a represents the a-th significant region feature;
[0018] Step 2.2: Perform three linear mappings on the significant region features v of the image respectively, and denote the obtained linear representations as Q, K, and V respectively. The specific formulas are as follows:
[0019] Q = vW Q + b Q
[0020] K = vW K + b K
[0021] V = vW V + b V
[0022] where W Q 、W K 、W V represent linear transformation matrices; b Q 、b K 、b V represent biases.
[0023] Step 2.3: Use the multi-visual semantic information interaction module to model the relationship between the significant region features of the image, and then obtain the global image fusion feature.
[0024] Furthermore, the above-mentioned Step 2.3 uses the multi-visual semantic information interaction module to model the relationship between the significant region features of the image, and then obtain the global image fusion feature. The specific steps are as follows:
[0025] The multi-visual semantic information interaction module consists of 3xNxR linear layers, NxR Layer Norm layers, NxR multi-head attention mechanisms, and NxR AoA layers;
[0026] Step 2.3.1: Adopt the multi-head attention mechanism to enable the features of the significant regions of the image to selectively focus on the features of other relevant regions, thereby obtaining local feature relationships. The specific formula is as follows:
[0027] f multi_head_att (Q, K, V) = Concat(head1, head2,..., head H )
[0028]
[0029] where f multi_head_att represents the multi-head attention function; Concat represents the vector concatenation operation; head j represents the j-th head attention function, which is implemented using the scaled dot-product attention function; H represents the number of heads; represents the scaling factor; Q j , K j , V j represent the linear representations of the j-th head; softmax represents the normalized exponential function;
[0030] Step 2.3.2: Use the AoA mechanism to determine the correlation between the local feature relationships and the features of the significant regions of the image, enabling the significant features of each image to selectively focus on the features of other regions that are truly relevant to it. The specific formula is as follows:
[0031]
[0032] where σ is the sigmoid activation function; represents element-wise multiplication, represents the linear transformation matrix; b e , b j represent the biases;
[0033] Step 2.3.3: Repeat Step 2.3.1 and Step 2.3.2 N times to obtain the high-level local feature relationship f AoAS ;
[0034] Step 2.3.4: Perform residual connection and normalization on the features of the significant regions of the image and the high-level local feature relationship to obtain enhanced image features. The specific formula is as follows:
[0035] v = LayerNorm(v + f AoAS (f multi_head_att , Q, K, V))
[0036] where LayerNorm is the layer normalization function;
[0037] Step 2.3.5: Repeat steps 2.3.1 to 2.3.4 for R times to generate multi-layer enhanced image features;
[0038] Step 2.3.6: Use vector concatenation operation to fuse the multi-layer enhanced image features to obtain multi-layer enhanced image fusion features. The specific formula is as follows:
[0039]
[0040] where [.,.] represents the vector concatenation operation, and v′ R represents the R-th layer enhanced image feature; represents the multi-layer enhanced image fusion feature;
[0041] Step 2.3.7: Generate global image fusion features by performing average pooling on the multi-layer enhanced image fusion features. The specific formula is as follows:
[0042]
[0043] where represents the global image fusion feature; a represents the number of channels of the multi-layer enhanced image fusion feature.
[0044] Furthermore, the multi-modal interaction information network in step 3 is composed of a single multi-head attention layer, an AoA layer, an embedding layer, and U long short-term memory networks, and specifically includes the following steps:
[0045] Step 3.1: Input the word vectors ∏ corresponding to all words in the corpus into the word embedding layer to obtain the word embedding vector W ∏ ∏ represented in one-hot encoding;
[0046] Step 3.2: Use the word embedding vector at the current time step, the global image fusion feature, and the context information at the previous time step as the inputs of U long short-term memory networks, and then obtain multiple complementary information on the interaction information between the global image fusion feature and the word embedding vector. The specific formula is as follows:
[0047]
[0048] where represents the U-th group of complementary information at the current time step; represents the U-th group of cell states at the current time step; W ∏ represents the word embedding matrix; Π t represents the input word at the current time step; represents the U-th group of context information at the previous time step; represents the U-th group of complementary information at the previous time step; represents the U-th group of cell states at the previous time step;
[0049] Step 3.3: Perform a vector concatenation operation on multiple multimodal interaction information for fusion, and map it to the same vector space through an embedding layer to generate multimodal interaction information fusion features. The specific formula is as follows:
[0050]
[0051] where p t represents the multimodal interaction information fusion feature at the current time step; [.,.] represents the vector concatenation operation, and W h represents the mapping matrix; b h represents the bias;
[0052] Step 3.4: Adopt the multi-head attention mechanism and the AoA mechanism to determine the correlation between the multimodal interaction information fusion feature and the image salient region feature, so as to obtain the context vector for generating the word sequence. The specific formula is as follows:
[0053]
[0054]
[0055]
[0056] where C t represents the context information at the current time step; W p represents the linear transformation matrix; represents the multi-head attention function; Concat represents the vector concatenation operation; head j represents the j-th head attention function, which is implemented by using the scaled dot-product attention function; H represents the number of heads; represents the scaling factor; K j , V j represent the linear representations of the j-th head; softmax represents the normalized exponential function.
[0057] Beneficial effects:
[0058] The present invention solves the problem that the current model does not comprehensively understand the image semantic information. The multi-visual semantic information interaction module fully explores the relationship between visual semantic information in the encoder part, and the multimodal interaction information network fully models the relationship between visual semantic information and text semantic information in the decoder part. Through this method, not only can words closer to the real description be generated, but also the semantic structure information of the generated sentences can more accurately express the image semantic information. Brief description of the drawings
[0059] Figure 1 This is the overall flowchart of the image description method based on multi-interaction information fusion of the present invention;
[0060] Figure 2 This is the schematic diagram of the multi-visual semantic information interaction module of the present invention;
[0061] Figure 3 This is the schematic diagram of the multi-modal interaction information network of the present invention. Detailed implementation manners
[0062] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0063] As Figure 1 shown, in order to achieve a more comprehensive understanding of the image semantics, the present invention proposes an image description method based on multi-interaction information fusion, including the following steps:
[0064] Step 1: Preprocess the data set and the real text description of the image, specifically including the following steps:
[0065] Step 1.1: Divide the data set in sequence, where 92% is divided into the training set, 4% is divided into the validation set, and the remaining 4% is divided into the test set.
[0066] Step 1.2: Convert the text of the 5 real descriptions corresponding to each picture in the data set to lowercase.
[0067] Step 1.3: Statistically analyze each word of the real description converted to lowercase to obtain a corpus, and this corpus is in <unk>As the end flag, remove the words that appear less than 5 times in the corpus.
[0068] Step 1.4: Count the length L of the true text description of each image, L = {L1, L2,..., L i}, and set the length of the true text description of each image to argmax(L) + 2. For those with a true text description length less than argmax(L) + 2, pad them with tokens.
[0069] Step 2: Extract the global image fusion features of the images in the dataset, which specifically includes the following steps:
[0070] Step 2.1: Use Faster R-CNN to extract all the salient region features of the training set images, denoted as v = [v1, v2,..., v a}.
[0071] Among them, v a represents the a-th salient region feature.
[0072] Step 2.2: Perform three linear mappings on the salient region features v of the image respectively, and denote the obtained linear representations as Q, K, and V. The specific formulas are as follows:
[0073] Q = vW Q + b Q
[0074] K = vW K + b K
[0075] V = vW V + b V
[0076] Among them, W Q 、W K 、W V represent linear transformation matrices; b Q 、b K 、b V represent biases.
[0077] Step 2.3: Use the multi-visual semantic information interaction module to model the relationship between the salient region features of the image, and then obtain the global image fusion features, which specifically includes the following steps:
[0078] As Figure 2 shown, the multi-visual semantic information interaction module consists of 3xNxR linear layers, NxR Layer Norm layers, NxR multi-head attention mechanisms, and NxR AoA layers. In this embodiment, N is taken as 6 and R is taken as 1.
[0079] Step 2.3.1: The multi-head attention mechanism is adopted to enable the features of the significant regions of the image to selectively pay attention to the features of other relevant regions, so as to obtain the local feature relationship. The specific formula is as follows:
[0080] f multi_head_att (Q, K, V) = Concat(head1, head2,..., head H )
[0081]
[0082] where f multi_head_att represents the multi-head attention function; Concat represents the vector concatenation operation; head j represents the j-th head attention function; H represents the number of heads, and H is taken as 8 in this embodiment; the scaled dot-product attention function is used to implement represents the scaling factor; Q j , K j , V j represent the linear representations of the j-th head; softmax represents the normalized exponential function.
[0083] Step 2.3.2: The AoA mechanism is used to determine the correlation between the local feature relationship and the features of the significant regions of the image, so that the significant features of each image can selectively pay attention to the features of other regions that are truly relevant to it. The specific formula is as follows:
[0084]
[0085] where σ is the sigmoid activation function; represents element-wise multiplication, represents the linear transformation matrix; b e , b j represent the biases;
[0086] Step 2.3.3: Repeat Step 2.3.1 and Step 2.3.2 N times to obtain the high-level local feature relationship f AoAS , and N is taken as 6 in this embodiment.
[0087] Step 2.3.4: Perform residual connection and normalization on the features of the significant regions of the image and the high-level local feature relationship to obtain the enhanced image features. The specific formula is as follows:
[0088] v′ = LayerNorm(v + f AoAS (f multi_head_att , Q, K, V))
[0089] where LayerNorm is the layer normalization function.
[0090] Step 2.3.5: Repeat Step 2.3.1, Step 2.3.2, Step 2.3.3, and Step 2.3.4 for R times to generate multi-layer enhanced image features. In this embodiment, R is taken as 1.
[0091] Step 2.3.6: Use vector concatenation operation to fuse the multi-layer enhanced image features to obtain multi-layer enhanced image fusion features. The specific formula is as follows:
[0092]
[0093] where [.,.] represents the vector concatenation operation, and v′ R represents the R-th layer of enhanced image features; represents the multi-layer enhanced image fusion features.
[0094] Step 2.3.7: Generate global image fusion features by performing average pooling on the multi-layer enhanced image fusion features. The specific formula is as follows:
[0095]
[0096] where represents the global image fusion features, and a represents the number of channels of the multi-layer enhanced image fusion features.
[0097] Step 3: Use the multi-modal interaction information network to mine the relationship between the global image fusion features and the text semantic information to obtain the context information at this time step. The specific steps are as follows:
[0098] As Figure 3 shown, the multi-modal interaction information network in Step 3 is composed of a single multi-head attention layer, an AoA layer, an embedding layer, and U long short-term memory networks. In this embodiment, U is taken as 3.
[0099] Step 3.1: Input the word vectors ∏ corresponding to all words in the corpus into the word embedding layer to obtain the word embedding vector W ∏ ∏ represented by one-hot encoding.
[0100] Step 3.2: Use the word embedding vector at the current time step, the global image fusion features, and the context information at the previous time step as the inputs of U long short-term memory networks, and then obtain multiple complementary information on the interaction information between the global image fusion features and the word embedding vector. The specific formula is as follows:
[0101]
[0102] where represents the U-th group of complementary information at the current time step; represents the U-th group of cell states at the current time step; W ∏ represents the word embedding matrix; ∏ t represents the input word at the current time step; represents the U-th group of context information at the previous time step; represents the U-th group of complementary information at the previous time step; represents the U-th group of cell states at the previous time step.
[0103] Step 3.3: Perform a vector concatenation operation on multiple multimodal interaction information for fusion, and map it to the same vector space through an embedding layer to generate a multimodal interaction information fusion feature. The specific formula is as follows:
[0104]
[0105] where p t represents the multimodal interaction information fusion feature at the current time step; [.,.] represents the vector concatenation operation, and W h represents the mapping matrix; b h represents the bias.
[0106] Step 3.4: Adopt a multi-head attention mechanism and an AoA mechanism to determine the correlation between the multimodal interaction information fusion feature and the image salient region feature, so as to obtain a context vector for generating a word sequence. The specific formula is as follows:
[0107]
[0108]
[0109]
[0110]
[0111]
[0112] where Ct represents the context information at the current time step; W K , W V , W p represent linear transformation matrices; b K , b V represent biases; represents the multi-head attention function; Concat represents the vector concatenation operation; head j represents the j-th head attention function; H represents the number of heads, and H takes 8 in this embodiment; it is implemented using a scaled dot-product attention function; represents the scaling factor; K j , V j represent the linear representations of the j-th head; softmax represents the normalized exponential function.
[0113] Step 4: Use a linear unit with semantic decoding to decode the context information to generate the probability distribution of the output word sequence. The specific formula is as follows:
[0114] y t = softmax(W C C t + b C )
[0115] where y t is the probability distribution of the output word sequence at the current time step; W c represents the linear transformation matrix; b C represents the bias.
[0116] To better illustrate the effectiveness of this method, an experimental verification is carried out on the image description method based on multi-interaction information fusion. The experimental environment is as follows:
[0117] Hardware configuration: NIADIA Geforce RTX 2080Ti graphics card (11G video memory).
[0118] Software configuration: Ubuntu 18.04 64-bit operating system, Python 3.6, Pytorch 1.2.0 and Torchversion 0.4.0 deep learning framework.
[0119] In this experiment, the effectiveness of the model is verified by using the mainstream evaluation metrics BLEU@N, METOR, ROUGE_L, CIDEr-D, and SPICE on the commonly used MS COCO dataset in the field of image description generation. The MS COCO dataset is divided, where 113,287 images are divided into the training set, 5,000 images are divided into the validation set, and the remaining 5,000 images are divided into the test set.
[0120] The Cross-Entropy Loss function is used to train the model. The comparison results of the model of the present invention with the Up-Down, RFNet, and AoA models are shown in Table 1. Compared with AoA, the present invention has improved by 0.3% in the evaluation metrics BLEU@3 and ROUGE_L, 0.7% in BLEU@4, 0.2% in METOR, and 1.4% in CIDEr-D.
[0121] Table 1 Comparison table of evaluation metrics after cross-entropy loss training
[0122]
[0123] The model is trained using the policy gradient algorithm SCST in reinforcement learning. The comparison results of the model in this paper with the Up-Down, RFNet, and AoA models after optimization by the SCST algorithm are shown in Table 2. Compared with AoA, the present invention has improvements in the evaluation metrics BLEU@1, BLEU@2, BLEU@3, BLEU@4, ROUGE_L, and CIDEr-D. Among them, the improvement in CIDEr-D is 1%.
[0124] Table 2 Comparison table of evaluation metrics after policy gradient learning
[0125]
[0126] It can be seen that the present invention can not only generate words closer to the real description, but also the semantic structure information of the generated sentences can more accurately express the semantic information of the images.
[0127] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable researchers familiar with the field to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.< / unk> < / unk>
Claims
1. An image description method based on multi-interaction information fusion, characterized in that It includes the following steps: Step 1: Preprocess the dataset and the real text descriptions of the images; Step 2: Extract the global image fusion features of the images in the dataset; Step 3: Use the multi-modal interaction information network to mine the relationship between the global image fusion features and the text semantic information, and obtain the context information at the current time step; The multi-modal interaction information network consists of a single multi-head attention layer, an AoA layer, an embedding layer, and U long short-term memory networks, and specifically includes the following steps: Step 3.1: Input the word vectors ∏ corresponding to all words in the corpus into the word embedding layer to obtain the word embedding vector W represented by one-hot encoding ∏ ∏; Step 3.2: Take the word embedding vector at the current time step, the global image fusion features, and the context information at the previous time step as the inputs of the U long short-term memory networks, and then obtain multiple complementary information on the interaction information between the global image fusion features and the word embedding vector. The specific formula is as follows: Among them, represents the U-th group of complementary information at the current time step; represents the U-th group of cell states at the current time step; W Π represents the word embedding matrix; Π t represents the input word at the current time step; represents the U-th group of context information at the previous time step; represents the U-th group of complementary information at the previous time step; represents the U-th group of cell states at the previous time step; Step 3.3: Perform a vector concatenation operation on multiple multi-modal interaction information for fusion, and map it to the same vector space through the embedding layer to generate multi-modal interaction information fusion features. The specific formula is as follows: Among them, p t represents the multi-modal interaction information fusion feature at the current time step; [.,.] represents the vector concatenation operation, and W h represents the mapping matrix; b h represents the bias; Step 3.4: Adopt the multi-head attention mechanism and the AoA mechanism to determine the correlation between the multi-modal interaction information fusion features and the image salient region features, so as to obtain the context vector for generating the word sequence. The specific formula is as follows: Among them, C t represents the context information at the current time step; W p represents the linear transformation matrix; represents the multi-head attention function; Concat represents the vector concatenation operation; head j represents the j-th head attention function, which is implemented using the scaled dot-product attention function; H represents the number of heads; represents the scaling factor; K j , V j represent the linear representations of the j-th head; softmax represents the normalized exponential function; Step 4: Use the linear unit of semantic decoding to decode the context information to generate the probability distribution of the output word sequence.
2. The method for image description based on multi-interaction information fusion according to claim 1, wherein The preprocessing in Step 1 specifically includes the following steps: Step 1.1: Divide the dataset in sequence, where 92% is divided into the training set, 4% is divided into the validation set, and the remaining 4% is divided into the test set; Step 1.2: Convert the text of the 5 real descriptions corresponding to each picture in the dataset to lowercase; Step 1.3: Statistically analyze each word in the real description converted to lowercase to obtain a corpus, and the corpus is in <unk>is the end flag, and remove the words that appear less than 5 times in the corpus;< / unk> Step 1.4: Count the true text description lengths L = {L1, L2, …, L i}, and set the true text description length of each image to argmax(L) + 2. For those with true text description lengths less than argmax(L) + 2, pad them with tokens.
3. The image description method based on multi-interaction information fusion according to claim 1, characterized in that The step 2 of extracting the global image fusion features of the images in the dataset is specifically as follows: Step 2.1: Use the object detection model to extract all the significant region features of the training set images, denoted as v = {v1, v2, …, v a}, where v a represents the a-th significant region feature; Step 2.2: Perform three linear mappings on the salient region features v of the image respectively, and denote the obtained linear representations as Q, K, and V respectively. The specific formula is as follows: Q = vW Q + b Q K = vW K + b K V = vW V + b V Among them, W Q , W K , W V represent linear transformation matrices; b Q , b K , b V represent biases; Step 2.3: Use the multi-visual semantic information interaction module to model the relationship between the image salient region features, and then obtain the global image fusion features.
4. The image description method based on multi-interaction information fusion according to claim 3, wherein The step 2.3 of using the multi-visual semantic information interaction module to model the relationship between the image salient region features, and then obtain the global image fusion features is specifically as follows: The multi-visual semantic information interaction module consists of 3×N×R linear layers, N×R Layer Norm layers, N×R multi-head attention mechanisms, and N×R AoA layers; Step 2.3.1: Adopt the multi-head attention mechanism to enable the image salient region features to selectively pay attention to other relevant region features to each other, so as to obtain the local feature relationship. The specific formula is as follows: f multi_head_att (Q, K, V) = Concat(head1, head2, …, head H ) Among them, f multi_head_att represents the multi-head attention function; Concat represents the vector concatenation operation; head j represents the j-th head attention function, which is implemented using the scaled dot-product attention function; H represents the number of heads; represents the scaling factor; Q j , K j , V j represent the linear representations of the j-th head; softmax represents the normalized exponential function; Step 2.3.2: Use the AoA mechanism to determine the correlation between the local feature relationship and the image salient region features, so that the salient features of each image can selectively pay attention to other region features that are truly relevant to it. The specific formula is as follows: where, σ is the sigmoid activation function; denotes element-wise multiplication, denotes the linear transformation matrix; b e and b j denotes the bias; Step 2.3.3: Repeat Step 2.3.1 and Step 2.3.2 for N times to obtain the high-level local feature relationship f AoAS ; Step 2.3.4: Perform a residual connection and normalization on the image salient region features and the high-level local feature relationship to obtain the enhanced image features. The specific formula is as follows: v' = LayerNorm(v + f AoAS (f multi_head_att , Q, K, V)) Among them, LayerNorm is the layer normalization function; Step 2.3.5: Repeat Step 2.3.1 to Step 2.3.4 for R times to generate multi-layer enhanced image features; Step 2.3.6: Use vector concatenation operation to fuse the multi-layer enhanced image features to obtain multi-layer enhanced image fusion features. The specific formula is as follows: where [.,.] represents the vector concatenation operation, and v' R represents the enhanced image feature of the R-th layer; represents the multi-layer enhanced image fusion feature; Step 2.3.7: Generate global image fusion features by performing average pooling on the multi-layer enhanced image fusion features. The specific formula is as follows: Among them, represents the global image fusion feature; a represents the number of channels of the multi-layer enhanced image fusion feature.
Citation Information
Patent Citations
An image description method that integrates visual context
CN110991515B
Image description generation method that integrates visual common sense and enhances multi-layer global features
CN113378919B
Image description generation method fusing visual common sense and enhancing multilayer global features
CN113378919A
Image description generation model method and device based on Transform structure and computer equipment
CN114266905A