Image description model training method based on context reasoning and image description method

By introducing a context reasoning mechanism and feature memory pool in the image description model, the problem of insufficient understanding of visual information in the image description model is solved, and more accurate image description is achieved.

CN115578726BActive Publication Date: 2025-06-27SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211106022.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2025-06-27
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

The existing image description model lacks understanding of visual information, resulting in inaccurate description content and reduces the accuracy of image description.

Method used

The image description model training method based on context reasoning is adopted, and the context reasoning process combined with semantic alignment is performed to optimize the image description model to improve the understanding of visual semantic information by hiding the state feature memory pool and visual attention feature memory pool.

Benefits of technology

By semantic alignment of visual features and language features, the model can better understand visual semantic information and improve the accuracy of image description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115578726B_ABST
    Figure CN115578726B_ABST
Patent Text Reader

Abstract

The present invention provides a method for training an image description model based on context reasoning and an image description method. Among them, the model training method includes: obtaining the image region features and text data of the image data in the current training data; respectively performing a context reasoning process combined with semantic alignment on each word in the text data to obtain the semantic reasoning information and semantic alignment deviation corresponding to each word; obtaining the cross-entropy loss and semantic alignment loss of the current text data, and optimizing the image description model based on the cross-entropy loss and semantic alignment loss; updating the training data and the corresponding semantic label information, and repeatedly executing the above steps based on the updated training data and the corresponding semantic label information until exiting, so as to obtain the trained image description model, thereby enhancing the model reasoning ability and improving the accuracy of the model in understanding multi-modal semantics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image description, and particularly to an image description model training method and an image description method based on context reasoning. One or more embodiments of this specification simultaneously relate to an image description model training method and an image description method, a device, and a computer storage medium based on context reasoning. Background Art

[0002] Image description is a process in which a machine automatically generates a description statement of the content of a given image, similar to "translating" the image into a language description; its essence is to segment the image and then use data, symbols, and formal languages to represent the features of different sub-regions after segmentation; therefore, the understanding of the content of each sub-region image, that is, representing the semantic information contained in the visual features in the semantic space of the language modality, plays a crucial role in the accuracy of the entire image description task.

[0003] Currently, the commonly used image description models are usually models based on the encoder-decoder structure; among them, the encoder usually uses a convolutional neural network pre-trained on a large-scale image classification dataset to extract visual features from the image; the decoder usually uses a recurrent neural network or a self-attention mechanism network, such as a long short-term memory network (LSTM) and a Transformer, to generate a natural statement describing the content of the image, that is, a sequence of words, based on the extracted visual features of the image. However, the currently mainly used image description models usually only perform word-level implicit semantic alignment on the semantic features of a single word and the visual features of the corresponding sub-region image through an attention mechanism, that is, calculating visual attention features related to semantics, without considering the semantic alignment relationship at the sentence segment level between multiple visual features associated with each word in the text statement and the semantic features of multiple words, resulting in the semantic information in the visual features from the image not being fully understood and extracted by the model, thereby reducing the model's understanding ability of visual semantic information. Therefore, there are problems such as insufficient understanding of visual information and inaccurate description content, which in turn lead to insufficient accuracy of image description. Summary of the Invention

[0004] In view of the above-mentioned drawbacks existing in the prior art, the purpose of the present invention is to provide an image description method, a device, and a computer storage medium based on context reasoning, which are used to solve the problems of insufficient understanding of visual information and inaccurate description content in the existing image description models, so as to improve the accuracy of image description.

[0005] To achieve the above and other related objectives, the present invention provides, in a first aspect, a method for training an image description model based on context reasoning. Based on each group of training data and the semantic label information corresponding to the training data, a model training process is performed on the image description model to obtain a trained image description model. Wherein, the training data includes image data and text data associated with the image data; the text data includes each pre-calibrated word; the training process includes: obtaining the image data in the current training data, and extracting the image region features of the image data; obtaining the text data in the current training data; for each word in the text data, respectively performing a context reasoning process combined with semantic alignment to obtain the semantic reasoning information and semantic alignment deviation corresponding to each word; constructing the cross-entropy loss of the current text data based on the semantic reasoning information corresponding to each word and the semantic label information of the current image; and obtaining the semantic alignment loss of the current text data based on the semantic alignment deviation corresponding to each word; obtaining the total model loss of the current text data based on the cross-entropy loss and the semantic alignment loss, so as to optimize the image description model based on the total model loss; updating the training data and the corresponding semantic label information, and repeating the above steps based on the updated training data and the corresponding semantic label information.

[0006] In an embodiment of the present invention, the image description model includes a hidden state feature memory pool and a visual attention feature memory pool. Then, a single context reasoning process combined with semantic alignment includes: obtaining the hidden state feature and visual attention feature of the current word; based on the hidden state feature and the visual attention feature of the current word, correspondingly updating the hidden state feature memory pool and the visual attention feature memory pool to respectively obtain the current hidden state feature sequence and the current visual attention feature sequence; obtaining the hidden state semantic feature and visual semantic feature of the current word based on the current hidden state feature sequence and the current visual attention feature sequence; performing semantic alignment on the hidden state semantic feature and the visual semantic feature to obtain the semantic alignment deviation of the current word; obtaining the semantic reasoning feature of the current word based on the visual attention feature and semantic feature of the current word; and determining the semantic information with the highest correlation degree in a preset semantic table based on the semantic reasoning feature of the current word as the semantic reasoning information of the current word.

[0007] In one embodiment of the present invention, the implementation method for obtaining the hidden state feature and visual attention feature of the current word includes: obtaining the hidden state feature corresponding to the previous word and obtaining the word embedding feature of the current word; based on the word embedding feature of the current word, the image region feature, and the hidden state feature corresponding to the previous word, obtaining the hidden state feature of the current word; and based on the hidden state feature and the image region feature of the current word, using the attention mechanism to obtain the visual attention feature of the current word.

[0008] In one embodiment of the present invention, the semantic alignment of the hidden state semantic feature and the visual semantic feature to obtain the semantic alignment deviation of the current word includes: inputting the hidden state semantic feature and the visual semantic feature into a semantic alignment function to obtain the semantic alignment deviation of the current word, which is:

[0009]

[0010] where L aln is the semantic alignment deviation of the current word; Smooth-L1 is a distribution adjustment function.

[0011] In one embodiment of the present invention, the method for obtaining the hidden state semantic feature and visual semantic feature of the current word based on the current hidden state feature sequence and the current visual attention feature sequence includes: constructing a hidden state feature query vector of the current word based on the sequence feature of the current visual attention feature sequence and the word embedding feature of the current word; and constructing a visual attention feature query vector of the current word based on the sequence feature of the current hidden state feature sequence and the word embedding feature of the current word; performing temporal and semantic enhancement on the current hidden state feature sequence to obtain an enhanced hidden state feature sequence; and performing temporal and semantic enhancement on the current visual attention feature sequence to obtain an enhanced visual attention feature sequence; based on the enhanced hidden state feature sequence and the hidden state feature query vector, using the attention mechanism to obtain the hidden state semantic feature of the current word; and based on the enhanced visual attention feature and the visual attention feature query vector, using the attention mechanism to obtain the visual semantic feature of the current word.

[0012] In an embodiment of the present invention, the implementation manner of obtaining the hidden state semantic feature of the current word by using the attention mechanism based on the enhanced hidden state feature sequence and the hidden state feature query vector includes: based on the enhanced hidden state feature sequence, obtaining an interaction-enhanced hidden state feature sequence by using the self-attention mechanism, and obtaining the hidden state semantic feature by using the traditional attention mechanism based on the interaction-enhanced hidden state feature sequence and the hidden state feature query vector; and the implementation manner of obtaining the visual semantic feature of the current word by using the attention mechanism based on the enhanced visual attention feature and the visual attention feature query vector includes: based on the enhanced visual attention feature sequence, obtaining an interaction-enhanced visual attention feature sequence by using the self-attention mechanism, and obtaining the visual semantic feature by using the traditional attention mechanism based on the interaction-enhanced visual attention feature sequence and the visual attention feature query vector.

[0013] In an embodiment of the present invention, the performing temporal and semantic enhancement on the current hidden state feature sequence includes: based on the sequence position of each hidden state feature in the current hidden state feature sequence, using a first position encoder to obtain the temporal information of each hidden state feature; superimposing the temporal information of each hidden state feature and the word embedding feature of the current word onto the corresponding hidden state feature; and the performing temporal and semantic enhancement on the current visual attention feature sequence includes: based on the sequence position of each visual attention feature in the current visual attention feature sequence, using a second position encoder to obtain the temporal information of each visual attention feature; superimposing the temporal information of each visual attention feature and the word embedding feature of the current word onto the corresponding visual attention feature.

[0014] In an embodiment of the present invention, the first position encoder and the second position encoder are the same and both include:

[0015]

[0016] where p i represents the relative position of the feature in the corresponding feature sequence; j is the dimension of the position encoding representation, when j is odd, f() is sin(), and when j is even, f() is cos().

[0017] In an embodiment of the present invention, the semantic feature of the current word includes the context reasoning feature; then obtaining the semantic reasoning feature of the current word based on the visual attention feature and the semantic feature of the current word includes: obtaining the context reasoning feature of the current word; and obtaining the semantic reasoning feature of the current word based on the visual attention feature of the current word and the context reasoning feature.

[0018] In an embodiment of the present invention, obtaining the context reasoning feature of the current word includes: obtaining a multi-modal feature sequence of the current word based on the enhanced hidden state feature sequence and the enhanced visual attention feature sequence; constructing a multi-modal feature query vector of the current word based on the hidden state query vector of the current word and the hidden state feature of the current word; and obtaining the context reasoning feature of the current word by using a traditional attention mechanism based on the multi-modal feature sequence and the multi-modal feature query vector of the current word.

[0019] In an embodiment of the present invention, the implementation manner of constructing the multi-modal feature query vector of the current word includes: using a gated linear unit, taking the hidden state feature query vector of the current word and the hidden state feature of the current word as inputs of the gated linear unit, and constructing the multi-modal feature query vector of the current word.

[0020] The present invention provides an image description method in a second aspect, which is characterized by including: constructing each training data set based on sample data of image description; a single set of the training data includes image data and text data associated with the image data; training a preset model by using the above-mentioned context reasoning-based image description model training method based on each training data set and semantic label information corresponding to the training data to obtain a trained image description model; and performing context semantic information reasoning on the image to be described by using the trained image description model to obtain the semantic information of the input image.

[0021] The present invention provides an electronic device in a third aspect, including: a processor and a memory; the memory is used for storing a computer program, and the processor is used for executing the computer program stored in the memory so that the electronic device executes the above-mentioned context reasoning-based image description model training method or the above-mentioned image description method.

[0022] The present invention provides a computer storage medium in a fourth aspect, where the computer storage medium stores a computer program, and the computer program is executed by a processor to execute the above-mentioned context reasoning-based image description model training method or the above-mentioned image description method.

[0023] As described above, the present invention provides a method for training an image description model based on context reasoning, an image description method, a device, and a computer storage medium. By setting up a hidden state feature memory pool and a visual attention feature memory pool, during the context reasoning process for each word, based on the hidden state feature information stored in the hidden state feature memory pool and the visual attention feature information stored in the visual attention feature memory pool, a semantic alignment deviation corresponding to each word is obtained. Based on the cross-loss of semantic reasoning for each word and the semantic alignment deviation, the image description model is optimized to obtain a final image description model, so that the model can align the semantic information extracted from the visual feature sequence and the semantic information extracted from the language time feature sequence, making the semantic features and visual semantic features as close as possible in the feature space, helping the model to better understand the visual semantic information, and thus improving the accuracy of image description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It shows a schematic flowchart of the method for training an image description model based on context reasoning provided by the present invention in an embodiment;

[0025] Figure 2 It shows a schematic flowchart of the process of performing context reasoning combined with semantic alignment for a single word in the present invention in an embodiment;

[0026] Figure 3 It shows a schematic flowchart of step S201 of the present invention in an embodiment;

[0027] Figure 4 It shows a schematic flowchart of step S202 of the present invention in an embodiment;

[0028] Figure 5 It shows a schematic flowchart of the method for training an image description model based on context reasoning provided by the present invention in another embodiment;

[0029] Figure 6 It shows a schematic flowchart of the process of obtaining the context reasoning feature of the current word in the present invention in an embodiment;

[0030] Figure 7 A schematic flowchart of the image description method provided by the present invention in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0032] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0033] For the convenience of understanding the technical solution of this application, the following terms are explained as follows:

[0034] Image description represents the visual content in an image in the form of text language.

[0035] Context reasoning is to perform comprehensive semantic information reasoning on the current text based on the semantic information of the current text and the image visual feature information, and based on the semantic information of the context.

[0036] To solve the problems existing in the prior art, the present invention provides a method for training an image description model based on context reasoning in the first aspect. Based on each group of training data and the semantic label information corresponding to the training data, a model training process is performed on the image description model to obtain a trained model; that is, after performing the model training process on the image description model based on a single group of training data, the training data is updated, and the model training process is re-executed based on the new training data, and this process is repeated until exiting, so as to obtain the final image description model.

[0037] Among them, the image description model is used to obtain the semantic information of the text through context reasoning for the text data associated with the image, as the description information of the image.

[0038] A single group of the training data includes image data and text data associated with the image data; each word pre-calibrated is included in the text data; each word is arranged in sequence in the text data.

[0039] The semantic label information is the actual semantic information corresponding to each word in the text data.

[0040] Please refer to Figure 1, which is a schematic flow diagram of the method for training an image description model based on context reasoning in an embodiment of the present invention.

[0041] In this embodiment, the image description model includes a pre-constructed feature memory pool, including a hidden state memory pool and a visual attention feature memory pool; wherein, the hidden state memory pool is used to store the hidden state features of words, and the visual attention feature memory pool is used to store the visual attention features of words.

[0042] Among them, the hidden state feature is the semantic feature of a word combined with the image partition feature; the visual attention feature is the image feature combined with the semantic feature of the word.

[0043] As Figure 1 shown, when the method for training an image description model based on context reasoning is executed, it includes the following steps:

[0044] S100, obtain the image data in the current training data, and extract the image region features of the image data;

[0045] Specifically, partition the current image data to obtain each region of the figure;

[0046] Optionally, use an object detection model to partition the image data, and extract the set of region features of each partition in the image, which is:

[0047]

[0048] Among them, R is the set of region features of each partition of the image; m is the number of partitions; d v is the feature dimension of the region feature; in this embodiment, d v = 2048.

[0049] Map the region features of each partition to new region features through a fully connected layer; take the average of each new region feature to obtain the image region feature

[0050] Among them, d h is the feature dimension of the new region feature; optionally, the feature dimension of this new region feature is d h = 1024.

[0051] S200, obtain the text data in the current training data; for each word in the text data, respectively perform a context reasoning process combined with semantic alignment to obtain the semantic reasoning information and semantic alignment deviation corresponding to each word;

[0052] In this embodiment, for a single word, perform the context reasoning process combined with semantic alignment (hereinafter referred to as "context reasoning process"), asFigure 2 As shown, it includes:

[0053] S201, obtaining the hidden state feature and visual attention feature of the current word;

[0054] Specifically, as Figure 3 shown, this step includes the following sub-steps:

[0055] S201A, obtaining the hidden state feature corresponding to the previous word;

[0056] Wherein, the previous word is the word in the current text data that is in the previous position of the current word, that is, the word corresponding to the context reasoning process in the previous cycle;

[0057] The hidden state feature of the previous word is the hidden state feature obtained when performing the context reasoning process on the previous word.

[0058] Specifically, before performing the context reasoning process of the current cycle, in the hidden state feature sequence stored in the hidden state memory pool, extract the hidden state feature corresponding to the previous word.

[0059] S201B, obtaining the word embedding feature of the current word; based on the word embedding feature of the current word, the image region feature, and the hidden state feature corresponding to the previous word, obtaining the hidden state feature of the current word;

[0060] In a specific embodiment, using a word embedding function, the word embedding feature extracted for the current word is:

[0061] e t = embed(w t )

[0062] Wherein, w t is the current word, embed() represents the word embedding function, e t is the word embedding feature of the current word, and

[0063] Concatenate the word embedding feature of the current word and the image region feature to obtain the concatenated feature of the current word, which is:

[0064] x t = [e t ; I g

[0065] Wherein, x t is the concatenated feature of the current word w t ; [;] represents the vector concatenation operation.

[0066] ​Based on the splicing features of the current word and the hidden state features of the previous word, a long short-term memory neural network is used to obtain the hidden state features of the current word, which are:

[0067] h t = LSTM(x t , h t-1 )

[0068] where h t represents the hidden state features of the current word; LSTM() is the long short-term memory neural network.

[0069] S201C. Based on the hidden state features of the current word and the image region features, an attention mechanism is used to obtain the visual attention features of the current word;

[0070] Using the traditional attention mechanism, feature interaction is performed on the hidden state features of the current word and the image region features.

[0071] Specifically, taking the hidden state features h t of the current word as the query vector, taking the image region features R as the key vector and the value vector, and inputting them into the traditional attention mechanism to calculate the visual attention features related to the hidden state features; that is, according to the correlation between the query vector and the key vector, weights are assigned to the value vectors, and the weighted sum of the weighted value vectors is calculated to obtain the visual attention features of the current word, which are:

[0072] a t = attention(R, h t )

[0073] where attention() is the traditional attention mechanism; a t represents the visual attention features of the current word.

[0074] Optionally, before performing feature interaction on the hidden state features of the current word and the image region features using the traditional attention mechanism, it further includes:

[0075] Using the self-attention mechanism to perform internal feature interaction on the image region features to obtain new image region features, and based on the new image region features, subsequent steps are performed, which are:

[0076] R s = self-attention(R)

[0077] where R s represents the new image region features.

[0078] S202. Update the hidden state feature memory pool based on the hidden state features of the current word to obtain the current hidden state feature sequence; and update the visual attention feature memory pool based on the visual attention features of the current word to obtain the current visual attention feature sequence. Based on the current hidden state feature sequence and the current visual attention feature sequence, obtain the hidden state semantic features and visual semantic features of the current word; and perform semantic alignment on the hidden state semantic features and the visual semantic features to obtain the semantic alignment deviation of the current word.

[0079] In this embodiment, as Figure 4 shown, step S202 specifically includes the following sub-steps:

[0080] S202A. Update the hidden state feature memory pool based on the hidden state features of the current word to obtain the current hidden state feature sequence; and update the visual attention feature memory pool based on the visual attention features of the current word to obtain the current visual attention feature sequence.

[0081] Specifically, input the hidden state feature h t of the current word to the top layer of the hidden state feature memory pool, and delete the hidden state feature stored in the bottom layer of the hidden state feature memory pool to update the hidden state feature memory pool. Take the feature sequence stored in the updated hidden state feature memory pool as the current hidden state feature sequence, that is:

[0082] H t ={h t ,h t-1 ,...,h t-k+1};

[0083] where h t , h t-1 ……, h t-k+1 are the respective hidden features in the updated hidden state feature memory pool; k is the number of features stored in the memory pool.

[0084] Also, input the visual attention feature a t of the current word to the top layer of the visual attention feature memory pool, and delete the visual attention feature stored in the bottom layer of the visual attention feature memory pool to obtain the updated visual attention feature memory pool. Take the feature sequence stored in the updated visual attention feature memory pool as the current visual attention feature sequence, that is:

[0085] A t ={a t ,a t-1 ,...,at-k+1}

[0086] Among them, a t 、a t-1 ……、a t-k+1 are the visual attention features in the updated visual attention feature memory pool; k is the number of features stored in the memory pool.

[0087] In this embodiment, the capacity of the hidden state feature memory pool and the visual attention feature memory pool is the same, which is used to store the features generated corresponding to the iterative execution of the context reasoning process of combining semantic alignment k times respectively.

[0088] Optionally, k is 8.

[0089] S202B, obtain the sequence feature of the current visual attention feature sequence; based on this sequence feature and the word embedding feature of the current word, construct the hidden state feature query vector of the current word; and obtain the sequence feature of the current hidden state feature sequence; based on this sequence feature and the word embedding feature of the current word, construct the visual attention feature query vector of the current word;

[0090] Optionally, the sequence feature of the current visual attention feature sequence includes the mean value of the current visual attention feature sequence, that is, the mean value of each feature in the current visual attention feature sequence.

[0091] Specifically, obtaining the mean value of the visual attention feature sequence is the mean value of each visual attention feature in the current visual attention feature sequence; based on this mean value and the word embedding feature of the current word, construct the first gate vector, which is:

[0092] g H =σ(W H (e t +mean(A t )))

[0093] Among them, g H is the first gate vector, which is used to dynamically balance the semantic information obtained from the word embedding feature and the mean value of the attention feature sequence; σ is sigmoid, which limits the feature between 0 and 1, is the Hadamard product; mean(A t ) is the mean value of the visual attention feature sequence.

[0094] Based on the first gate vector, the mean value of the visual attention feature sequence and the word embedding feature of the previous word, construct the hidden state feature query vector of the current word, which is:

[0095]

[0096] Among them, is the hidden state feature query vector of the current word.

[0097] Optionally, the sequence feature of the current hidden state feature sequence includes the mean value of the current hidden state feature sequence, that is, the mean value of each feature in the current hidden state feature sequence.

[0098] Specifically, obtain the mean value of each hidden state feature in the current hidden state feature sequence as the mean value of the hidden state feature sequence; based on this mean value and the word embedding feature of the current word, construct the second gate vector as:

[0099] g A =σ(W A (e t +mean(H t ))

[0100] where g A is the second gate vector, which is used to dynamically balance the semantic information obtained from the word embedding feature and the mean value of the hidden state feature sequence; σ is sigmoid, which limits the feature between 0 and 1, is the Hadamard product; mean(H t ) is the mean value of the hidden state feature sequence.

[0101] Based on the second gate vector, the mean value of the hidden state feature sequence, and the word embedding feature of the previous word, construct the visual attention feature query vector of the current word as:

[0102]

[0103] where is the visual attention feature query vector of the current word.

[0104] S202C, perform temporal and semantic enhancement on the current hidden state feature sequence to obtain an enhanced hidden state feature sequence; and perform temporal and semantic enhancement on the current visual attention feature sequence to obtain an enhanced visual attention feature sequence;

[0105] In this embodiment, for the current hidden state feature sequence, obtain the temporal information of each hidden state feature based on its sequence position in the current hidden state feature sequence; optionally, use a pre-constructed position encoder to convert the sequence position of each hidden state feature into temporal information;

[0106] Based on the temporal information of each of the hidden state features and the word embedding feature of the current word, perform temporal and semantic enhancement on each of the hidden state features to obtain each of the hidden state features after temporal and semantic enhancement, which is:

[0107]

[0108] Wherein, is the enhanced hidden state feature; p i is the hidden state feature h i at the relative position in the current hidden state feature sequence H t ; W he is the model parameter; pe() is the relative position encoder, which is used to obtain the temporal information of each of the hidden state features.

[0109] In a specific embodiment, pe() is calculated using the following calculation formula, which is:

[0110]

[0111] Wherein, p i represents the relative position of each feature in the corresponding feature sequence; j is the dimension of the position encoding representation. When j is odd, f() is sin(), and when it is even, it is cos(). The above enhancement process is performed on each of the hidden state features in the current hidden state feature memory pool to obtain the enhanced hidden state feature sequence, which is:

[0112]

[0113] For the current visual attention feature sequence, based on the sequence position of each of the visual attention features in the current visual attention feature sequence, obtain the temporal information of each of the visual attention features; optionally, use a pre-constructed position encoder to convert the sequence position of each of the visual attention features into temporal information;

[0114] Based on the temporal information of each of the visual attention features and the word embedding feature of the current word, perform temporal and semantic enhancement on each of the visual attention features to obtain each of the visual attention features after temporal and semantic enhancement, which is:

[0115]

[0116] Wherein, is the enhanced visual attention feature; p i represents the relative position of the visual attention feature a i in the current visual attention feature sequence A t ; Wae are model parameters; pe() is a relative position encoder for obtaining the temporal information of each of the visual attention features.

[0117] In a specific embodiment, pe() is calculated using the following calculation formula:

[0118]

[0119] where p i represents the relative position of each feature in the corresponding feature sequence; j is the dimension of the position encoding representation. When j is odd, f() is sin(), and when j is even, it is cos().

[0120] For each of the visual attention features in the current visual attention feature sequence, the above enhancement process is performed to obtain the enhanced visual attention feature sequence, which is

[0121] S202D. Based on the enhanced hidden state feature sequence and the hidden state feature query vector, an attention mechanism is used to obtain the hidden state semantic feature of the current word; and based on the enhanced visual attention feature and the visual attention feature query vector, an attention mechanism is used to obtain the visual semantic feature of the current word;

[0122] In this embodiment, the enhanced hidden state feature sequence is input into the self-attention mechanism to perform the interaction of the internal features of each of the enhanced hidden state features, and an interaction-enhanced hidden state feature sequence is obtained, which is:

[0123]

[0124] where is the interaction-enhanced hidden state feature sequence.

[0125] Based on the interaction-enhanced hidden state feature sequence and the hidden state feature query vector, a traditional attention mechanism is used to aggregate features to obtain the interaction-enhanced hidden state semantic feature, which is:

[0126]

[0127] where the is the interaction-enhanced hidden state semantic feature.

[0128] And, the enhanced visual attention feature sequence is input into the self-attention mechanism to perform the interaction of the internal features of each of the enhanced visual attention features, and an interaction-enhanced visual attention feature sequence is obtained, which is

[0129]

[0130] Among them, is the sequence of visual attention features enhanced by the interaction.

[0131] Based on the sequence of visual attention features enhanced by the interaction and the visual attention feature query vector, traditional attention mechanism is used to aggregate features to obtain the visual semantic features after interaction enhancement, which is:

[0132]

[0133] Among them, the is the visual semantic feature after interaction enhancement.

[0134] S202E, based on the hidden state semantic feature and the visual semantic feature, perform semantic alignment of the current word to obtain the semantic alignment deviation of the current word;

[0135] Specifically, input the hidden state semantic feature and the visual semantic feature into the semantic alignment function to obtain the semantic alignment deviation between the two as the semantic alignment deviation of the current word, which is:

[0136]

[0137] Among them, L aln is the semantic alignment deviation of the current word; Smooth-L1 is a distribution adjustment function composed of the combination of L1 loss and L2 loss, which is used to solve the non-smooth problem of L1 loss in the range of -1 to 1.

[0138] It should be noted that the above-mentioned step S202C can also be executed before step S202B or executed simultaneously with step S202B, which is not limited here.

[0139] S203, based on the visual attention feature and semantic feature of the current word, obtain the semantic reasoning feature of the current word; based on the semantic reasoning feature of the current word, determine the semantic information with the highest correlation degree in the semantic table as the semantic reasoning information of the current word;

[0140] In this embodiment, the hidden state feature h t is used as the semantic feature of the current word;

[0141] Specifically, based on the visual attention feature of the current word and the hidden state feature, use the gated linear unit to calculate the semantic reasoning feature of the current word, which is:

[0142] out t = GLU(W out [at ; h t )

[0143] Among them, out t is the semantic inference feature of the current word; W out is the model parameter and satisfies

[0144] Based on the pre-constructed word semantic table, using the softmax function for solving multi-classification problems, select the semantic information with the highest feature correlation weight in the semantic table as the semantic inference information of the current word.

[0145] y t = softmax(W v out t )

[0146] Among them, d w is the vocabulary size, softmax() is the softmax function, used to select the semantic information with the highest weight in the semantic table; y t is the semantic inference information of the current word.

[0147] S204, update the current word, and repeat the above steps based on the updated current word.

[0148] In this embodiment, according to the arrangement order of each word in the text data, use the word next in line to the current word as the new current word, and return to step S201; based on this new current word, re-execute the above steps S201 to step S204.

[0149] Repeat the above process until exiting, so as to obtain the semantic alignment deviation corresponding to each word and obtain the semantic inference information corresponding to each word.

[0150] S300, based on the semantic inference information of each word in the current text data and the semantic label information of the current image, construct the cross-entropy loss of the current text data; and based on the semantic alignment deviation of each word in the current text data, obtain the semantic alignment loss of the current text data; based on the cross-entropy loss and the semantic alignment loss, obtain the total model loss of the current text data, so as to optimize the image description model based on this total model loss;

[0151] Specifically, based on the semantic inference information of each word and the corresponding semantic label information, construct the cross-entropy loss corresponding to each word; and accumulate the cross-entropy losses corresponding to each word to obtain the cross-entropy loss of the current text data, which is:[[]]

[0152]

[0153] Among them, L xe is the cross-entropy loss of the current text data, which is used to maximize the probability of the model outputting the semantic label information of the current image, that is, to constrain the distribution probability between the semantic inference information and the semantic label information to be similar. p() represents the probability distribution function, I is the current image, l is the total number of words in the current text data, represents the word sequence composed of the first word to the (t-1)th word in the text data; is the tth word in the text data.

[0154] In addition, the semantic alignment losses corresponding to each word are accumulated to obtain the semantic alignment loss of the current text data, which is:

[0155]

[0156] Among them, L ALN is the semantic alignment loss of the current text data.

[0157] Based on the cross-entropy loss and the semantic alignment loss of the current text data, the total loss of the model of the current text data is obtained, which is:

[0158] L = L xe + λL ALN

[0159] Among them, L is the total loss of the model of the current text data, and λ is a tuning parameter used to balance the proportion of the cross function and the semantic alignment loss in the total loss of the model; optionally, λ is set to 1.

[0160] S400. Update the training data and the corresponding semantic label information, and repeat the above steps based on the updated training data and the corresponding semantic label information until exiting, so as to obtain the trained image description model.

[0161] In this embodiment, obtain the next training data and the corresponding semantic label information as the new training data and the corresponding semantic label information, and return to step S100; based on the new training data and the corresponding semantic label information, re-execute the above steps S100 to S400;

[0162] Repeat the above process until exiting, so as to obtain the trained image description model.

[0163] In this embodiment, the method for training an image description model based on context reasoning sets up a hidden state feature memory pool and a visual attention feature memory pool. When performing context reasoning on each word, semantic alignment deviations corresponding to each word are obtained based on the hidden state feature information stored in the hidden state feature memory pool and the visual attention feature information stored in the visual attention feature memory pool. The image description model is optimized based on the cross-loss of semantic reasoning of each word and the semantic alignment deviation to obtain the final image description model. Thus, the model can align the semantic information extracted from the visual feature sequence with the semantic information extracted from the language time feature sequence, making the semantic features and visual semantic features as close as possible in the feature space, helping the model to more fully understand the visual semantic information, and thereby improving the accuracy of image description.

[0164] Please refer to Figure 5 , which shows a schematic flowchart of the method for training an image description model based on context reasoning provided by the present invention in another embodiment.

[0165] As Figure 5 shown, in this embodiment, the context reasoning process combined with semantic alignment in the method for training an image description model based on context reasoning is basically the same as the execution process shown in Figure 2 , except that the semantic feature of the current word further includes the context reasoning feature of the current word. When performing the above step S203, it can also be:

[0166] S203’, obtain the context reasoning feature of the current word; based on the visual attention feature of the current word and the context reasoning feature, obtain the semantic reasoning feature of the current word; based on the semantic reasoning feature of the current word, determine the semantic information with the highest correlation degree in the semantic table as the semantic reasoning information of the current word.

[0167] Among them, the context reasoning feature is a multi-modal semantic feature of the current word based on context information, and is used to represent the context comprehensive information associated with the word embedding information of the current word and the current image feature information.

[0168] In this embodiment, the specific implementation manner of obtaining the context reasoning feature of the current word, as Figure 6 shown, includes the following sub-steps:

[0169] S601, based on the enhanced hidden state feature sequence and the enhanced visual attention feature sequence, obtain the multi-modal feature sequence of the current word.

[0170] Specifically, the enhanced hidden state feature sequence and the enhanced visual attention feature sequence are concatenated to obtain the multimodal feature sequence of the current word, which is:

[0171]

[0172] where is the multimodal feature sequence of the current word.

[0173] Optionally, when this sub-step is executed, it further includes:

[0174] Based on the multimodal feature sequence of the current word, a self-attention mechanism is used to perform feature interaction to obtain a new multimodal feature sequence, and subsequent steps are executed based on the new multimodal feature sequence.

[0175] In a specific embodiment, the self-attention mechanism is used to perform feature interaction twice, which is:

[0176]

[0177] where self-attention2() indicates that the self-attention layer of this self-attention mechanism has two layers and is used to implement feature interaction twice; represents the new multimodal feature sequence.

[0178] S602. Based on the hidden state query vector of the current word and the hidden state feature of the current word, construct the multimodal feature query vector of the current word;

[0179] Specifically, a gated linear unit is used to take the hidden state feature query vector of the current word and the hidden state feature of the current word as the input of the gated linear unit to construct the multimodal feature query vector of the current word, which is:

[0180]

[0181] where is the multimodal feature query vector of the current word; GLU() is the gated linear unit function; this function equally divides the input features into two parts according to the dimension size, one part is used as information, and the other part is processed by the sigmoid function as a gate, and the output is the product of the two; W M is the model parameter.

[0182] S603. Based on the multimodal feature sequence and the hidden state query vector of the current word, use the traditional attention mechanism to obtain the context inference feature of the current word, which is:

[0183]

[0184] Among them, is the context inference feature of the current word.

[0185] Based on the context inference feature and the visual attention feature obtained above, a gated linear unit is used to calculate the semantic inference feature of the current word, which is:

[0186]

[0187] Among them, out t is the semantic inference feature of the current word; a t is the visual attention feature of the current word; W out is a model parameter and satisfies

[0188] Based on the pre-constructed word semantic table, a normalized exponential function for solving multi-classification problems is used to select the semantic information with the highest feature correlation weight in the semantic table as the semantic inference information of the current word.

[0189] y t = softmax(W v out t )

[0190] Among them, d w is the vocabulary size, softmax() is the normalized exponential function, which is used to select the semantic information with the highest weight in the semantic table; y t is the semantic inference information of the current word.

[0191] In the method for training an image description model based on context inference in this embodiment, by obtaining the context inference feature of the current word and based on the visual attention feature and the context inference feature of the current word, the semantic inference feature of the current word is obtained. Therefore, in the context inference process of each word, the correlation between the image visual feature and the text language feature can be fully exploited, thereby enhancing the model inference ability and further improving the accuracy of image description.

[0192] To solve the problems existing in the prior art, the present invention provides an image description method in a second aspect, which is used to implement the image description process of the image to be described to obtain the description information of the image.

[0193] Please refer to Figure 7 , which shows the schematic flow chart of the image description method provided by the present invention in an embodiment.

[0194] As Figure 7 shown, the image description method includes the following steps:

[0195] S10. Construct training data for each group based on sample data of image descriptions. Each group of the training data includes image data and text data associated with the image data;

[0196] Among them, the text data includes each pre - calibrated word; each word is arranged in order in the text data.

[0197] S20. Train a preset image description model based on each group of the training data and semantic label information corresponding to each group of the training data to obtain a trained image description model;

[0198] Among them, the semantic label information is the actual semantic information corresponding to each word in the text data.

[0199] Specifically, use the above - mentioned image description model training method based on context reasoning to train the preset image description model to obtain a trained image description model.

[0200] Among them, the image description model includes a pre - constructed feature memory pool, including a hidden state memory pool and a visual attention feature memory pool; the hidden state memory pool is used to store the hidden state features of words; the visual attention feature memory pool is used to store the visual attention features of words.

[0201] The hidden state feature is the semantic feature of a word combined with image partition features; the visual attention feature is the image feature combined with the semantic feature of a word.

[0202] S30. For the image to be described, use the trained image description model to perform a context reasoning process combined with semantic alignment on the image to obtain the semantic information of the input image.

[0203] To solve the problems existing in the prior art, in the third aspect of the present invention, an electronic device is further provided. The electronic device includes: a processor, a memory, a transceiver, a communication interface, and a system bus; the memory and the communication interface are connected to the processor and the transceiver through the system bus and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate with other devices, and the processor and the transceiver are used to run the computer programs to enable the processing device to execute each step of the above - mentioned image description model training method based on context reasoning or the above - mentioned image description method.

[0204] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0205] In addition, in the fourth aspect of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the program is called by a processor, it implements each step of the above-mentioned method for training an image description model based on context reasoning or the above-mentioned image description method.

[0206] Among them, the computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example (but not limited to), an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device.

[0207] The computer-readable program described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0208] In summary, the present invention provides a method for training an image description model based on context reasoning, an image description method, a device, and a computer storage medium. By setting up a hidden state feature memory pool to store the current hidden state feature information, and by setting up a visual attention feature memory pool to store the current visual attention feature information; during the context reasoning process corresponding to each word, based on the corresponding hidden state feature information and visual attention feature information, the semantic alignment deviation corresponding to each word is obtained, and based on the cross-loss of semantic reasoning and the semantic alignment deviation of each word, the image description model is optimized, so that the model aligns the semantic information extracted from the visual feature sequence and the semantic information extracted from the language temporal feature sequence, and further constrains the visual semantics and the language semantics in the same semantic space, further enhancing the model's ability to understand and capture visual semantic information, and can significantly improve the model's ability to understand multi-modal semantics and perform effective reasoning, and improve the accuracy of image description. In addition, by obtaining the context reasoning feature of the current word, and based on the visual attention feature of the current word and the context reasoning feature, the semantic reasoning feature of the current word is obtained, so that during the context reasoning process of each word, the correlation between the image visual feature and the text language feature can be fully mined, thereby enhancing the model's reasoning ability and further improving the accuracy of image description.

[0209] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for training an image description model based on context reasoning, characterized in that Based on each group of training data and the semantic label information corresponding to the training data, perform a model training process on the image description model to obtain a trained image description model; wherein, the training data includes image data and text data associated with the image data; the text data includes each pre-calibrated word; the image description model includes a hidden state feature memory pool and a visual attention feature memory pool; The training process includes: Obtain the image data in the current training data, and extract the image region features of the image data; Obtain the text data in the current training data; for each word in the text data, respectively perform a context reasoning process combined with semantic alignment to obtain the semantic reasoning information and semantic alignment deviation corresponding to each word; Based on the semantic reasoning information corresponding to each word and the semantic label information of the current image, construct the cross-entropy loss of the current text data; based on the semantic alignment deviation corresponding to each word, obtain the semantic alignment loss of the current text data; Based on the cross-entropy loss and the semantic alignment loss, obtain the total model loss of the current text data, so as to optimize the image description model based on the total model loss; Update the training data and the corresponding semantic label information, and repeat the above steps based on the updated training data and the corresponding semantic label information; Wherein, a single context reasoning process combined with semantic alignment includes: Obtain the hidden state feature and visual attention feature of the current word; Based on the hidden state feature and the visual attention feature of the current word, correspondingly update the hidden state feature memory pool and the visual attention feature memory pool to respectively obtain the current hidden state feature sequence and the current visual attention feature sequence; Based on the current hidden state feature sequence and the current visual attention feature sequence, correspondingly obtain the hidden state semantic feature and visual semantic feature of the current word; perform semantic alignment on the hidden state semantic feature and the visual semantic feature to obtain the semantic alignment deviation of the current word; Based on the visual attention feature and semantic feature of the current word, obtain the semantic reasoning feature of the current word; specifically, the semantic feature of the current word includes the hidden state feature of the current word; input the visual attention feature and the hidden state feature of the current word into a gated linear unit to calculate and output the semantic reasoning feature of the current word; Based on the semantic reasoning feature of the current word, determine the semantic information with the highest degree of association in a preset semantic table as the semantic reasoning information of the current word.

2. The method for training an image description model based on context reasoning according to claim 1, wherein The implementation manner of obtaining the hidden state feature and visual attention feature of the current word includes: Obtain the hidden state feature corresponding to the previous word and obtain the word embedding feature of the current word; Based on the word embedding feature of the current word and the corresponding image region feature, and the hidden state feature corresponding to the previous word, obtain the hidden state feature of the current word; and, Based on the hidden state features of the current word and the corresponding image region features, an attention mechanism is used to obtain the visual attention features of the current word.

3. The method for training an image description model based on context reasoning according to claim 1, wherein Performing semantic alignment on the hidden state semantic features and the visual semantic features to obtain the semantic alignment deviation of the current word, including: Inputting the hidden state semantic features and the visual semantic features into a semantic alignment function to obtain the semantic alignment deviation of the current word, specifically: Among them, L aln is the semantic alignment deviation of the current word; Smooth-L1 is the distribution adjustment function; is the semantic feature of the hidden state after interaction enhancement; is the visual semantic feature after interaction enhancement.

4. The method for training an image description model based on context reasoning according to claim 2, wherein Based on the current hidden state feature sequence and the current visual attention feature sequence, obtaining the hidden state semantic features and visual semantic features of the current word, including: Based on the sequence features of the current visual attention feature sequence and the word embedding features of the current word, constructing a hidden state feature query vector for the current word; based on the sequence features of the current hidden state feature sequence and the word embedding features of the current word, constructing a visual attention feature query vector for the current word; Performing temporal and semantic enhancement on the current hidden state feature sequence to obtain an enhanced hidden state feature sequence; performing temporal and semantic enhancement on the current visual attention feature sequence to obtain an enhanced visual attention feature sequence; Based on the enhanced hidden state feature sequence and the hidden state feature query vector, using an attention mechanism to obtain the hidden state semantic features of the current word; based on the enhanced visual attention features and the visual attention feature query vector, using an attention mechanism to obtain the visual semantic features of the current word.

5. The method for training an image description model based on context reasoning according to claim 4, wherein The implementation manner of using an attention mechanism to obtain the hidden state semantic features of the current word based on the enhanced hidden state feature sequence and the hidden state feature query vector includes: Based on the enhanced hidden state feature sequence, using a self-attention mechanism to obtain an interaction-enhanced hidden state feature sequence, and based on the interaction-enhanced hidden state feature sequence and the hidden state feature query vector, using a traditional attention mechanism to obtain the hidden state semantic features; and, The implementation manner of using an attention mechanism to obtain the visual semantic features of the current word based on the enhanced visual attention features and the visual attention feature query vector includes: Based on the enhanced visual attention feature sequence, using a self-attention mechanism to obtain an interaction-enhanced visual attention feature sequence, and based on the interaction-enhanced visual attention feature sequence and the visual attention feature query vector, using a traditional attention mechanism to obtain the visual semantic features.

6. The method for training an image description model based on context reasoning according to claim 4, wherein Performing temporal and semantic enhancement on the current hidden state feature sequence includes: Based on the sequence positions of the hidden state features in the current hidden state feature sequence, using a first position encoder to obtain the temporal information of each hidden state feature; superimposing the temporal information of each hidden state feature and the word embedding features of the current word onto the corresponding hidden state feature; and, Performing temporal and semantic enhancement on the current visual attention feature sequence includes: Based on the sequence positions of the visual attention features in the current visual attention feature sequence, use a second position encoder to obtain the temporal information of each visual attention feature; superimpose the temporal information of each visual attention feature and the word embedding feature of the current word into the corresponding visual attention feature.

7. The method for training an image description model based on context reasoning according to claim 6, wherein The first position encoder and the second position encoder have the same structure, both including: Among them, pe() is a relative position encoder for obtaining the temporal information of each of the visual attention features; p i represents the relative position of the feature in the corresponding feature sequence; j is the dimension of the position encoding representation. When j is odd, f() is sin(), and when j is even, f() is cos(); d h is the feature dimension of the new region feature.

8. The method for training an image description model based on context reasoning according to claim 5, wherein, The semantic feature of the current word further includes a context reasoning feature; Then, the obtaining of the semantic reasoning feature of the current word based on the visual attention feature and semantic feature of the current word further includes: Obtain the context reasoning feature of the current word; Based on the visual attention feature of the current word and the context reasoning feature, obtain the semantic reasoning feature of the current word.

9. The method for training an image description model based on context reasoning according to claim 8, wherein The obtaining of the context reasoning feature of the current word includes: Based on the enhanced hidden state feature sequence and the enhanced visual attention feature sequence, obtain the multimodal feature sequence of the current word; Based on the hidden state query vector of the current word and the hidden state feature of the current word, construct the multimodal feature query vector of the current word; Based on the multimodal feature sequence and the multimodal feature query vector of the current word, use a traditional attention mechanism to obtain the context reasoning feature of the current word.

10. The method for training an image description model based on context reasoning according to claim 9, wherein The implementation manner of constructing the multimodal feature query vector of the current word includes: Use a gated linear unit, take the hidden state feature query vector of the current word and the hidden state feature of the current word as the inputs of the gated linear unit, and use the calculated output result as the multimodal feature query vector of the current word.

11. An image description method, characterized in that, Include: Construct each training data set based on the sample data of the image description; A single set of the training data includes image data and text data associated with the image data; Based on each training data set and the semantic label information corresponding to the training data, use the context reasoning-based image description model training method described in any one of claims 1 to 10 to train a preset model to obtain a trained image description model; For the image to be described, use the trained image description model to perform context semantic information reasoning to obtain the semantic information of the input image to be described.

12. An electronic device, characterized in that, Include: A processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the context reasoning-based image description model training method described in any one of claims 1 to 10 or the image description method described in claim 11.

13. A computer storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the context reasoning-based image description model training method described in any one of claims 1 to 10 or the image description method described in claim 11.

Citation Information

Patent Citations

  • Image description model training method based on modal enhancement and image description method

    CN115238118A