Image subtitle generation method, device, equipment and readable storage medium

By combining the image subtitle generation method with cyclic attention and self-attention mechanism, the problem of difficulty in capturing information and insufficient regional correlation in the prior art is solved, and a higher generation accuracy is achieved.

CN114550159BActive Publication Date: 2025-05-13CHINA UNIV OF PETROLEUM (BEIJING)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210188782.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-13
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

In the existing image subtitle generation method, information that is farther away is difficult to be captured. Only the correlation degree of the internal area of ​​the image is paid attention to, and the most related area cannot be identified, resulting in low generation accuracy.

Method used

Using a method combining cyclic attention mechanism and self-attention mechanism, image features are extracted through cyclic convolution neural networks, and the multi-head self-attention mechanism and fully connected feedforward network in the serialized variator model are used for encoding and decoding, and combining long and short-term memory networks and gated linear units of the sequence memory model for attention adjustment and output generation.

Benefits of technology

By identifying the most relevant areas, the accuracy of image subtitles is significantly improved, taking into account the advantages of cyclic attention and self-attention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550159B_ABST
    Figure CN114550159B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating image subtitles, comprising: extracting features from a target image using a recurrent convolutional neural network, obtaining an original image feature set and sending it to a serialized transformer encoder of a serialized transformer model for encoding; adjusting the attention of the original image feature set based on a target state variable of a long short-term memory network, inputting the attention-adjusted features and the target state variable into a gated linear unit to obtain a target memory; inputting the encoded image feature set, the current time step memory, and the current time step output word vector into a serialized transformer decoder, performing weighted summation on the obtained target decoder output and the target memory, and generating target subtitles according to the weighted summation results of each time step. The present invention can identify the most relevant area according to the external state, greatly improving the accuracy of image subtitle generation. The present invention also discloses a device, an apparatus, and a storage medium, which have corresponding technical effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image subtitle generation method, device, equipment and computer-readable storage medium. Background Art

[0002] Image captioning is used to extract natural language content from visual content, which is a subfield of sequence modeling tasks. The goal of this technology is to automatically add captions to images, which can be used for image search tasks later. In order to generate high-quality captions, it is necessary to understand different elements (objects, actions, scenes) and organize them with appropriate grammatical structures.

[0003] In the prior art, image caption generation is generally performed through the following two types of models: one is a model that does not include an attention mechanism, and generates image captions based on a template consisting of the output of object detection and attribute classification. However, this method is limited by the type of template. The other type is a model that includes an attention mechanism. Models that include an attention mechanism are further divided into models based on CNN+RNN and models based on Transformer.

[0004] The CNN+RNN based model uses convolutional neural network (CNN) and recurrent neural network (RNN) as encoder and decoder respectively. The encoder CNN extracts visual features from the input image, and the decoder RNN translates the visual features into corresponding words. The recurrent attention will selectively pay more attention to some image areas based on the state vector of the decoder RNN. However, the inherent sequential nature of recurrent neural networks excludes the possibility of parallel training. For two features that are weakly connected, the recurrent neural network needs to accumulate information through several time steps to connect them. Therefore, the farther the information is, the more difficult it is to capture, resulting in low accuracy in image caption generation.

[0005] The transformer-based model relies entirely on self-attention to calculate the representation of its input and output, without using sequence-aligned recurrent neural networks. Unlike recurrent attention, the self-attention in the transformer is a completely self-sustaining mechanism. The self-attention mechanism avoids the above-mentioned problems in the prior art to a certain extent by directly parallelizing the calculation of the correlation between the internal regions of the image. However, for image subtitle generation, only noting the correlation between the internal regions of the image cannot identify the most relevant regions, resulting in low accuracy in image subtitle generation.

[0006] In summary, how to effectively solve the problems in existing image subtitle generation methods, such as the farther the information is, the harder it is to capture, and only paying attention to the correlation of the internal areas of the image, failing to identify the most relevant areas, resulting in low accuracy of image subtitle generation, is an issue that technicians in this field urgently need to solve. Summary of the invention

[0007] The purpose of the present invention is to provide an image subtitle generation method, which takes into account the advantages of the recurrent attention mechanism and the self-attention mechanism, can identify the most relevant area according to the external state, and greatly improves the accuracy of image subtitle generation; another purpose of the present invention is to provide an image subtitle generation device, equipment and computer-readable storage medium.

[0008] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0009] A method for generating image captions, comprising:

[0010] Using a recurrent convolutional neural network to extract features from the received target image to be generated with subtitles, and obtaining a set of original image features;

[0011] Sending the original image feature set to a serialization transformer model, so as to use a serialization transformer encoder in the serialization transformer model to perform an encoding operation on the original image feature set to obtain an encoded image feature set;

[0012] Obtaining a current time step memory of a sequence memory model and a current time step output word vector of a serializer transformer decoder in the serializer transformer model;

[0013] Determine the target state variable of the long short-term memory network in the sequence memory model by combining the original image feature set, the current time step memory, and the current time step output word vector;

[0014] Using the sequence memory model to perform attention adjustment on the original image feature set based on the target state variable to obtain attention-adjusted features;

[0015] Inputting the attention-adjusted feature and the target state variable into a gated linear unit in the sequence memory model to obtain a target memory output by the sequence memory model;

[0016] Inputting the encoded image feature set, the current time step memory, and the current time step output word vector into a serializer decoder in the serializer model to obtain a target decoder output of the serializer decoder;

[0017] Performing a weighted summation on the target memory and the target decoder output, and using a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result;

[0018] Determine the word corresponding to the maximum value among the predicted probabilities as the target word, and determine the word vector of the target word as the target output word vector;

[0019] Determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector, and repeatedly perform the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network until a stop sign is detected;

[0020] The target words are combined in series to obtain a target caption corresponding to the target image.

[0021] In a specific embodiment of the present invention, the target state variable of the long short-term memory network in the sequence memory model is determined by combining the original image feature set, the current time step memory, and the current time step output word vector, including:

[0022] Calculating the average pooling of the original image feature set;

[0023] Performing a sum calculation on the average pooling and the current time step memory to obtain a fusion vector;

[0024] Performing vector concatenation on the word vector output at the current time step and the fusion vector to obtain a concatenated vector;

[0025] Inputting the concatenated vector into the long short-term memory network to obtain a target state variable of the long short-term memory network;

[0026] Repeating the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network includes:

[0027] Repeat the step of performing sum calculation on the average pooling and the current time step memory to obtain a fusion vector.

[0028] In a specific embodiment of the present invention, the encoding operation of the original image feature set is performed using the serialized transformer encoder in the serialized transformer model, including:

[0029] The original image feature set is encoded using the multi-head self-attention mechanism and the fully connected feed-forward network of the serialized transformer encoder.

[0030] In a specific embodiment of the present invention, the encoded image feature set, the current time step memory, and the current time step output word vector are input into the serializer transformer decoder in the serializer transformer model to obtain the target decoder output of the serializer transformer decoder, including:

[0031] The encoded image feature set, the current time step memory, and the current time step output word vector are input into a serializer transformer decoder including a multi-head self-attention mechanism and a fully connected feedforward network to obtain a target decoder output of the serializer transformer decoder.

[0032] An image caption generating device, comprising:

[0033] A feature extraction module is used to extract features of a received target image to be generated with subtitles using a recurrent convolutional neural network to obtain a feature set of an original image;

[0034] An encoding module, used for sending the original image feature set to a serialization transformer model, so as to use a serialization transformer encoder in the serialization transformer model to perform an encoding operation on the original image feature set to obtain an encoded image feature set;

[0035] A memory and word vector acquisition module, used to acquire the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model;

[0036] A state variable determination module, used to determine the target state variable of the long short-term memory network in the sequence memory model by combining the original image feature set, the current time step memory, and the current time step output word vector;

[0037] an attention adjustment module, configured to use the sequence memory model to perform attention adjustment on the original image feature set based on the target state variable to obtain attention-adjusted features;

[0038] A memory acquisition module, used for inputting the attention-adjusted feature and the target state variable into a gated linear unit in the sequence memory model to obtain a target memory output by the sequence memory model;

[0039] A decoder output acquisition module, used for inputting the encoded image feature set, the current time step memory and the current time step output word vector into the serializer decoder in the serializer model to obtain a target decoder output of the serializer decoder;

[0040] A prediction probability determination module, used for performing a weighted summation on the target memory and the target decoder output, and using a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result;

[0041] An output word vector determination module, used to determine the word corresponding to the maximum value among the predicted probabilities as a target word, and determine the word vector of the target word as a target output word vector;

[0042] A repeated execution module is used to determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector, and repeatedly execute the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network until a stop sign is detected;

[0043] The subtitle acquisition module is used to combine the target words in series to obtain the target subtitles corresponding to the target image.

[0044] In a specific implementation of the present invention, the state variable determination module includes:

[0045] An average pooling calculation submodule, used to calculate the average pooling of the original image feature set;

[0046] A fusion vector obtaining submodule is used to perform sum calculation on the average pooling and the current time step memory to obtain a fusion vector;

[0047] A vector splicing submodule, used for performing vector splicing on the word vector output at the current time step and the fusion vector to obtain a spliced ​​vector;

[0048] A state variable acquisition submodule, used for inputting the concatenated vector into the long short-term memory network to obtain a target state variable of the long short-term memory network;

[0049] The repeated execution module is specifically a module that repeatedly executes the step of performing sum calculation on the average pooling and the current time step memory to obtain a fusion vector.

[0050] In a specific embodiment of the present invention, the encoding module is specifically a module that utilizes the multi-head self-attention mechanism and the fully connected feed-forward network of the serialized transformer encoder to perform encoding operations on the original image feature set.

[0051] In a specific embodiment of the present invention, the decoder output acquisition module is specifically a module that inputs the encoded image feature set, the current time step memory, and the current time step output word vector into a serializer transformer decoder including a multi-head self-attention mechanism and a fully connected feedforward network to obtain the target decoder output of the serializer transformer decoder.

[0052] An image caption generating device, comprising:

[0053] Memory for storing computer programs;

[0054] A processor is used to implement the steps of the image subtitle generation method as described above when executing the computer program.

[0055] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the image subtitle generation method as described above are implemented.

[0056] The image subtitle generation method provided by the present invention uses a recurrent convolutional neural network to extract features from a received target image for generating subtitles, and obtains an original image feature set; sends the original image feature set to a serializer transformer model, so as to use a serializer transformer encoder in the serializer transformer model to perform encoding operations on the original image feature set, and obtain an encoded image feature set; obtains a current time step memory of the sequence memory model and a current time step output word vector of the serializer transformer decoder in the serializer transformer model; combines the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network in the sequence memory model; uses the sequence memory model to perform attention adjustment on the original image feature set based on the target state variable, and obtains attention-adjusted features; inputs the attention-adjusted features and the target state variable into a gated linear unit in the sequence memory model, and obtains the target memory output by the sequence memory model. memory; input the encoded image feature set, the current time step memory and the current time step output word vector into the serializer decoder in the serializer model to obtain the target decoder output of the serializer decoder; perform weighted summation on the target memory and the target decoder output, and use the classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result; determine the word corresponding to the maximum value in each prediction probability as the target word, and determine the word vector of the target word as the target output word vector; determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector, and repeatedly perform the steps of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network until a stop sign is detected; and combine the target words in series to obtain the target subtitle corresponding to the target image.

[0057] It can be seen from the above technical solution that an outward-inward attention mechanism for image subtitle generation is proposed. By deploying a long short-term memory network with outward-inward attention in a sequence memory model. Combined with the original image feature set, the current time step memory, and the current time step output word vector, the target state variable of the long short-term memory network in the sequence memory model is determined, and then the attention adjustment is performed on the target state variable and the original image feature set, and the result is sent to the gated linear unit for output. A serialized transformer model including a serialized transformer encoder and a decoder is pre-trained. For the serialized transformer encoder, the attention mechanism therein is a multi-head self-attention mechanism using self-attention. For the serialized transformer decoder, the output of the sequence memory model and the decoder output are weighted and summed according to a certain ratio to balance and adjust the final output. Compared with only considering a recurrent neural network or only considering the self-attention mechanism. The present invention takes into account the advantages of the recurrent attention mechanism and the self-attention mechanism, can identify the most relevant area according to the external state, and greatly improves the accuracy of image subtitle generation.

[0058] Correspondingly, the present invention also provides an image subtitle generation device, equipment and computer-readable storage medium corresponding to the above-mentioned image subtitle generation method, which have the above-mentioned technical effects and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0060] Figure 1 A flowchart of an implementation of the image subtitle generation method in an embodiment of the present invention;

[0061] Figure 2 Another implementation flow chart of the method for generating image subtitles in an embodiment of the present invention;

[0062] Figure 3 A diagram showing the structure of a self-attention mechanism applied to the image caption generation task;

[0063] Figure 4 A diagram showing the structure of an outside-in attention mechanism applied to the image captioning task;

[0064] Figure 5 A schematic diagram of the outward-inward attention mechanism;

[0065] Figure 6 is a structural block diagram of an image subtitle generating device according to an embodiment of the present invention;

[0066] Figure 7 is a structural block diagram of an image subtitle generating device in an embodiment of the present invention;

[0067] Figure 8 A schematic diagram of the specific structure of an image subtitle generating device provided in this embodiment. DETAILED DESCRIPTION

[0068] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0069] See also Figure 1 , Figure 1 This is a flowchart of an implementation of a method for generating image subtitles in an embodiment of the present invention. The method may include the following steps:

[0070] S101: Using a recurrent convolutional neural network to extract features of a received target image for generating subtitles, to obtain an original image feature set.

[0071] When a target image for which subtitles need to be generated is received, a recurrent convolutional neural network R-CNN is used to extract features of the received target image for which subtitles need to be generated, and obtain a set of original image features.

[0072] S102: Send the original image feature set to the serialization transformer model, so as to use the serialization transformer encoder in the serialization transformer model to perform an encoding operation on the original image feature set to obtain an encoded image feature set.

[0073] A serializer transformer model including a serializer transformer encoder is pre-trained. After the original image feature set of the target image is extracted using the recurrent convolutional neural network, the original image feature set is sent to the serializer transformer model, so that the serializer transformer encoder in the serializer transformer model performs an encoding operation on the original image feature set to obtain an encoded image feature set.

[0074] S103: Obtain the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model.

[0075] Pre-train the sequence memory model, and the serialization transformer model also includes a serialization transformer decoder. Get the current time step memory of the sequence memory model and the current time step output word vector of the serialization transformer decoder in the serialization transformer model.

[0076] S104: Determine the target state variable of the long short-term memory network in the sequence memory model by combining the original image feature set, the current time step memory, and the current time step output word vector.

[0077] After using the recurrent convolutional neural network to extract the original image feature set of the target image, and obtaining the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model, the target state variable of the long short-term memory network LSTM in the sequence memory model is determined by combining the original image feature set, the current time step memory, and the current time step output word vector.

[0078] S105: Using the sequence memory model to adjust the attention of the original image feature set based on the target state variable to obtain attention-adjusted features.

[0079] After determining the target state variables of the long short-term memory network in the sequence memory model, the sequence memory model is used to adjust the attention of the original image feature set based on the target state variables to obtain the attention-adjusted features.

[0080] S106: Input the attention-adjusted features and the target state variables into the gated linear unit in the sequence memory model to obtain the target memory output by the sequence memory model.

[0081] After attention adjustment is performed on the original image feature set to obtain the attention-adjusted features, the attention-adjusted features and the target state variables are input into the gated linear unit GLU in the sequence memory model to obtain the target memory output by the sequence memory model.

[0082] S107: Input the encoded image feature set, the current time step memory, and the current time step output word vector into the serializer transformer decoder in the serializer transformer model to obtain the target decoder output of the serializer transformer decoder.

[0083] Get the output word vector of the current time step. After obtaining the target memory output by the sequence memory model, input the encoded image feature set, the current time step memory and the output word vector of the current time step into the serializer transformer decoder in the serializer transformer model to obtain the target decoder output of the serializer transformer decoder.

[0084] S108: performing weighted summation on the target memory and the target decoder output, and using a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result.

[0085] After the target decoder outputs, the target memory and the target decoder output are weighted and summed, and a classifier is used to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result.

[0086] S109: Determine the word corresponding to the maximum value in each predicted probability as the target word, and determine the word vector of the target word as the target output word vector.

[0087] After using the classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted sum result, the prediction probabilities are sorted, and the largest prediction probability is selected according to the sorting result. The word corresponding to the maximum value in each prediction probability is determined as the target word, and the word vector of the target word is determined as the target output word vector.

[0088] S110: Determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector.

[0089] After the target output word vector is determined, the target state variable is determined as the current state variable, the target memory is determined as the current time step memory, and the target output word vector is determined as the current time step output word vector.

[0090] S111: Determine whether a stop sign is detected, if not, return to step S104, if yes, execute step S112.

[0091] A stop sign is set in advance. When the execution of this time step is completed, it is determined whether the stop sign is detected. If not, it means that the words constituting the target image subtitles have not been completely determined, and further word determination is required, and the process returns to step S104. If so, it means that the words constituting the target image subtitles have been completely determined, and step S112 is executed.

[0092] S112: Combine the target words in series to obtain target subtitles corresponding to the target image.

[0093] When the stop sign is detected, it means that the words constituting the subtitle of the target image have been completely determined, and the target words are combined in series to obtain the target subtitle corresponding to the target image.

[0094] It can be seen from the above technical solution that an outward-inward attention mechanism for image subtitle generation is proposed. By deploying a long short-term memory network with outward-inward attention in a sequence memory model. Combined with the original image feature set, the current time step memory, and the current time step output word vector, the target state variable of the long short-term memory network in the sequence memory model is determined, and then the attention adjustment is performed on the target state variable and the original image feature set, and the result is sent to the gated linear unit for output. A serialized transformer model including a serialized transformer encoder and a decoder is pre-trained. For the serialized transformer encoder, the attention mechanism therein is a multi-head self-attention mechanism using self-attention. For the serialized transformer decoder, the output of the sequence memory model and the decoder output are weighted and summed according to a certain ratio to balance and adjust the final output. Compared with only considering a recurrent neural network or only considering the self-attention mechanism. The present invention takes into account the advantages of the recurrent attention mechanism and the self-attention mechanism, can identify the most relevant area according to the external state, and greatly improves the accuracy of image subtitle generation.

[0095] It should be noted that, based on the above embodiment, the embodiment of the present invention also provides corresponding improved solutions. In the subsequent embodiments, the same steps or corresponding steps as those in the above embodiment can be referenced to each other, and the corresponding beneficial effects can also be referenced to each other, which will not be repeated one by one in the following improved embodiments.

[0096] See also Figure 2 , Figure 2 FIG. 4 is another implementation flow chart of the method for generating image subtitles in an embodiment of the present invention. The method may include the following steps:

[0097] S201: Using a recurrent convolutional neural network to extract features from a received target image for generating subtitles, to obtain an original image feature set.

[0098] The target image is feature extracted using a recurrent convolutional neural network to obtain the original image feature set, which can be expressed as A = {a1, a2, …, an}.

[0099] S202: Send the original image feature set to the serializer transformer model to utilize the multi-head self-attention mechanism and fully connected feed-forward network of the serializer transformer encoder to encode the original image feature set to obtain an encoded image feature set.

[0100] The multi-head self-attention mechanism of the serialized transformer encoder and the fully connected feed-forward network are used to encode the original image feature set to obtain the encoded image feature set. The encoded image feature set can be expressed as N .

[0101] The serialized transformer encoder is composed of N = 6 blocks cascaded together, each of which contains a multi-head self-attention mechanism and a fully connected feed-forward network, as follows:

[0102]

[0103]

[0104]

[0105] AddNorm represents the residual connection layer normalization operation. In the multi-head attention sublayer, the key vector (key), query vector (query), and value vector (value) in A are converted to the same dimension. Then, the output of the encoder is normalized. It is fed into a fully connected feed-forward network consisting of two linear transformations with a ReLu in between to apply the position information. and are the parameters of the linear transformation that can be learned and optimized, After passing through the linear layer, the activation function is performed and then passed through another linear layer to obtain the intermediate result. represents the output of the (l+1)th block. Because the encoder has a cyclic structure, the input of the current block is the output of the previous block. After N blocks are cascaded, the output of the last block (the Nth block) is O N , which is also the output of the entire encoder. And define the input O of the first block 0 =A+PE(A), where A is the image feature set extracted by the recurrent convolutional neural network, and PE is the position encoding information of the image. The position encoding information PE (Position Embedding) of the image is described as follows:

[0106] The image is divided into patches of size K*K (grid division, K=14 in the present invention), and then the image grid of size [K*K] is stretched into a one-dimensional sequence of length [K*K], and position encoding is used for this sequence, that is:

[0107]

[0108] Where F is the feature dimension of the image, and pos_img is the corresponding sequence number of the patch where each pixel is located in the one-dimensional sequence described above. a can be an integer multiple of 2 in the position code, and 2 can be selected.

[0109] The self-attention mechanism is also generally called scaled dot product attention in the transformer. In theory, an attention module can be described as mapping a query vector and a set of key-value pairs to the output space. Self-attention is a special type of attention that is used to obtain the internal connections of the input sequence. In self-attention, Q (query), K (key), and V (value) all come from the same input. It first measures Q = {q1,…,q n} and K={k1,…,k m}, and use the similarity to calculate V = {v1,…,v m The weighted average sum of} can be expressed as:

[0110]

[0111] definition And d is the scaling factor.

[0112] In order to allow the model to pay attention to different subspaces at the same time, the translator (Transformer) uses a multi-head attention mechanism (self-attention is implemented h times in parallel). On each head, Q (query), K (key), and V (value) are linearly projected onto d k ,d k ,d v dimensional space. It can be expressed as:

[0113] MultiHead(Q,K,V)=Concat(H1,...,H h )W O ;

[0114]

[0115] in, is the head projection matrix, is the linear transformation matrix, and d m =d k *h is the initialization dimension of keys, values, and queries, d k is the dimension of the single-head attention Key. When multi-head attention is applied and the number of heads is h, the dimension becomes d k *h.

[0116] Two different multi-head attention structures are defined. The two structures are defined as follows:

[0117] Υ=multiheadS (κ);

[0118] Υ=multihead O (κ);

[0119] The first is a multi-head self-attention mechanism structure (Multi-head) in which each head is a self-attention mechanism, and the second is a multi-head outside-inAttention mechanism in which each head is an outside-in attention mechanism (Multi-head Outside-inAttention). κ represents input and Υ represents output.

[0120] S203: Obtain the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model.

[0121] Get the current time step memory of the sequence memory model The current time step output word vector S of the serializer transformer decoder in the serializer transformer model t-1 .

[0122] S204: Calculate the average pooling of the original image feature set.

[0123] Calculate the average pooling of the original image feature set. The average pooling of the original image feature set can be expressed as

[0124] S205: performing sum calculation on the average pooling and the current time step memory to obtain a fusion vector.

[0125] After calculating the average pooling of the original image feature set And get the current time step memory of the sequence memory model After that, the average pooling and the current time step memory are summed to obtain the fusion vector

[0126] S206: Perform vector concatenation on the word vector output at the current time step and the fusion vector to obtain a concatenated vector.

[0127] After obtaining the current time step output word vector of the serialized transformer decoder in the serialized transformer model and calculating the fusion vector, output the word vector S for the current time step t-1 and fusion vector Perform vector concatenation to obtain a concatenated vector. Vector concatenation can be performed using the following formula:

[0128]

[0129] Thus, the concatenated vector is obtained.

[0130] S207: Input the concatenated vector into the long short-term memory network to obtain the target state variable of the long short-term memory network.

[0131] After obtaining the concatenated vector, the concatenated vector is input into the long short-term memory network to obtain the target state variable h of the long short-term memory network. t The target state variable can be calculated using the following formula:

[0132] h t =LSTM(x t ,h t-1 );

[0133] Thus the target state variable is obtained.

[0134] Transformers are not sequential. Therefore, most methods add a "position encoding" module to enable the model to utilize the sequential information of the sequence. However, this method still has shortcomings compared to recurrent neural networks. Therefore, a long short-term memory network (LSTM) layer is added to the proposed serialization transformer structure to enhance the model's sequence modeling capabilities.

[0135] The long short-term memory network model is a type of time series model. It can improve a function of the ordinary time series model, that is, when the input sequence or text is too long, it can have a longer memory, that is, long-term dependency. A long short-term memory network is composed of a long string of gates. They are input gate (current cell state), forget gate (0: forget all previous ones; 1: pass all previous ones), output gate (select output), New memory cell (get new memory cell). The four different gates cooperate and inhibit each other to make the whole model work. Among them, the input gate is mainly used to complete the work of the input interface, the forget gate is mainly used to control the legacy of the model's judgment information, the output gate mainly controls the output of the model, and the New memory cell is the "brain" of the whole model, which can control the operation of the whole model.

[0136] S208: Using the sequence memory model to adjust the attention of the original image feature set based on the target state variable to obtain attention-adjusted features.

[0137] After obtaining the target state variable h of the long short-term memory network t Afterwards, the sequence memory model is used to adjust the attention of the original image feature set based on the target state variable to obtain the attention-adjusted feature set. It can be obtained by the following formula Send to the Multi-head Outside-in sublayer:

[0138]

[0139] This results in a more representative attention-adjusted feature.

[0140] S209: Input the attention-adjusted features and the target state variables into the gated linear unit in the sequence memory model to obtain the target memory output by the sequence memory model.

[0141] After obtaining the attention-adjusted features, the attention-adjusted features and the target state variable h t Input the gated linear unit in the sequence memory model to obtain the target memory output by the sequence memory model The attention-adjusted features and target state variables can be input into the gated linear unit in the sequence memory model by the following formula:

[0142]

[0143] Thus, the target memory output by the sequence memory model is obtained

[0144] S210: Input the encoded image feature set, the current time step memory, and the current time step output word vector into a serializer decoder including a multi-head self-attention mechanism and a fully connected feedforward network to obtain a target decoder output of the serializer decoder.

[0145] The serialization transformer decoder contains a multi-head self-attention mechanism and a fully connected feed-forward network. N , Current time step memory And the output word vector S of the current time step t-1 Input to the serialization transformer decoder to get the target decoder output of the serialization transformer decoder

[0146] Similarly, the serialized transformer decoder can also be formed by cascading N blocks (such as N=6), and each cycle of l in the following formula is regarded as passing through one block:

[0147]

[0148] for l in range(0,N):

[0149]

[0150]

[0151]

[0152]

[0153] Among them, S <t ={s0,s1,…,s t-1} is the feature vector set of the entire model historical output sequence, AddNorm is to perform residual connection and layer normalization on the data, FFN is to perform linear transformation and nonlinear activation on the data in sequence (the nonlinear activation function can use tanh function), Multihead S That is, each head is a multi-head attention (Multi-head Attention) with self-attention.

[0154] For sequence S <t Use positional encoding, that is:

[0155]

[0156] Among them, a is the sequence feature dimension, pos_seq is a one-dimensional sequence S for each word <t The corresponding sequence number within.

[0157] The input of the l+1th block at the tth time step includes: the output of the encoder O N , the output sequence of the sequence memory module The output sequence of the decoder ( That is, the set of feature vectors of the entire historical output sequence of the model).

[0158] S211: performing weighted summation on the target memory and the target decoder output, and using a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result.

[0159] Get the target decoder output of the serialized transformer decoder Afterwards, the target memory and the target decoder output A weighted sum is performed, and a classifier is used to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted sum result.

[0160] The overall output of the decoder at the tth time step is the output of the last block, that is, the output of the Nth block Compare it with the current time step output of the sequence memory module The data is fused according to the ratio of 1-β to β to obtain the matrix ξ t , and send it to the linear layer + softmax to calculate the probability of each possible word at the tth time step. The formula is as follows:

[0161]

[0162] p t =softmax(ξ t W p );

[0163] Regarding the above processes and variables, and It is the feature vector of the word (token) at the beginning of the sentence. in and Softmax represents the logistic regression classification function. A represents the features extracted by the convolutional neural network VGG model, and its dimension is [2048,49]. The final result obtained by weighting is a linearly variable parameter that can be learned and optimized, and |Ξ| is the size of the dictionary.

[0164] S212: Determine the word corresponding to the maximum value in each predicted probability as the target word, and determine the word vector of the target word as the target output word vector.

[0165] S213: Determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector.

[0166] S214: Determine whether a stop sign is detected, if not, return to step S205, if yes, execute step S215.

[0167] S215: Combine the target words in series to obtain target subtitles corresponding to the target image.

[0168] This embodiment is different from the embodiment 1 corresponding to the technical solution claimed for protection by the independent claim 1, and further adds the technical solution claimed for protection by the dependent claims 2 to 4. Of course, according to different actual conditions and requirements, the technical solutions claimed for protection by the dependent claims can be flexibly combined without affecting the integrity of the solutions, so as to better meet the requirements of different usage scenarios. This embodiment only provides one of the solutions with the most solutions and the best effects. Due to the complexity of the situation, it is impossible to list all possible solutions one by one. Those skilled in the art should be able to realize that there can be many examples based on the basic method principles provided by this application in combination with actual conditions, and they should all be within the scope of protection of this application without sufficient creative work.

[0169] See also Figure 3 , Figure 3 A diagram showing the structure of a self-attention mechanism applied to the image caption generation task. Figure 3 As can be seen in the figure, the self-attention mechanism can only obtain associations within the image area. Its operation is to generate three vector sets of query vectors, key vectors, and value vectors in the image, and calculate the weighted sum of the value vectors according to the calculated similarity distribution of the query vectors and key vectors.

[0170] See also Figure 4 , Figure 4 A diagram showing the structure of an out-to-in attention mechanism applied to the image captioning task. Figure 4 As can be seen in , the outside-in attention mechanism treats the external state as an additional query vector. Figure 3 The difference between the method and the self-attention method is that it can not only obtain the correlation between the internal areas of the image, but also the correlation between the internal areas of the image and the external state. In this way, the model can obtain the most relevant areas at each time step like self-attention.

[0171] See also Figure 5 , Figure 5 This is a schematic diagram of the outward-inward attention mechanism. The inward-inward attention mechanism is the Multihead in the sequence memorization module. O The sublayer can be expressed as follows:

[0172] (1) Perform self-attention on the image to obtain the query vector (Queries), key vector (Keys), and value vector (Values) corresponding to the current attention sub-layer.

[0173] (2) Take the inner product of the query vector (Queries) and the key vector (Keys) to get the attention image matrix AttentionMap of attention α 1.

[0174] (3) Attention Map 1 Do a dot product with the image value vector (Values) to get the image feature X after attention adjustment.

[0175] (4) Assume that the external state variable h of the long short-term memory network in the sequence memory model at the previous time step is t-1 For Query.

[0176] (5) After adding X and Query, the activation function (Softmax) is normalized to obtain the attention of each pixel The matrix AttentionMap 2 .

[0177] (6) Attention Map 2 Do a dot product with X to get the feature matrix C.

[0178] (7) Repeat 6 times and concatenate the results and transform them linearly into

[0179] In this way, the external state h can be taken into account at the same time t-1 As well as the connections between internal regions, it inherits the advantages of multi-head attention and recurrent attention.

[0180] Corresponding to the above method embodiment, the present invention further provides an image subtitle generating device. The image subtitle generating device described below and the image subtitle generating method described above can be referred to each other.

[0181] See also Figure 6 , Figure 6 : is a structural block diagram of an image subtitle generating device in an embodiment of the present invention, and the device may include:

[0182] The feature extraction module 601 is used to extract features of the received target image to be generated with subtitles by using a recurrent convolutional neural network to obtain a set of original image features;

[0183] The encoding module 602 is used to send the original image feature set to the serialization transformer model, so as to use the serialization transformer encoder in the serialization transformer model to perform an encoding operation on the original image feature set to obtain an encoded image feature set;

[0184] The memory and word vector acquisition module 603 is used to acquire the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model;

[0185] The state variable determination module 604 is used to determine the target state variable of the long short-term memory network in the sequence memory model by combining the original image feature set, the current time step memory, and the current time step output word vector;

[0186] An attention adjustment module 605 is used to adjust the attention of the original image feature set based on the target state variable using a sequence memory model to obtain attention-adjusted features;

[0187] A memory acquisition module 606 is used to input the attention-adjusted features and the target state variable into the gated linear unit in the sequence memory model to obtain the target memory output by the sequence memory model;

[0188] The decoder output obtaining module 607 is used to input the encoded image feature set, the current time step memory and the current time step output word vector into the serializer decoder in the serializer model to obtain the target decoder output of the serializer decoder;

[0189] A prediction probability determination module 608 is used to perform a weighted summation on the target memory and the target decoder output, and use a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result;

[0190] An output word vector determination module 609 is used to determine the word corresponding to the maximum value among the predicted probabilities as the target word, and determine the word vector of the target word as the target output word vector;

[0191] Repeating the execution module 610, which is used to determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector, and repeatedly perform the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network until a stop sign is detected;

[0192] The subtitle obtaining module 611 is used to combine the target words in series to obtain the target subtitles corresponding to the target image.

[0193] It can be seen from the above technical solution that an outward-inward attention mechanism for image subtitle generation is proposed. By deploying a long short-term memory network with outward-inward attention in a sequence memory model. Combined with the original image feature set, the current time step memory, and the current time step output word vector, the target state variable of the long short-term memory network in the sequence memory model is determined, and then the attention adjustment is performed on the target state variable and the original image feature set, and the result is sent to the gated linear unit for output. A serialized transformer model including a serialized transformer encoder and a decoder is pre-trained. For the serialized transformer encoder, the attention mechanism therein is a multi-head self-attention mechanism using self-attention. For the serialized transformer decoder, the output of the sequence memory model and the decoder output are weighted and summed according to a certain ratio to balance and adjust the final output. Compared with only considering a recurrent neural network or only considering the self-attention mechanism. The present invention combines the advantages of the recurrent attention mechanism and the self-attention mechanism, can identify the most relevant area according to the external state, and greatly improves the accuracy of image subtitle generation.

[0194] In a specific implementation of the present invention, the state variable determination module 604 includes:

[0195] The average pooling calculation submodule is used to calculate the average pooling of the original image feature set;

[0196] The fusion vector acquisition submodule is used to sum the average pooling and the current time step memory to obtain the fusion vector;

[0197] The vector concatenation submodule is used to concatenate the word vector and the fusion vector output at the current time step to obtain a concatenated vector;

[0198] A state variable acquisition submodule is used to input the concatenated vector into the long short-term memory network to obtain the target state variable of the long short-term memory network;

[0199] The repeated execution module is specifically a module for repeatedly executing the step of summing the average pooling and the current time step memory to obtain the fusion vector.

[0200] In a specific embodiment of the present invention, the encoding module 602 is specifically a module that utilizes the multi-head self-attention mechanism of the serialized transformer encoder and the fully connected feed-forward network to perform encoding operations on the original image feature set.

[0201] In a specific embodiment of the present invention, the decoder output acquisition module 607 is specifically a module that inputs the encoded image feature set, the current time step memory, and the current time step output word vector into a serializer transformer decoder including a multi-head self-attention mechanism and a fully connected feedforward network to obtain the target decoder output of the serializer transformer decoder.

[0202] Corresponding to the above method embodiment, see Figure 7 , Figure 7 This is a schematic diagram of an image subtitle generating device provided by the present invention, and the device may include:

[0203] A memory 332, for storing computer programs;

[0204] The processor 322 is configured to implement the steps of the image subtitle generation method of the above method embodiment when executing a computer program.

[0205] For details, please refer to Figure 8 , Figure 8 A specific structural diagram of an image subtitle generating device provided in this embodiment, which may have relatively large differences due to different configurations or performances, may include a processor (central processing units, CPU) 322 (for example, one or more processors) and a memory 332, wherein the memory 332 stores one or more computer applications 342 or data 344. The memory 332 may be a temporary storage or a permanent storage. The program stored in the memory 332 may include one or more modules (not shown in the figure), each of which may include a series of instruction operations in the data processing device. Furthermore, the processor 322 may be configured to communicate with the memory 332, and execute a series of instruction operations in the memory 332 on the image subtitle generating device 301.

[0206] The image subtitle generating device 301 may further include one or more power supplies 326 , one or more wired or wireless network interfaces 350 , one or more input and output interfaces 358 , and / or one or more operating systems 341 .

[0207] The steps in the image subtitle generation method described above can be implemented by the structure of the image subtitle generation device.

[0208] Corresponding to the above method embodiment, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps can be implemented:

[0209] A recurrent convolutional neural network is used to extract features of a received target image to be generated for subtitles, and an original image feature set is obtained; the original image feature set is sent to a serializer transformer model, and the serializer transformer encoder in the serializer transformer model is used to encode the original image feature set, and an encoded image feature set is obtained; the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model are obtained; the target state variable of the long short-term memory network in the sequence memory model is determined by combining the original image feature set, the current time step memory, and the current time step output word vector; the sequence memory model is used to adjust the attention of the original image feature set based on the target state variable, and an attention-adjusted feature is obtained; the attention-adjusted feature and the target state variable are input into a gated linear unit in the sequence memory model, and a target memory output by the sequence memory model is obtained; the encoded image The image feature set, the current time step memory and the current time step output word vector are input to the serialized transformer decoder in the serialized transformer model to obtain the target decoder output of the serialized transformer decoder; the target memory and the target decoder output are weighted summed, and the classifier is used to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result; the word corresponding to the maximum value in each prediction probability is determined as the target word, and the word vector of the target word is determined as the target output word vector; the target state variable is determined as the current state variable, the target memory is determined as the current time step memory, and the target output word vector is determined as the current time step output word vector, and the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network is repeated until a stop sign is detected; the target words are combined in series to obtain the target subtitle corresponding to the target image.

[0210] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0211] For an introduction to the computer-readable storage medium provided by the present invention, please refer to the above method embodiment, and the present invention will not be elaborated here.

[0212] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the devices, equipment and computer-readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0213] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the technical solution and core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, the present invention can also be improved and modified without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A method for generating image subtitles, characterized in that: include: Using a recurrent convolutional neural network to extract features from the received target image to be generated with subtitles, and obtaining a set of original image features; Sending the original image feature set to a serialization transformer model, so as to use a serialization transformer encoder in the serialization transformer model to perform an encoding operation on the original image feature set to obtain an encoded image feature set; Obtaining a current time step memory of a sequence memory model and a current time step output word vector of a serializer transformer decoder in the serializer transformer model; Determine the target state variable of the long short-term memory network in the sequence memory model by combining the original image feature set, the current time step memory, and the current time step output word vector; Using the sequence memory model to perform attention adjustment on the original image feature set based on the target state variable to obtain attention-adjusted features; Inputting the attention-adjusted feature and the target state variable into a gated linear unit in the sequence memory model to obtain a target memory output by the sequence memory model; Inputting the encoded image feature set, the current time step memory, and the current time step output word vector into a serializer decoder in the serializer model to obtain a target decoder output of the serializer decoder; Performing a weighted summation on the target memory and the target decoder output, and using a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result; Determine the word corresponding to the maximum value among the predicted probabilities as the target word, and determine the word vector of the target word as the target output word vector; Determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector, and repeatedly perform the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network until a stop sign is detected; The target words are combined in series to obtain a target caption corresponding to the target image.

2. The image subtitle generation method according to claim 1, characterized in that: Combining the original image feature set, the current time step memory, and the current time step output word vector, determining the target state variable of the long short-term memory network in the sequence memory model includes: Calculating the average pooling of the original image feature set; Performing a sum calculation on the average pooling and the current time step memory to obtain a fusion vector; Performing vector concatenation on the word vector output at the current time step and the fusion vector to obtain a concatenated vector; Inputting the concatenated vector into the long short-term memory network to obtain a target state variable of the long short-term memory network; Repeating the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network includes: Repeat the step of performing sum calculation on the average pooling and the current time step memory to obtain a fusion vector.

3. The image subtitle generation method according to claim 1, characterized in that: The serialization transformer encoder in the serialization transformer model is used to perform an encoding operation on the original image feature set, including: The original image feature set is encoded using the multi-head self-attention mechanism and the fully connected feed-forward network of the serialized transformer encoder.

4. The image subtitle generation method according to claim 1, characterized in that: Inputting the encoded image feature set, the current time step memory, and the current time step output word vector into the serializer transformer decoder in the serializer transformer model to obtain a target decoder output of the serializer transformer decoder, including: The encoded image feature set, the current time step memory, and the current time step output word vector are input into a serializer transformer decoder including a multi-head self-attention mechanism and a fully connected feedforward network to obtain a target decoder output of the serializer transformer decoder.

5. An image subtitle generating device, characterized in that: include: A feature extraction module is used to extract features of a received target image to be generated with subtitles using a recurrent convolutional neural network to obtain a feature set of an original image; An encoding module, used for sending the original image feature set to a serialization transformer model, so as to use a serialization transformer encoder in the serialization transformer model to perform an encoding operation on the original image feature set to obtain an encoded image feature set; A memory and word vector acquisition module, used to acquire the current time step memory of the sequence memory model and the current time step output word vector of the serializer transformer decoder in the serializer transformer model; A state variable determination module, used to determine the target state variable of the long short-term memory network in the sequence memory model by combining the original image feature set, the current time step memory, and the current time step output word vector; an attention adjustment module, configured to use the sequence memory model to perform attention adjustment on the original image feature set based on the target state variable to obtain attention-adjusted features; A memory acquisition module, used for inputting the attention-adjusted feature and the target state variable into a gated linear unit in the sequence memory model to obtain a target memory output by the sequence memory model; A decoder output acquisition module, used for inputting the encoded image feature set, the current time step memory and the current time step output word vector into the serializer decoder in the serializer model to obtain a target decoder output of the serializer decoder; A prediction probability determination module, used for performing a weighted summation on the target memory and the target decoder output, and using a classifier to determine the prediction probability corresponding to each word in the preset dictionary according to the weighted summation result; An output word vector determination module, used to determine the word corresponding to the maximum value among the predicted probabilities as a target word, and determine the word vector of the target word as a target output word vector; A repeated execution module is used to determine the target state variable as the current state variable, determine the target memory as the current time step memory, and determine the target output word vector as the current time step output word vector, and repeatedly execute the step of combining the original image feature set, the current time step memory, and the current time step output word vector to determine the target state variable of the long short-term memory network until a stop sign is detected; The subtitle acquisition module is used to combine the target words in series to obtain target subtitles corresponding to the target image.

6. The image subtitle generating device according to claim 5, characterized in that: The state variable determination module comprises: An average pooling calculation submodule, used to calculate the average pooling of the original image feature set; A fusion vector obtaining submodule is used to perform sum calculation on the average pooling and the current time step memory to obtain a fusion vector; A vector splicing submodule, used for performing vector splicing on the word vector output at the current time step and the fusion vector to obtain a spliced ​​vector; A state variable acquisition submodule, used for inputting the concatenated vector into the long short-term memory network to obtain a target state variable of the long short-term memory network; The repeated execution module is specifically a module that repeatedly executes the step of performing sum calculation on the average pooling and the current time step memory to obtain a fusion vector.

7. The image subtitle generating device according to claim 5, characterized in that: The encoding module is specifically a module that utilizes the multi-head self-attention mechanism and the fully connected feed-forward network of the serialized transformer encoder to perform encoding operations on the original image feature set.

8. The image subtitle generating device according to claim 5, characterized in that: The decoder output acquisition module is specifically a module that inputs the encoded image feature set, the current time step memory, and the current time step output word vector into a serializer transformer decoder including a multi-head self-attention mechanism and a fully connected feedforward network to obtain the target decoder output of the serializer transformer decoder.

9. An image subtitle generation device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the image subtitle generation method according to any one of claims 1 to 4 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the image subtitle generation method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • A character recognition method based on a gating cascade attention mechanism

    CN109919174A

  • An image description generation method and device based on a deep residual network and attention

    CN109948691A