Image description generation system, training method, generation method and electronic equipment
By using Transformer decoder and directed acyclic graph in the image description generation system, the problems of high inference delay and low description quality in the prior art are solved, and efficient and smooth image text description generation is achieved.
Patent Information
- Application Number
- CN202510357550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has high inference delay and high computational overhead in image description generation, especially in low efficiency when generating longer texts. Non-autoregressive decoders have challenges in description quality, and the generated description is not smooth and cannot generate high-quality image text descriptions.
The image description generation system is adopted, including an image encoder, an embedded mapping layer, a Transformer decoder, a path selection module and a linear classifier. Through cross-modal semantic calculation and the construction of directed acyclic graphs, non-autoregressive decoding attributes are realized, and the inference speed and the fluency of generating descriptions are improved.
It realizes the generation of high-quality image text descriptions at a faster speed, improves the inference speed and description fluency, and enhances the efficiency and scalability of the image description generation system.
Smart Images

Figure CN120219769A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image description, and more specifically, relates to an image description generation system, a training method, a generation method and an electronic device. Background Art
[0002] In today's digital age, image understanding and description technologies have become research hotspots in the fields of computer vision and artificial intelligence, and have broad application prospects. Image description generation aims to enable a computer to automatically generate natural language descriptions to accurately interpret the content contained in an image. This technology has shown great practical value in many fields such as image search engines, military intelligence generation, and disaster warning and assessment.
[0003] In recent years, the "encoder-decoder" framework in deep learning has been widely used in the field of image description generation. Among them, the encoder extracts visual semantic features of an image, and the decoder receives the visual semantic features of the image to generate a target description.
[0004] The decoder can be divided into the following two types in terms of generation methods: autoregressive decoding and non-autoregressive decoding.
[0005] The autoregressive decoder generates a description word by word through the output generated historically and the visual semantic features, and can capture sequence dependencies and context information. The generated text is usually more coherent. Therefore, although this method has advantages in generating high-quality descriptions, its inference latency is high and the computational cost is large. Especially when generating longer texts, the inference process is particularly slow, affecting its efficiency and scalability in emergency and large-scale application scenarios.
[0006] Different from the autoregressive decoder, the non-autoregressive decoder receives the visual semantic feature vector and generates a complete sentence at one time, significantly reducing the inference latency. However, the non-autoregressive method still faces challenges in the quality of the generated description. Specifically, since the non-autoregressive model independently predicts each word during training, and each image has multiple description labels, then for the same position in the output, the model may see multiple labels. In this case, the non-autoregressive model may learn a result that mixes multiple labels, generating illogical descriptions. And during inference, since all words are generated simultaneously without considering the order relationship between words, it may lead to the generation of unsmooth descriptions and the inability to generate high-quality image text descriptions. Summary of the Invention
[0007] Aiming at the above defects or improvement requirements of the prior art, the present invention provides an image description generation system, a training method, a generation method and an electronic device, and its purpose is to generate high-quality image text descriptions at a relatively fast speed.
[0008] To achieve the above object, in a first aspect, the present invention provides an image description generation system, including:
[0009] An image encoder, configured to divide an image to be described into image patches, extract visual features of each image patch, and further obtain a visual feature sequence;
[0010] An embedding mapping layer, configured to map and encode the visual feature sequence into a corresponding semantic information sequence by using a visual encoder of a vision-language model;
[0011] A Transformer decoder, including a plurality of cascaded decoding blocks, configured to use the semantic information sequence as the input of the multi-head attention module in the first decoding block, and use the visual feature sequence as the K value and V value of the cross-attention module in each decoding block to implement cross-modal semantic calculation between the semantic information sequence and the visual feature sequence, so as to obtain intermediate hidden states of each candidate word;
[0012] A path selection module, configured to calculate a transition probability matrix based on an intermediate hidden vector formed by intermediate hidden states of all candidate words; perform a lower triangular mask on the state transition matrix, and construct a corresponding directed acyclic graph based on the masked state transition matrix; select candidate paths with a path length of a preset length from the directed acyclic graph, and use the candidate path with the maximum probability as the optimal path; wherein, the state transition matrix includes: the transition probability between intermediate hidden states of any two candidate words; the nodes in the directed acyclic graph are intermediate hidden states of candidate words, and the weight of an edge is the transition probability between two nodes on the edge; the probability of a candidate path is the product of the transition probabilities between all adjacent two nodes on the candidate path;
[0013] A linear classifier, configured to map each intermediate hidden state on the optimal path to a corresponding word to obtain an image text description.
[0014] Further preferably, the above image encoder includes:
[0015] A feature extraction module, configured to extract image features of N different scales of an image patch; sequentially denote each image feature as a first image feature to an Nth image feature in descending order of scale; N≥2;
[0016] A feature fusion module, configured to align and add the first image feature and the second image feature to obtain a first fused image feature; align and add the ith fused image feature and the (i + 2)th image feature to obtain the (i + 1)th fused image feature; i = 1, 2,..., N - 2; use the (N - 1)th fused image feature as the visual feature of the image patch;
[0017] Wherein, the feature fusion module aligns two image features of different scales in the following manner:
[0018] For the large-scale image feature B and the small-scale image feature S among the image features at two different scales respectively, perform channel dimension transformation on the large-scale image feature B to make it consistent with the channel dimension of the small-scale image feature S; perform upsampling on the small-scale image feature S to make it consistent with the scale of the large-scale image feature B.
[0019] Further preferably, the transition probability matrix is:
[0020]
[0021] Wherein, H is the intermediate hidden vector formed by the intermediate hidden states of all candidate words; d is the length of the intermediate hidden vector; W Q and W K are both pre-trained learnable parameters.
[0022] Further preferably, the above linear classifier includes a cascaded linear layer and a softmax layer;
[0023] After each intermediate hidden state on the optimal path undergoes a linear transformation through the linear layer, it is input into the softmax layer to obtain the probabilities that each intermediate hidden state on the optimal path is respectively mapped to different words in the preset vocabulary set, and the word corresponding to the maximum probability is used as the word corresponding to this intermediate hidden state.
[0024] In a second aspect, the present invention provides a training method for the above image description generation system, including:
[0025] Input each image sample in the pre-collected training set into the above image description generation system to obtain the corresponding image text description; wherein, the training set includes: image samples and one or more corresponding text description labels;
[0026] For each image sample in the training set, calculate the difference loss between its image text description and each corresponding text description label respectively, and then sum them to obtain the corresponding image description loss; construct a first training objective with the goal of minimizing the sum of the image description losses of all image samples;
[0027] Based on the total training objective including the first training objective, train the image description generation system.
[0028] Further preferably, the above training method further includes:
[0029] For each image sample in the training set, calculate the difference loss between its global semantic feature and the corresponding text embedding feature as the corresponding global alignment loss; construct a second training objective with the goal of minimizing the sum of the global alignment losses of all image samples; wherein, the global semantic feature of the image sample is the feature obtained by averaging the semantic information sequence of the image sample; the semantic information sequence is the sequence output by the embedding mapping layer in the image caption generation system when the image sample is input into the image caption generation system; the text embedding feature is the average of the features obtained by passing each text description label of the image sample through the language encoder in the vision-language model;
[0030] The above total training objective further includes: a second training objective; while training the image caption generation system based on the total training objective, the parameters in the above language encoder are also adjusted.
[0031] Further preferably, the above training method further includes:
[0032] For each image sample in the training set, calculate the corresponding word-level value function; wherein, the word-level value function of the k-th image sample is:
[0033]
[0034] Denote the expectation of logf k,n,h over all e L (e k,n,j ,W2); n = 1, 2, …, N k ; j = 1, 2, …, L k,n ; N k is the number of text description labels of the k-th image sample; L k,n is the number of words in the n-th text description label of the k-th image sample; e k,n,j is the j-th word in the n-th text description label of the k-th image sample; f L (e k,n,j ,W2) is the feature obtained by passing e k,n,j through the language encoder in the vision-language model; W2 is the learnable parameter in the language encoder of the vision-language model; Denote the expectation of log(1 - f k,j' ) over all s L (s k,j' ,W2); j' = 1, 2, …, b k ; b k is the number of semantic information in the semantic information sequence of the k-th image sample; s k,j'is the j'-th semantic information in the semantic information sequence of the k-th image sample; the semantic information sequence is the sequence output by the embedding mapping layer in the image description generation system when the image sample is input into the image description generation system; f L (s k,j' , W2) is the feature obtained by s k,j' through the language encoder in the vision-language model;
[0035] Construct a third training objective with as the target; K is the number of image samples in the training set; W1 is the learnable parameter of the vision encoder in the embedding mapping layer;
[0036] The above total training objective further includes: a third training objective; while training the image description generation system based on the total training objective, the parameters in the above language encoder are also adjusted.
[0037] In a third aspect, the present invention provides an image description generation method, including: inputting the image to be described into the image description generation system provided in the first aspect of the present invention to obtain a corresponding image text description.
[0038] In a fourth aspect, the present invention provides an electronic device, including: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the method provided in the second aspect or the third aspect of the present invention.
[0039] In a fifth aspect, the present invention further provides a computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method provided in the second aspect or the third aspect of the present invention.
[0040] In a sixth aspect, the invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, they implement the method provided in the second aspect or the third aspect of the present invention.
[0041] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0042] 1. The present invention provides an image description generation system. The visual features of an image are mapped into a space where vision and language are comparable. After obtaining a semantic information sequence, cross-modal semantic calculation of the semantic information sequence and the visual feature sequence is achieved through a Transformer decoder, thereby obtaining the intermediate hidden states of each candidate word. Then, a corresponding directed acyclic graph is constructed. A candidate path with a path length equal to a preset length is selected from the directed acyclic graph, and the candidate path with the highest probability is used as the optimal path, which is directly mapped into an image text description by a linear classifier. The present invention makes full use of the visual information and the semantic information contained in the image, learns the sequential relationship between words by introducing a directed acyclic graph, improves the fluency of the generated description, and the present invention can directly generate an image text description based on the optimal path at one time, has a non-autoregressive decoding property, improves the inference speed, and can generate high-quality image text descriptions at a relatively fast speed.
[0043] 2. Further, considering that the scenes of some images are relatively complex and have multi-scale characteristics, in order to extract as rich image information as possible, the image encoder in the image description generation system provided by the present invention first extracts image features of different scales from image patches, and then fuses the image features of different levels in a top-down and lateral connection manner, thereby obtaining the visual features of the image patches with rich information, which can further improve the quality of the image text description.
[0044] 3. The present invention provides a training method for an image description generation system, with a first training objective of minimizing the sum of the image description losses of all image samples, and training the image description generation system based on the total training objective including the first training objective, so that the generalization of the image description generation system is better, and the quality of the image text description generated by the image description generation system is improved.
[0045] 4. Further, the training method for the image description generation system provided by the present invention constructs a second training objective of minimizing the sum of the global alignment losses of all image samples, and then obtains a total training objective that also includes the second training objective, which can further enhance the semantic information input to the Transformer decoder, so that the semantic information sequence contains more language information at the overall sentence level, helps the decoder map and transform the global semantics into a descriptive sentence, and further improves the quality of the image text description generated by the image description generation system.
[0046] 5. Further, the training method for the image description generation system provided by the present invention, with The third training objective is targeted, and then the total training objective including the third training objective is obtained, which can further optimize the local text information in the semantic information sequence, so that the semantic information sequence contains more word-level language information, helping the decoder to map and transform the local semantics into descriptive vocabulary, and further improving the quality of the image text description generated by the image description generation system. Brief Description of the Drawings
[0047] Figure 1 FIG. 6 is a schematic structural diagram of an image description generation system provided by an embodiment of the present invention. Detailed Embodiments
[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0049] To achieve the above object, in a first aspect, the present invention provides an image description generation system, including:
[0050] An image encoder, configured to divide an image to be described into image blocks, and extract visual features of each image block, so as to obtain a visual feature sequence;
[0051] An embedding mapping layer, configured to map and encode the visual feature sequence into a corresponding semantic information sequence by using a visual encoder of a vision-language model;
[0052] A Transformer decoder, including a plurality of cascaded decoding blocks, configured to use the semantic information sequence as the input of the multi-head attention module in the first decoding block, and use the visual feature sequence as the K value and V value of the cross-attention module in each decoding block to implement cross-modal semantic calculation of the semantic information sequence and the visual feature sequence, so as to obtain intermediate hidden states of each candidate word;
[0053] A path selection module, which is used to calculate a transition probability matrix based on an intermediate hidden vector formed by intermediate hidden states of all candidate words; perform a lower triangular mask on the state transition matrix, and construct a corresponding directed acyclic graph based on the masked state transition matrix; select candidate paths with a path length of a preset length from the directed acyclic graph, and use the candidate path with the highest probability as the optimal path; wherein, the state transition matrix includes: the transition probability between intermediate hidden states of any two candidate words; the nodes in the directed acyclic graph are the intermediate hidden states of candidate words, and the weight of an edge is the transition probability between two nodes on the edge; the probability of a candidate path is the product of the transition probabilities between all adjacent two nodes on the candidate path.
[0054] A linear classifier, which is used to map each intermediate hidden state on the optimal path to a corresponding word to obtain an image text description.
[0055] It should be noted that the above visual language model can be any existing visual language model, such as the clip model, etc., which is not limited here; the visual encoder can adopt resnet, vision-transformer, etc. in the clip model, which is not limited here.
[0056] Preferably, in an alternative implementation, the above Transformer decoder is a bidirectional Transformer decoder; the bidirectional Transformer decoder removes the mask mechanism of self-attention, and can simply and conveniently generate an image text description based on the optimal path at one time, further improving the efficiency of image description generation.
[0057] Considering that the scenes of some images are relatively complex and have multi-scale characteristics, in order to extract as rich image information as possible, preferably, in an alternative implementation, the above image encoder includes:
[0058] A feature extraction module, which extracts image features of N different scales of an image block; sequentially records each image feature in descending order of scale as the first image feature to the Nth image feature; N≥2;
[0059] A feature fusion module, which aligns the first image feature and the second image feature with each other and then adds them to obtain a first fused image feature; aligns the ith fused image feature and the (i + 2)th image feature with each other and then adds them to obtain the (i + 1)th fused image feature; i = 1, 2,..., N - 2; uses the (N - 1)th fused image feature as the visual feature of the image block.
[0060] Wherein, the feature fusion module aligns two image features of different scales in the following way:
[0061] For the large-scale image feature B and the small-scale image feature S in the image features of two different scales respectively, perform a channel dimension transformation on the large-scale image feature B to make it consistent with the channel dimension of the small-scale image feature S; perform upsampling on the small-scale image feature S to make it consistent with the scale of the large-scale image feature B.
[0062] It should be noted that the above feature extraction module can be any existing feature extraction module, such as CNN, Swin Transformer, etc., which is not limited here.
[0063] In an optional implementation manner, the transition probability matrix is:
[0064]
[0065] Among them, H is the intermediate hidden vector composed of the intermediate hidden states of all candidate words; d is the length of the intermediate hidden vector; W Q and W K are both pre-trained learnable parameters, which are obtained by training during the training process of the image caption generation system.
[0066] It should be noted that the linear classifier can be any existing linear classifier, which is not limited here. Preferably, in an optional implementation manner, the above linear classifier includes a cascaded linear layer and a softmax layer;
[0067] After each intermediate hidden state on the optimal path undergoes a linear transformation through the linear layer, it is input into the softmax layer to obtain the probability that each intermediate hidden state on the optimal path is mapped to different words in the preset vocabulary set, and the word corresponding to the maximum probability is used as the word corresponding to this intermediate hidden state.
[0068] To further illustrate the image caption generation system provided by the present invention, a specific embodiment will be described in detail below:
[0069] This embodiment takes a remote sensing image as an example. As Figure 1 shown, the image caption generation system provided by this embodiment includes: an image encoder, an embedding mapping layer, a Transformer decoder, a path selection module, and a linear classifier;
[0070] 1) Image encoder
[0071] The image encoder is used to divide the image to be described into image patches and extract the visual features of each image patch, thereby obtaining a visual feature sequence.
[0072] Considering the multi-scale characteristics of remote sensing images, the above image encoder in this embodiment includes a feature extraction module and a feature fusion module.
[0073] In this embodiment, Swin Transformer is used as the feature extraction module; image features of N (N = 4 in this embodiment) different scales are extracted through the hierarchical structure in Swin Transformer. First, the input remote sensing image is divided into multiple small blocks. These small blocks will be sent to the four-level hierarchical structure of Swin Transformer for processing. In the first-level hierarchical structure, after changing the number of channels through the linear embedding layer, the small blocks are sent to the first encoding block of Swin-Transformer for processing to obtain the first image feature; the first image feature represents the initial high-resolution feature. In the second-level hierarchical structure, the first image feature is downsampled and then input to the second encoding block of Swin-Transformer for processing to obtain the second image feature; the resolution of the second image feature is reduced, but the richness of the features increases. In the third-level hierarchical structure, the second image feature is downsampled and then input to the third encoding block of Swin-Transformer for processing to obtain the third image feature to further enhance the abstraction and semantic information of the features. In the fourth-level hierarchical structure, the third image feature is downsampled and then input to the fourth encoding block of Swin-Transformer for processing to obtain the fourth image feature, and the abstraction and semantic information of the features are further enhanced. Among them, the output of the i-th encoding block of Swin-Transformer can be expressed as:
[0074] X i = SwinTransformerBlock i (upSample(X i-1 ))
[0075] Among them, i = 1, 2, 3, 4; H and W respectively represent the height and width of the input remote sensing image, C represents the number of output channels; SwinTransformerBlock i (·) represents the operation of the i-th encoding block of Swin-Transformer; upSample(·) represents the downsampling operation.
[0076] The image features are sequentially recorded as the first image feature to the N-th image feature in descending order of scale; N ≥ 2;
[0077] The feature fusion module fuses image features at different levels in a top-down and lateral connection manner. Specifically, the first image feature and the second image feature are aligned with each other and then added to obtain the first fused image feature; the i-th fused image feature and the (i + 2)-th image feature are aligned with each other and then added to obtain the (i + 1)-th fused image feature; i = 1, 2, …, N - 2; the (N - 1)-th fused image feature is used as the visual feature of the image patch.
[0078] In this embodiment, the feature fusion module aligns two image features at different scales in the following way:
[0079] For the large-scale image feature B and the small-scale image feature S in the two image features at different scales respectively, the large-scale image feature B is subjected to channel dimension transformation to be consistent with the channel dimension of the small-scale image feature S; the small-scale image feature S is upsampled to be consistent with the scale of the large-scale image feature B.
[0080] In this embodiment, the large-scale image feature B is subjected to channel dimension transformation through 1×1 convolution; the upsampling of the small-scale image feature S is to upsample the small-scale image feature S by a factor of two.
[0081] 2) Embedding mapping layer
[0082] The embedding mapping layer is used to map and encode the visual feature sequence into a visual and language comparable space by using the visual encoder of the vision-language model to obtain the corresponding semantic information sequence. The vision-language model in this embodiment is the clip model.
[0083] 3) Transformer decoder
[0084] The Transformer decoder includes a plurality of cascaded decoding blocks, which are used to take the semantic information sequence as the input of the multi-head attention module in the first decoding block, and take the visual feature sequence as the K value and V value of the cross-attention module in each decoding block to realize the cross-modal semantic calculation of the semantic information sequence and the visual feature sequence, so as to obtain the intermediate hidden states of each candidate word.
[0085] The following operations are performed in the first decoding block of the Transformer decoder: perform multi-head attention calculation on the semantic information sequence, normalize the obtained attention scores to obtain language attention information, use the language attention information as the Q value, and perform cross-attention calculation with the visual feature sequence as the K value and V value to obtain the first-layer cross-modal attention information. Further input the first-layer cross-modal attention information into the feed-forward layer to extract cross-modal deep semantic information. Finally, after normalizing the cross-modal deep semantic information, input it into the first decoding block of the Transformer decoder; the following operations are performed in the second decoding block of the Transformer decoder: perform multi-head attention calculation on the information input by the first decoding block of the Transformer decoder, normalize the obtained attention scores to obtain language attention information, use the language attention information as the Q value, and perform cross-attention calculation with the visual feature sequence as the K value and V value to obtain the second-layer cross-modal attention information. Further input the second-layer cross-modal attention information into the feed-forward layer to extract cross-modal deep semantic information. Finally, after normalizing the cross-modal deep semantic information, input it into the third decoding block of the Transformer decoder; and so on. The output of the last decoding block in the Transformer decoder is the intermediate hidden state of each candidate word.
[0086] 4) Path selection module
[0087] The path selection module is used to calculate the transition probability matrix based on the intermediate hidden vector composed of the intermediate hidden states of all candidate words; perform a lower triangular mask on the state transition matrix, and construct a corresponding directed acyclic graph based on the masked state transition matrix; select candidate paths with a path length of a preset length from the directed acyclic graph, and use the candidate path with the highest probability as the optimal path; where the state transition matrix includes: the transition probability between the intermediate hidden states of any two candidate words; the nodes in the directed acyclic graph are the intermediate hidden states of candidate words, and the edges are the transition probabilities between the two nodes on the edge; the probability of a candidate path is the product of the transition probabilities between all adjacent two nodes on the candidate path.
[0088] Among them, the transition probability matrix is:
[0089]
[0090] Among them, H is the intermediate hidden vector composed of the intermediate hidden states of all candidate words, denoted as H = [h1, h2, …, h L ; d is the length of the intermediate hidden vector; W Q and W K are both pre-trained learnable parameters.
[0091] The element in the \(i\)-th row and \(j\)-th column of the transition probability matrix represents the probability that the intermediate hidden state of the \(i\)-th candidate word transfers to the intermediate hidden state of the \(j\)-th candidate word; \(i = 1, 2, \ldots, L\), \(j = 1, 2, \ldots, L\). Apply a lower triangular mask to the state transition matrix, making all elements below the main diagonal of the state transition matrix equal to 0, thereby constructing a corresponding directed acyclic graph.
[0092] The probability of candidate path A is: M is the preset length; is the probability that the intermediate hidden state of the \(i\)-th candidate word on candidate path A transfers to the intermediate hidden state of the \(i + 1\)-th candidate word.
[0093] 5) Linear classifier
[0094] The linear classifier is used to map each intermediate hidden state on the optimal path to the corresponding word, obtaining the image text description.
[0095] The linear classifier in this embodiment includes a cascaded linear layer and a softmax layer;
[0096] After each intermediate hidden state on the optimal path undergoes a linear transformation through the linear layer, it is input into the softmax layer, obtaining the probability that each intermediate hidden state on the optimal path is mapped to different words in the preset vocabulary set, and taking the word corresponding to the maximum probability as the word corresponding to this intermediate hidden state.
[0097] For the vector formed by each intermediate hidden state on the optimal path After passing through the linear layer and the softmax layer in sequence, a probability matrix is obtained where \(P\in R\) D×M ; where D is the number of words in the preset vocabulary set; represents the probability that the vector formed by the \(i\)-th intermediate hidden state on the optimal path is mapped to the \(j\)-th word in the preset vocabulary set; \(i = 1, 2, \ldots, M\), \(j = 1, 2, \ldots, D\).
[0098] It should be noted that the training method of the above image description generation system can adopt a conventional training method, such as an end-to-end training method. Preferably, on the second aspect, the present invention provides a training method for the above image description generation system, including:
[0099] Input each image sample in the pre-collected training set into the above image description generation system to obtain the corresponding image text description; where the training set includes: image samples and corresponding one or more text description labels;
[0100] For each image sample in the training set, after calculating the difference loss between its image text description and each corresponding text description label respectively and then summing them up, the corresponding image description loss is obtained; a first training objective is constructed with the goal of minimizing the sum of the image description losses of all image samples.
[0101] Based on the total training objective including the first training objective, the image description generation system is trained.
[0102] It should be noted that the difference loss between the image text description and the text description label can be measured by means such as cross-entropy loss, log loss function, similarity, etc., which is not limited here.
[0103] In order to further enhance the semantic information input to the Transformer decoder, so that the semantic information sequence contains more language information at the overall sentence level. Preferably, in an alternative implementation, the above training method further includes:
[0104] For each image sample in the training set, calculate the difference loss between its global semantic feature and the corresponding text embedding feature as the corresponding global alignment loss; construct a second training objective with the goal of minimizing the sum of the global alignment losses of all image samples; wherein, the global semantic feature of the image sample is the feature after averaging the semantic information sequence of the image sample; the semantic information sequence is the sequence output by the embedding mapping layer in the image description generation system when the image sample is input into the image description generation system; the text embedding feature is the average value of the features obtained by passing each text description label of the image sample through the language encoder in the vision-language model.
[0105] The above total training objective further includes: the second training objective; while training the image description generation system based on the total training objective, the parameters in the above language encoder are also adjusted.
[0106] It should be noted that the difference loss between the above global semantic feature and the corresponding text embedding feature can be measured by means such as the L2 loss function, similarity calculation, etc., which is not limited here.
[0107] In order to further optimize the local text information in the semantic information sequence, so that the semantic information sequence contains more language information at the word level. Preferably, in an alternative implementation, the above training method further includes:
[0108] For each image sample in the training set, calculate the corresponding word-level value function; wherein, the word-level value function of the k-th image sample is:
[0109]
[0110] Denote the expectation for all e k,n,j of logf L (e k,n,j , W2); n = 1, 2, …, N k ; j = 1, 2, …, L k,n ; N k is the number of text description labels of the k-th image sample; L k,n is the number of words in the n-th text description label of the k-th image sample; e k,n,j is the j-th word in the n-th text description label of the k-th image sample; f L (e k,n,j , W2) is the feature obtained by e k,n,j after passing through the language encoder in the vision-language model; W2 is the learnable parameter in the language encoder of the vision-language model; Denote the expectation for all s k,j' of log(1 - f L (s k,j' , W2)); j' = 1, 2, …, b k ; b k is the number of semantic information in the semantic information sequence of the k-th image sample; s k,j' is the j'-th semantic information in the semantic information sequence of the k-th image sample; the semantic information sequence is the sequence output by the embedding mapping layer in the image description generation system when the image sample is input into the image description generation system; f L (s k,j' , W2) is the feature obtained by s k,j' after passing through the language encoder in the vision-language model;
[0111] Construct the third training objective with as the target; K is the number of image samples in the training set; W1 is the learnable parameter of the vision encoder in the embedding mapping layer;
[0112] The above total training objective further includes: the third training objective; while training the image description generation system based on the total training objective, the parameters in the above language encoder are also adjusted.
[0113] In summary, the present invention uses a Transformer decoder to achieve non-autoregressive decoding attributes, improving the inference speed; and introduces a directed acyclic graph to obtain the sequential dependence relationship of sentence vocabulary, improving the fluency of the generated description; in addition, by using the global alignment loss and the local adversarial loss to enhance the semantic information of the decoder input, it can reduce the difficulty of the process of mapping and converting image features into descriptive text by the decoder, further improving the accuracy of the generated description.
[0114] The related technical solutions are the same as the image description generation system provided in the first aspect of the present invention, which will not be elaborated here.
[0115] In a third aspect, the present invention provides an image description generation method, including: inputting the image to be described into the image description generation system provided in the first aspect of the present invention to obtain a corresponding image text description.
[0116] The related technical solutions are the same as the image description generation system provided in the first aspect of the present invention, which will not be elaborated here.
[0117] In a fourth aspect, the present invention provides an electronic device, including: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the method provided in the second or third aspect of the present invention.
[0118] The related technical solutions are the same as the image description generation system provided in the first aspect of the present invention, which will not be elaborated here.
[0119] In a fifth aspect, the present invention further provides a computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method provided in the second or third aspect of the present invention.
[0120] The related technical solutions are the same as the image description generation system provided in the first aspect of the present invention, which will not be elaborated here.
[0121] In a sixth aspect, the invention further provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the method provided in the second or third aspect of the present invention.
[0122] The related technical solutions are the same as the image description generation system provided in the first aspect of the present invention, which will not be elaborated here.
[0123] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An image description generation system, characterized in that: include: An image encoder is used to divide the image to be described into image blocks and extract the visual features of each image block to obtain a visual feature sequence; An embedding mapping layer, for using a visual encoder of a visual language model to map and encode the visual feature sequence into a corresponding semantic information sequence; A Transformer decoder, comprising a plurality of cascaded decoding blocks, for using the semantic information sequence as an input of a multi-head attention module in a first decoding block, and using the visual feature sequence as a K value and a V value of a cross-attention module in each decoding block, so as to realize cross-modal semantic calculation of the semantic information sequence and the visual feature sequence, thereby obtaining an intermediate hidden state of each candidate word; A path selection module, used for calculating a transition probability matrix based on an intermediate hidden vector formed by the intermediate hidden states of all candidate words; Performing lower triangular masking on the state transfer matrix, and constructing a corresponding directed acyclic graph based on the masked state transfer matrix; A candidate path with a preset length is selected from the directed acyclic graph, and the candidate path with the largest probability is used as the optimal path; wherein the state transition matrix includes: the transition probability between the intermediate hidden states of any two candidate words; the nodes in the directed acyclic graph are the intermediate hidden states of the candidate words, and the weight of the edge is the transition probability between the two nodes on the edge; the probability of the candidate path is the product of the transition probabilities between all two adjacent nodes on the candidate path; A linear classifier is used to map each intermediate hidden state on the optimal path to a corresponding vocabulary to obtain a text description of the image.
2. The image description generation system according to claim 1, characterized in that: The image encoder comprises: A feature extraction module extracts N image features of different scales of the image block; each image feature is recorded in order from large to small scale as the first image feature to the Nth image feature; N ≥ 2; A feature fusion module aligns the first image feature and the second image feature and adds them to obtain a first fused image feature; aligns the i-th fused image feature and the i+2-th image feature and adds them to obtain an i+1-th fused image feature; i=1,2,…,N-2; and uses the N-1-th fused image feature as a visual feature of the image block; The feature fusion module aligns the image features of two different scales in the following manner: For the large-scale image feature B and the small-scale image feature S of the two image features of different scales, the channel dimension of the large-scale image feature B is transformed to keep it consistent with the channel dimension of the small-scale image feature S; the small-scale image feature S is upsampled to keep it consistent with the scale of the large-scale image feature B.
3. The image description generation system according to claim 1, characterized in that: The transition probability matrix is: Where H is the intermediate hidden vector composed of the intermediate hidden states of all candidate words; d is the length of the intermediate hidden vector; W Q and W K All are pre-trained learnable parameters.
4. The image description generation system according to any one of claims 1 to 3, characterized in that: The linear classifier includes a cascaded linear layer and a softmax layer; Each intermediate hidden state on the optimal path is linearly transformed by the linear layer and then input into the softmax layer, so that each intermediate hidden state on the optimal path is mapped to the probability of different words in the preset vocabulary set, and the word corresponding to the maximum probability is used as the word corresponding to the intermediate hidden state.
5. The training method for an image description generation system according to any one of claims 1 to 4, characterized in that: include: Input each image sample in the pre-collected training set into the image description generation system to obtain the corresponding image text description; wherein the training set includes: image samples and corresponding one or more text description labels; For each image sample in the training set, respectively calculate the difference loss between its image text description and each corresponding text description label, and then sum them up to obtain the corresponding image description loss; construct a first training objective with the goal of minimizing the sum of the image description losses of all image samples; The image description generation system is trained based on an overall training objective including the first training objective.
6. The training method according to claim 5, characterized in that: Also includes: For each image sample in the training set, the difference loss between its global semantic feature and the corresponding text embedding feature is calculated as the corresponding global alignment loss; a second training target with the goal of minimizing the sum of the global alignment losses of all image samples is constructed; wherein the global semantic feature of the image sample is the feature after averaging the semantic information sequence of the image sample; the semantic information sequence is the sequence output by the embedding mapping layer in the image description generation system when the image sample is input into the image description generation system; the text embedding feature is the average value of the features obtained by the language encoder in the visual language model for each text description label of the corresponding image sample; The overall training goal also includes: the second training goal; while training the image description generation system based on the overall training goal, the parameters in the language encoder are also adjusted.
7. The training method according to claim 5 or 6, characterized in that: Also includes: For each image sample in the training set, the corresponding word level value function is calculated; wherein the word level value function of the kth image sample is: For all e k,n,j logf L (e k,n,j ,W2) Find the expectation; n=1,2,…,N k ; j = 1, 2, ..., L k,n ; N k is the number of text description labels of the kth image sample; L k,n is the number of words in the nth text description label of the kth image sample; e k,n,j is the jth word in the nth text description label of the kth image sample; f L (e k,n,j ,W2) is e k,n,j Features obtained through the language encoder in the visual language model; W2 is the learnable parameter in the language encoder in the visual language model; For all s k,j' log(1-f L (s k,j' ,W2)) Find the expectation; j'=1,2,…,b k ; b k is the amount of semantic information in the semantic information sequence of the kth image sample; s k,j' is the j'th semantic information in the semantic information sequence of the kth image sample; the semantic information sequence is the sequence output by the embedding mapping layer in the image description generation system when the image sample is input into the image description generation system; f L (s k,j' ,W2) is s k,j' Features obtained through the language encoder in the visual language model; Build with is the third training target of the target; K is the number of image samples in the training set; W1 is the learnable parameter of the visual encoder in the embedding mapping layer; The overall training goal also includes: the third training goal; while training the image description generation system based on the overall training goal, the parameters in the language encoder are also adjusted.
8. A method for generating an image description, characterized in that: include: The image to be described is input into the image description generation system described in any one of claims 1 to 4 to obtain a corresponding image text description.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor executes the method according to any one of claims 5 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to execute the method according to any one of claims 5 to 8.
Citation Information
Cited By
Training method of image description generation model and image description generation method
CN120612564A
A training method of an image description generation model and an image description generation method
CN120612564B