A training method of an image description generation model and an image description generation method
The multi-stage Transformer image description generation technology (VMFormer) guided by visual cues solves the problem of inconsistency between training and reasoning of the Transformer architecture in image description generation tasks, improves the accuracy and visual relevance of generated descriptions, while maintaining the advantages of parallel computing and enhancing the model's sensitivity to visual content.
Patent Information
- Application Number
- CN202511105742.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-08-08
AI Technical Summary
The existing Transformer architecture suffers from inconsistent inputs during training and inference in image description generation tasks, causing the generated descriptions to gradually deviate from the actual image content, especially in terms of the accuracy of key nouns and relational words. At the same time, traditional methods undermine the model's parallel computing advantages and ignore visual information guidance.
A multi-stage Transformer image description generation technology (VMFormer) guided by visual clues is adopted. Through the visual perception plan sampling module (VASS), the ratio of teacher guidance information and self-learning information is dynamically adjusted during the training process. A multi-level visual perception strategy is used to generate mixed information sequences, and the model parameters are optimized by combining cross entropy loss and self-critical sequence training.
It effectively solves the exposure bias problem, improves the quality and visual relevance of image descriptions, retains the advantages of parallel computing, enhances the model's sensitivity to visual content, and improves the accuracy and stability of generated descriptions.
Smart Images

Figure CN120612564B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing and natural language generation, in particular to a training method of an image description generation model and an image description generation method. BACKGROUND
[0002] With the continuous development of computer vision and natural language processing technology, image description generation task has gradually become an important bridge connecting visual understanding and language generation. The goal of image description generation model is to generate accurate, semantically coherent and visually relevant description sentences according to the input image, covering objects, attributes and their mutual relationships in the image. In recent years, models based on Transformer architecture have made significant progress in image description tasks due to their powerful self-attention mechanism and efficient parallel computing capability, significantly improving the quality and efficiency of description generation.
[0003] However, there is a key problem in the existing Transformer architecture in the image description generation task: the inconsistency of input between the training and inference stages, i.e. the so-called "exposure bias" problem. In the training process, the model usually adopts the "teacher forcing" strategy, using the real labeled sequence as the input of the decoder; while in the inference stage, the model can only rely on its own previous generated prediction results for sequence text generation. This inconsistency between training and inference will cause error accumulation, making the generated description gradually deviate from the actual image content, especially in the accuracy of key nouns (such as objects, attributes) and relationship words (such as verbs).
[0004] In addition, although traditional planned sampling methods can alleviate the exposure bias problem to some extent, these methods were originally designed for RNN (Recurrent Neural Network, RNN: Recurrent Neural Network) / LSTM (Long Short-Term Memory, LSTM: Long Short-Term Memory Network) sequence structure, which dynamically replaces the input label of each time step, which destroys the inherent sequence-level parallel computing advantage of the Transformer model. At the same time, these methods do not distinguish between visual keywords (such as key nouns describing objects, actions and attributes) and non-visual words (such as conjunctions), which may weaken the model's ability to capture core visual content by replacing important visual clues too early. Therefore, how to effectively solve the exposure bias problem while preserving the parallel computing advantage of the Transformer, and improve the model's ability to accurately describe visual content, has become a key problem to be solved in the current image description generation field. SUMMARY
[0005] In view of the above problems, the present invention provides a training method for an image description generation model and an image description generation method, which are used to solve at least one of the existing problems.
[0006] According to a first aspect of the present invention, a method for training an image description generation model is provided, comprising:
[0007] The image description generation model is used to perform the first-stage decoding process on the visual feature samples of the image samples and the label sequence with teacher guidance information to generate an initial prediction sequence with self-learning information;
[0008] Based on the visual clue guidance strategy, the teaching gate in the image description generation model that changes dynamically with training is used to perform multi-level visual perception planning sampling on the labeled sequence and the initial prediction sequence to obtain a mixed information sequence;
[0009] The image description generation model is used to perform a second-stage decoding process on the visual feature samples and the mixed guide sequence generated by the mixed information sequence and the label sequence to obtain an image description generation sample of the image sample;
[0010] The data processing process of the image description generation model is supervised by using a preset loss function, and the parameters of the image description generation model are iteratively optimized based on the loss value to obtain a trained image description generation model.
[0011] According to an embodiment of the present invention, the image description generation model performs a first-stage decoding process on the visual feature samples of the image samples and the label sequence with teacher guidance information to generate an initial prediction sequence with self-learning information, including:
[0012] The pre-trained visual encoder is used to segment the image sample into multiple image patch samples, and combined with the global image representation label of the image sample, the multiple image patch samples are mapped into image patch embedding representation to obtain the visual feature sample of the image sample;
[0013] The first Transformer decoder of the image description generation model is used to perform the first stage decoding processing on the tag sequence and visual feature samples in parallel to generate an initial prediction sequence.
[0014] According to an embodiment of the present invention, the first Transformer decoder using the image description generation model performs a first-stage decoding process on the tag sequence and the visual feature sample in parallel to generate an initial prediction sequence, including:
[0015] The first Transformer decoder is used to process the tag sequence based on the masked multi-head self-attention mechanism and layer normalization to obtain the tag sequence matrix;
[0016] According to the linear projection matrix, the first Transformer decoder is used to process the label sequence matrix and visual feature samples based on the multi-head self-attention mechanism and layer normalization to generate an initial prediction sequence.
[0017] According to an embodiment of the present invention, the above-mentioned visual cue-guided strategy utilizes the teaching gate that changes dynamically with training in the image description generation model to sequentially perform multi-level visual perception planning sampling on the labeled sequence and the initial prediction sequence, and the obtained mixed information sequence includes:
[0018] Using teaching gating, the labeled sequence and the initial prediction sequence are sampled based on sentence-level visual perception planning to obtain an initial mixed information sequence with mixed guidance information. The teaching gating is used to adjust the ratio between teacher-guided information and self-learning information.
[0019] The preset parts of speech in the mixed guidance information are used as visual clues, and the initial mixed information sequence and the tag sequence are sampled based on word-level visual perception plan using teaching gating to obtain the mixed information sequence.
[0020] According to an embodiment of the present invention, the above-mentioned use of teaching gating to perform sentence-level visual perception planning sampling on the labeled sequence and the initial prediction sequence to obtain the initial mixed information sequence with mixed guidance information includes:
[0021] Setting the teaching gate based on the planned sampling starting step number, the interval step number of linear increase of the teaching gate, the preset scaling factor, and the preset maximum value of the teaching gate;
[0022] The first screening ratio determined by the teaching gate is used to screen information between the teacher-guided information and the self-learning information to obtain an initial mixed information sequence with mixed guidance information. The initial mixed information sequence is used as the input sequence for the second-stage decoding processing, wherein the mixed guidance information integrates the teacher-guided information and the self-learning information.
[0023] According to an embodiment of the present invention, the above-mentioned preset part of speech in the mixed guidance information is used as a visual clue, and the initial mixed information sequence and the tag sequence are sampled based on word-level visual perception plan using teaching gating, and the obtained mixed information sequence includes:
[0024] The nouns, verbs and adjectives in the mixed guidance information are used as visual clues to screen the parts of speech of the initial mixed information sequence and the tag sequence, and the second screening ratio determined by the teaching gate is used to perform information screening on the initial mixed information sequence and the tag sequence in parallel to obtain a mixed information sequence.
[0025] According to an embodiment of the present invention, the image description generation model is used to perform a second-stage decoding process on the visual feature sample and the mixed guide sequence generated by the mixed information sequence and the label sequence to obtain an image description generation sample of the image sample, including:
[0026] By combining the visual clue guidance strategy and using the teaching gating of the image description generation model to mix the mixed information sequence and the label sequence, a hybrid guidance sequence is generated;
[0027] The second Transformer decoder of the image description generation model is used to perform a second-stage decoding process on the mixed guide sequence and visual feature samples to obtain an image description generation sample of the image sample, wherein the second Transformer decoder shares the parameters of the first Transformer decoder of the image description generation model.
[0028] According to an embodiment of the present invention, the second Transformer decoder using the image description generation model performs a second-stage decoding process on the mixed guide sequence and the visual feature sample to obtain an image description generation sample of the image sample, including:
[0029] The second Transformer decoder is used to process the mixed guidance sequence based on the masked multi-head self-attention mechanism and layer normalization to obtain the mixed guidance sequence matrix;
[0030] The second Transformer decoder is used to process the mixed guidance sequence matrix and visual feature samples based on the multi-head self-attention mechanism and layer normalization to obtain the initial image description generation sample;
[0031] The second Transformer decoder is used to activate the initial image description generation sample to obtain an image description generation sample of the image sample.
[0032] According to an embodiment of the present invention, the data processing process of the image description generation model is supervised by using a preset loss function, and the parameters of the image description generation model are iteratively optimized based on the loss value to obtain a trained image description generation model, including:
[0033] The cross-entropy loss function is used to supervise the difference between the image description samples output by the image description generation model and the teacher's guidance information, and the parameters of the image description generation model are iteratively optimized by minimizing the cross-entropy loss value;
[0034] The self-critical sequence training loss function is used to supervise the semantic similarity between the image description generated by the image description generation model and the teacher guide information, and the parameter of the image description generation model is iteratively updated through the self-critical sequence training loss value, so that the trained image description generation model is obtained.
[0035] According to a second aspect of the present application, an image description generation method is provided, comprising:
[0036] The pre-trained visual encoder is used for feature extraction of the target image to obtain the visual features of the target image, and the trained image description generation model is used for multi-stage decoding processing of the visual features to generate the image description text of the target image, wherein the trained image description generation model is trained according to the training method of the image description generation model.
[0037] The third aspect of the present disclosure provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.
[0038] The fourth aspect of the present disclosure also provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the method.
[0039] The training method of the image description generation model provided by the present application can effectively solve the technical problems such as exposure bias existing in the existing image description generation method, significantly improve the quality and visual relevance of the image description, retain the parallel computing advantage of the image description generation model, and enhance the sensitivity of the trained image description generation model to visual content, thereby improving the performance of the trained image description generation model and having wide scene practicability and high application value. BRIEF DESCRIPTION OF DRAWINGS
[0040] The above and other objects, features and advantages of the present application will become more apparent from the following description of embodiments of the present application taken with reference to the accompanying drawings, in which:
[0041] Figure 1 is a flowchart of the training method of the image description generation model according to an embodiment of the present application;
[0042] Figure 2 is a framework diagram of visual perception planning sampling according to an embodiment of the present application;
[0043] Figure 3is a framework schematic diagram of a multi-stage Transformer image description generation method based on visual clue guidance according to an embodiment of the present application;
[0044] Figure 4 is a block diagram of an electronic device suitable for implementing a training method and an image description generation method of an image description generation model according to an embodiment of the present application. DETAILED DESCRIPTION
[0045] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It should be understood, however, that the description is merely exemplary and is not intended to limit the scope of the present application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it would be apparent to those skilled in the art that the embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have been omitted or simply referenced in order not to obscure the concept of the present application.
[0046] The terms used herein are merely used to describe specific embodiments and are not intended to limit the present application. The terms "include", "comprise" and the like used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0047] All terms used herein, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of the specification, and should not be interpreted in an idealized or overly formal manner.
[0048] In the case of using expressions similar to "at least one of A, B, and C, etc.", in general, it should be interpreted to include at least one of the items, but not limited to the items (e.g., "a system having at least one of A, B, and C" should include a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.).
[0049] In view of the exposure bias problem existing in the existing Transformer architecture in the image description generation task, and the disadvantages of traditional methods in solving the problem, i.e., destroying the parallel computing advantage and ignoring the visual information guidance, the present application proposes a multi-stage Transformer image description generation technology (VMFormer) based on visual clue guidance. Through the technology, the image description generation model is trained, and the trained image description generation model is used to generate the description text of the target image, thereby improving the accuracy and visual relevance of the description of the target image, while retaining the parallel computing advantage of the image description generation model.
[0050] Figure 1 4 is a flowchart of a method for training an image description generation model according to an embodiment of the present invention.
[0051] like Figure 1 As shown, the training method of the image description generation model includes operations S110 to S140.
[0052] In operation S110, a first-stage decoding process is performed on a visual feature sample of an image sample and a tag sequence with teacher guidance information using an image description generation model to generate an initial prediction sequence with self-learning information.
[0053] The above-mentioned image description generation model includes a first Transformer decoder, a second Transformer decoder, a teaching gate, etc.
[0054] The first Transformer decoder in the image description generation model is used for the first stage of decoding operations; among them, visual feature samples can be extracted from image samples using the Vision Transformer (ViT) visual encoder in the pre-trained CLIP (Contrastive Language-Image Pre-training) model.
[0055] The teacher guidance information is the label information corresponding to the image sample. The first stage of decoding processing is performed on the above visual feature samples and label sequences to generate an initial prediction sequence with self-learning information.
[0056] In operation S120 , based on the visual cue guidance strategy, a multi-level visual perception plan sampling is performed on the labeled sequence and the initial prediction sequence using the teaching gate that changes dynamically with training in the image description generation model to obtain a mixed information sequence.
[0057] The teaching gate in the image description generation model can change with the training steps. The role of the teaching gate is to dynamically control the ratio of teacher guidance information and self-learning information.
[0058] In operation S130, the image description generation model is used to perform a second-stage decoding process on the visual feature sample and the mixed guide sequence generated by the mixed information sequence and the tag sequence to obtain an image description generation sample of the image sample.
[0059] In the training method of the model provided in the application, two-stage decoding processing is required, wherein each stage corresponds to a Transformer decoder: the first-stage decoding processing is completed by a first Transformer decoder; and the second-stage decoding processing is completed by a second Transformer decoder.
[0060] In operation S140, the data processing process of the image description generation model is supervised by using a preset loss function, and the parameters of the image description generation model are iteratively optimized based on the loss value to obtain a trained image description generation model.
[0061] The training method of the image description generation model provided in the application can effectively solve the technical problems such as exposure bias existing in the existing image description generation method, significantly improve the quality and visual relevance of the image description, retain the parallel computing advantage of the image description generation model, and enhance the sensitivity of the trained image description generation model to visual content, thereby improving the performance of the trained image description generation model and having wide scene practicability and high application value.
[0062] The training process of the above image description generation model provided in the application will be further described in detail through specific embodiments.
[0063] The training method of the image description generation model provided by the application first uses the Vision Transformer (ViT) in the pre-trained CLIP model as a visual encoder to divide an input image into multiple image blocks and map the image blocks to image block embeddings to extract image features and obtain a final visual feature representation; then, in a first Transformer decoding stage, the image features and a labeled token sequence (teacher guidance information) are input into a standard first Transformer decoder to generate an initial prediction sequence (self-learning information) in a parallel manner; secondly, the proportion of the teacher guidance information and the self-learning information is dynamically controlled by a visual perception sampling (VASS) module, a visual cue guided strategy is used, and a mixed information sequence is adaptively generated according to a training step, the visual-related keywords are preferentially retained in an early training stage, and the non-visual connecting words in the real labels are gradually replaced with the non-visual connecting words predicted by the model in a later training stage; then, the mixed information sequence is input again into a second Transformer decoder sharing parameters with the first stage (that is, the second Transformer decoder shares the parameters of the first Transformer decoder), and a final output sequence is generated; finally, a cross-entropy loss and a CIDEr (Consensus-based Image Description Evaluation) score optimization loss are used as the training target, the model is trained by minimizing the cross-entropy loss, and the non-differentiable CIDEr index is optimized by using a self-critical sequence training (SCST: Self-Critical Sequence Training) method to improve the quality and visual descriptiveness of the generated description.
[0064] In the inference stage (that is, the application stage of the image description generation model), the trained image description generation model is used to encode and decode an input image to generate an accurate and visually descriptive image description.
[0065] According to the embodiment of the application, the first-stage decoding processing of the visual feature sample of the image sample and the token sequence with teacher guidance information by the image description generation model to generate an initial prediction sequence with self-learning information includes: using a pre-trained visual encoder to divide the image sample into multiple image block samples, and combining the global image representation token of the image sample, the multiple image block samples are mapped to image block embedding representations to obtain the visual feature sample of the image sample; the first-stage decoding processing of the token sequence and the visual feature sample is performed by the first Transformer decoder of the image description generation model in a parallel manner to generate the initial prediction sequence.
[0066] According to an embodiment of the present invention, the first Transformer decoder of the image description generation model performs a first-stage decoding process on the tag sequence and the visual feature samples in parallel to generate an initial prediction sequence, including: using the first Transformer decoder to process the tag sequence based on a masked multi-head self-attention mechanism and layer normalization to obtain a tag sequence matrix; according to the linear projection matrix, using the first Transformer decoder to process the tag sequence matrix and the visual feature samples based on a multi-head self-attention mechanism and layer normalization to generate an initial prediction sequence.
[0067] The above embodiments involve the extraction of visual feature samples from image samples and the first stage decoding of the image description generation model (ie, the self-learning information generation process).
[0068] In the extraction stage of visual feature samples: the Vision Transformer (ViT) in the pre-trained CLIP model is used as a visual encoder to divide the input image sample into multiple image block samples, and these image block samples are mapped to image block embedding samples to obtain the initial visual feature sample representation. Specifically, given an input image sample , ViT CLIP-based image encoder divides the input image samples into image block samples, and these image block samples Mapping to image patch embedding , as shown in formula (1):
[0069] (1),
[0070] in, represents the hidden dimension of CLIP, Represents the real number space. ViT has L layers, and the input of the lth layer is Enter formula (2) as shown:
[0071] (2),
[0072] in, is the global image representation tag ([CLS] tag embedding). After passing through all ViT layers, the final visual feature sample representation is obtained , the visual feature sample representation will serve as the visual condition in the subsequent decoding stage. The embedding vector of each image block sample contains the visual information of the image block sample, providing a basis for subsequent feature fusion and description generation.
[0073] The first stage decoding process of the image description generation model: the extracted visual feature samples and the labeled tag sequence (teacher guidance information) are input into the standard Transformer decoder to generate the initial prediction sequence (self-learning information) in parallel. Specifically, given an image feature sample and labeled sequences (teacher-guided information) , the present invention uses a standard Transformer decoder to generate the initial prediction sequence (self-learning information) The purpose of this stage is to allow the model to initially learn the generation of descriptions under the guidance of the teacher, providing a reference for the subsequent generation of mixed information. This process can be expressed by formulas (3) and (4):
[0074] (3),
[0075] (4),
[0076] in, represents a multi-head self-attention mechanism with mask, represents the multi-head attention mechanism, used to integrate visual features, is the linear projection matrix, Layer normalization operation, Represents the decoder of an image description generation model.
[0077] According to an embodiment of the present invention, the above-mentioned visual clue guidance strategy uses the teaching gate that dynamically changes with training in the image description generation model to perform multi-level visual perception plan sampling on the labeled sequence and the initial prediction sequence in sequence to obtain a mixed information sequence, including: using the teaching gate to perform sentence-level visual perception plan sampling on the labeled sequence and the initial prediction sequence to obtain an initial mixed information sequence with mixed guidance information, wherein the teaching gate is used to adjust the ratio between the teacher's guidance information and the self-learning information; using the preset part of speech in the mixed guidance information as a visual clue, and using the teaching gate to perform word-level visual perception plan sampling on the initial mixed information sequence and the labeled sequence to obtain a mixed information sequence.
[0078] The following is a specific implementation method and combined with the attached Figure 2 The visual perception plan sampling process provided by the present invention is further described in detail.
[0079] Figure 2 4 is a framework diagram of visual perception plan sampling according to an embodiment of the present invention.
[0080] like Figure 2 As shown in Figure 3, the visual perception plan adopts two parts, including a dynamic hybrid information generation mechanism and a visual clue guidance strategy, where BOS represents the beginning of a sentence.
[0081] In the dynamic mixed information generation mechanism: dynamically control the ratio of teacher-guided information and self-learning information, use visual clues to guide the strategy, and adaptively generate mixed information sequences according to the training steps. Specifically, the VASS module uses a teaching gate that changes with the training rounds. To dynamically adjust the ratio of teacher-guided information and self-learning information. Teaching gate control The definition of is shown in formula (5):
[0082] (5),
[0083] in, represents the maximum value of the teaching gate, Indicates the starting number of planned sampling steps, represents the number of steps of linearly increasing teaching gates, and is a scaling factor, Indicates the training round.
[0084] Different from the traditional token-level planned sampling, this paper adopts a sentence-level sampling strategy to maintain the parallel computing advantage of the Transformer architecture. Specifically, for each sentence in the training batch, according to the probability Select teacher guidance information or self-learning information as the input of the second stage decoding. This process can be expressed by formula (6):
[0085] (6),
[0086] in, Indicates the number of The mixed information of the sentence, represents the self-learning information in the mixed information, Represents the teacher guidance information in the mixed information. The final generated mixed information sequence It will be input into the Transformer decoder of the second stage, effectively simulating the autoregressive generation process of the inference stage, thereby promoting consistency between the training and inference stages.
[0087] According to an embodiment of the present invention, the above-mentioned use of teaching gating to perform sentence-level visual perception plan sampling on the labeled sequence and the initial prediction sequence to obtain an initial mixed information sequence with mixed guidance information includes: setting the teaching gate based on the starting number of planned sampling steps, the number of interval steps of linear increase of the teaching gate, the preset scaling factor and the maximum value of the preset teaching gate; using the first screening ratio determined by the teaching gate to perform information screening between the teacher guidance information and the self-learning information to obtain the initial mixed information sequence with mixed guidance information, and using the initial mixed information sequence as the input sequence for the second stage decoding processing, wherein the mixed guidance information integrates the teacher guidance information and the self-learning information.
[0088] According to an embodiment of the present invention, the above-mentioned method uses the preset parts of speech in the mixed guidance information as visual clues, and uses the teaching gating to perform word-level visual perception plan sampling on the initial mixed information sequence and the tag sequence to obtain the mixed information sequence, including: using the nouns, verbs and adjectives in the mixed guidance information as visual clues to perform part-of-speech screening on the initial mixed information sequence and the tag sequence, and using the second screening ratio determined by the teaching gating to perform information screening on the initial mixed information sequence and the tag sequence in parallel to obtain the mixed information sequence.
[0089] In the visual clue guidance strategy stage: In the early stages of training, in order to enhance the model's ability to capture core visual content, the present invention prioritizes retaining visual basic keywords (such as "dog", "running", and "red") from real annotations. These keywords typically include nouns (NN), verbs (VB or VBZ), and adjectives (JJ), which are key elements in describing the core visual content in the image. In the later stages of training, when the model has sufficient ability to self-learn visual information, the present invention gradually replaces non-visual connectives (such as conjunctions and prepositions) to guide the model to master the combination logic of language details. This strategy not only improves the model's sensitivity to visual content, but also enhances the accuracy and stability of the generated description. Formally, this strategy can be expressed by formula (7):
[0090] (7),
[0091] in, Indicates the first Mixed guidance information of words, represents the self-learning information in the mixed guidance information, Indicates the teacher guidance information in the mixed guidance information. Indicates using the NLTK toolkit to perform part-of-speech tagging on the real tags. Indicates a noun, Indicates a verb, Represents attribute words. Specifically, this invention uses words labeled as nouns (NN), verbs (VB), and adjectives (JJ) as visual cues, while treating other part-of-speech tags (such as conjunctions) as non-visual. This allows the model to better focus on and learn the core visual content of the image during training, while gradually mastering the generation logic of language structure.
[0092] According to an embodiment of the present invention, the above-mentioned use of the image description generation model to perform a second-stage decoding process on the visual feature sample and the mixed guide sequence generated by the mixed information sequence and the label sequence to obtain the image description generation sample of the image sample includes: by combining the visual clue guidance strategy, using the teaching gating of the image description generation model to mix the information of the mixed information sequence and the label sequence to generate a mixed guide sequence; using the second Transformer decoder of the image description generation model to perform a second-stage decoding process on the mixed guide sequence and the visual feature sample to obtain the image description generation sample of the image sample, wherein the second Transformer decoder shares the parameters of the first Transformer decoder of the image description generation model.
[0093] According to an embodiment of the present invention, the second Transformer decoder of the image description generation model performs a second-stage decoding process on the mixed guide sequence and the visual feature sample to obtain the image description generation sample of the image sample, including: using the second Transformer decoder to process the mixed guide sequence based on the masked multi-head self-attention mechanism and layer normalization processing to obtain a mixed guide sequence matrix; using the second Transformer decoder to process the mixed guide sequence matrix and the visual feature sample based on the multi-head self-attention mechanism and layer normalization processing to obtain the initial image description generation sample; using the second Transformer decoder to activate the initial image description generation sample to obtain the image description generation sample of the image sample.
[0094] The above embodiment relates to the second-stage decoding process of the image description generation model (i.e., the visual perception mixed information generation process). The above process is further described in detail below through specific embodiments.
[0095] The mixed information sequence is input again into the Transformer decoder that shares parameters with the first stage (i.e., the second Transformer decoder shares the parameters of the first Transformer decoder) to generate the final output sequence. Specifically, the VASS module is gated by teaching Dynamically control ground truth annotation and model predictions The mixed ratio is combined with the visual clue guidance strategy to generate a mixed guidance sequence This process can be formalized using formula (8):
[0096] (8),
[0097] in, represents the dynamic sampling probability, gradually increasing the proportion of model-generated tags, while It is a set of visual vocabulary indicators used to distinguish visually salient marks and linguistic connectives.
[0098] Subsequently, the hybrid boot sequence It is input into the Transformer decoder that shares parameters with the first stage (that is, the second Transformer decoder shares the parameters of the first Transformer decoder) to generate the final output sequence (That is, the model's image description generates samples.) The specific process is shown in formulas (9) to (11):
[0099] (9),
[0100] (10),
[0101] (11),
[0102] in, Representation layer normalization operation, represents a multi-head self-attention mechanism with mask, represents the multi-head attention mechanism, used to integrate visual features, is a linear projection matrix that maps the decoder’s hidden states to the log-odds of the vocabulary, represents the intermediate output of the model, Represents the improved decoder in the image description generation model, which can output more accurate decoding information. The final output Used to calculate the loss of description generation and update model parameters.
[0103] This stage simulates the human cognitive process of describing an image by fusing teacher-guided information with self-learning information, improving the accuracy and visual relevance of generated descriptions. In the early stages of training, visually salient landmarks are prioritized to guide the model in capturing the core semantics of the image. Subsequently, non-visually salient landmarks are gradually replaced with model predictions, prompting the model to learn fluent and coherent language structures. This visually perceptual hybrid information generation strategy not only enhances the model's sensitivity to visual content but also improves the quality and stability of generated descriptions.
[0104] According to an embodiment of the present invention, the above-mentioned method of using a preset loss function to supervise the data processing process of the image description generation model, and iteratively optimizing the parameters of the image description generation model based on the loss value to obtain a trained image description generation model includes: using the cross-entropy loss function to supervise the difference between the image description generation samples output by the image description generation model and the teacher guidance information, and iteratively optimizing the parameters of the image description generation model by minimizing the cross-entropy loss value; using the self-critical sequence training loss function to supervise the semantic similarity between the image description generation samples output by the image description generation model and the teacher guidance information, and iteratively updating the parameters of the image description generation model by the self-critical sequence training loss value to obtain a trained image description generation model.
[0105] During the training process of the image description generation model, the training objective optimization mainly adopts the cross entropy loss function training model and the self-critical sequence training architecture.
[0106] The model is trained using cross entropy loss to minimize the difference between the description generated by the model and the true label, ensuring that the model learns to generate accurate descriptions. Specifically, given the target’s true labeled description sequence and the VMFormer model parameters of the present invention , the present invention minimizes the cross entropy loss , as shown in formula (12):
[0107] (12),
[0108] in, Indicates the length of the description sequence, Indicates that before the given Under the condition of true labels (i.e. description sequence, the same below), the model generates By minimizing the cross-entropy loss, the model can learn to generate descriptions that are as close as possible to the true annotations.
[0109] The self-critical sequence training (SCST) method is used to optimize the non-differentiable CIDEr metric to improve the quality and visual descriptiveness of the generated descriptions. The CIDEr metric can measure the semantic similarity between the generated description and the reference description, thereby improving the accuracy and relevance of the description. Specifically, the optimized objective function The definition is as shown in formula (13):
[0110] (13),
[0111] in, is the reward value of the CIDEr scoring function, is a description sequence sampled from the probability distribution of the model. The gradient of this loss function can be approximated as follows, as shown in formula (14):
[0112] (14),
[0113] in, is a description sequence sampled from the probability distribution of the model, is the baseline description sequence obtained by greedy decoding of the model, is the CIDEr score of the baseline description sequence. By optimizing the CIDEr metric, the model is able to generate descriptions that are more semantically similar to the reference description, thereby improving the quality and visual relevance of the description.
[0114] According to a second aspect of the present invention, an image description generation method is provided, comprising: using a pre-trained visual encoder to extract features of a target image to obtain visual features of the target image, and using a trained image description generation model to perform multi-stage decoding processing on the visual features to generate image description text of the target image, wherein the trained image description generation model is trained according to the above-mentioned image description generation model training method.
[0115] Figure 3 3 is a schematic diagram of a framework of a multi-stage Transformer image description generation method based on visual clue guidance according to an embodiment of the present invention.
[0116] like Figure 3 The image description generation method provided by the present invention mainly includes the training and inference processes of an image description generation model. The training process of the image description generation model mainly includes visual feature extraction, multi-stage Transformer decoding operations, sentence-level and word-level visual perception planning sampling operations guided by visual reality, etc. For example, during the model training process, a photo of an elephant and a woman is input, and "A woman smiles at a painted and dressed elephant" is also input as a text embedding. In this process, a feedforward neural network (FFN) is used, residual connections and layer normalization (Add&Norm) are used, softmax represents the activation function, and a masked multi-head self-attention mechanism (MaskMHA) and a multi-head attention mechanism (MHA) are used.
[0117] During actual training, the present invention uses ViT-L / 14 from CLIP as the image encoder to extract visual features, obtaining a 768-dimensional global feature and 1024-dimensional local patch features with 8×8 spatial resolution for effective visual representation. In the VMFormer encoder and decoder, the hidden layer dimension is 2048, the number of attention heads h is 8, and the number of layers N in the VMFormer encoder and decoder is set to 3. Furthermore, for visual perception plan sampling, the present invention sets the initial probability g of the teaching gate to 0 and increases it by 0.05 every five training epochs; for MSCOCO evaluation, the maximum probability g_max is set to 0.5. The entire image captioning architecture is primarily implemented using the PyTorch deep learning framework on four NVIDIA 4090 GPUs. During training, the present invention trains VMFormer using the cross-entropy loss for 15 epochs, using the Adam optimizer with β1 and β2 taking the default values of 0.9 and 0.999, respectively. The initial learning rate is set to 5e-4 and decays by 0.8 every three epochs. The number of warm-up steps is set to 20,000, and the mini-batch size is 40. When training with the self-critical training strategy, the present invention performs 15 more cycles to directly optimize the CIDEr-D score, with the initial learning rate set to 1e-5, and finally selects the model with the best performance on the validation set for the final test.
[0118] During the inference phase, the trained model is used to encode and decode the input image, generating accurate and visually descriptive image descriptions. This approach effectively addresses the issue of exposure bias while maintaining the parallel computing advantages of the Transformer, significantly improving the accuracy and visual relevance of generated descriptions.
[0119] In summary, this paper effectively solves the exposure bias problem existing in existing methods through a multi-stage decoding mechanism guided by visual cues, significantly improves the quality and visual relevance of image description, while retaining the parallel computing advantages of the model, and has important theoretical and application value.
[0120] Compared with the prior art, the present invention has the following beneficial effects:
[0121] (1) Effectively alleviate the problem of exposure bias: The VMFormer framework proposed in this paper adopts a sentence-level sampling strategy through the Visual Perception Planned Sampling (VASS) module to dynamically combine model predictions with true labels, effectively bridging the gap between training and inference, significantly reducing error accumulation, and thus generating more accurate image descriptions.
[0122] (2) Preserving the parallel computing advantage of the Transformer: Through a two-stage decoding mechanism, this paper maintains the parallel computing advantage of the Transformer architecture while guiding the strategy through visual cues to ensure that the model can accurately capture the core visual content. This design not only improves the training efficiency of the model, but also enhances the stability of the generated descriptions.
[0123] (3) Improving the accuracy and visual relevance of generated descriptions: The VASS module prioritizes visual keywords (such as nouns, verbs, and adjectives), strengthening the model's ability to capture core visual content. In the later stages of training, the non-visual connectives in the real labels are gradually replaced with non-visual connectives predicted by the model, guiding the model to grasp the combinatorial logic of language details, significantly improving the accuracy and visual relevance of generated descriptions.
[0124] (4) Enhancing the model’s sensitivity to visual content: This paper uses a visual cue guidance strategy to enable the model to better focus on and learn the core visual content in the image during training, thereby improving the model’s sensitivity to visual content and further enhancing the quality and stability of the generated description.
[0125] (5) Wide Applicability and Performance Improvement: The VASS module can be used as a plug-and-play component, seamlessly adapting to the Transformer architecture and pre-trained large language models, significantly improving the performance of existing image description methods. Experimental results on multiple benchmark datasets show that VMFormer outperforms existing methods in all indicators, verifying the effectiveness and superiority of the present invention.
[0126] Figure 4 4 is a block diagram of an electronic device suitable for implementing a training method for an image description generation model and an image description generation method according to an embodiment of the present invention.
[0127] like Figure 4 As shown, an electronic device 400 according to an embodiment of the present invention includes a processor 401, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 402 or a program loaded from a storage unit 408 into a random access memory (RAM) 403. Processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. Processor 401 may also include onboard memory for caching purposes. Processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0128] In the RAM 403, various programs and data required for the operation of the electronic device 400 are stored. The processor 401, the ROM 402, and the RAM 403 are connected to each other via the bus 404. The processor 401 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 402 and / or the RAM 403. It should be noted that the programs can also be stored in one or more memories other than the ROM 402 and the RAM 403. The processor 401 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.
[0129] According to the embodiments of the present application, the electronic device 400 can further include an input / output (I / O) interface 405, which is also connected to the bus 404. The electronic device 400 can further include one or more of the following components connected to the input / output (I / O) interface 405: an input part 406 including a keyboard, a mouse, etc.; an output part 407 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 408 including a hard disk, etc.; and a communication part 409 including a network interface card such as a LAN card, a modem, etc. The communication part 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output (I / O) interface 405 as necessary. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 410 as necessary, so that a computer program read out therefrom is installed in the storage part 408 as necessary.
[0130] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0131] According to embodiments of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, such as, for example, without limitation, a portable computer diskette, a hard disk, random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to embodiments of the present application, the computer readable storage medium can include the ROM 402 and / or the RAM 403 described above, and / or one or more other memories not expressly described above.
[0132] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functional processes, and operational processes, according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0133] Those skilled in the art will understand that features recited in various embodiments of the present application can be combined and / or integrated in various ways, even if such combinations or integrations are not expressly noted in the present application. In particular, features recited in various embodiments of the present application can be combined and / or integrated in ways that are not expressly noted in the present application, without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.
[0134] The embodiments of the present application described above are merely intended to illustrate the present application. These embodiments are merely for illustrative purposes, and are not intended to limit the scope of the present application. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Various alternatives and modifications can be made to the embodiments of the present application by those skilled in the art without departing from the scope of the present application, and such alternatives and modifications are intended to fall within the scope of the present application.
Claims
1. A training method for an image description generation model, characterized in that: The method comprises: The image description generation model is used to perform the first-stage decoding process on the visual feature samples of the image samples and the label sequence with teacher guidance information to generate an initial prediction sequence with self-learning information; Based on a visual cue guidance strategy, a teaching gate that changes dynamically with training in the image description generation model is used to perform multi-level visual perception planning sampling on the label sequence and the initial prediction sequence to obtain a mixed information sequence; performing a second-stage decoding process on the visual feature sample and the mixed guide sequence generated by the mixed information sequence and the tag sequence using the image description generation model to obtain an image description generation sample of the image sample; The data processing process of the image description generation model is supervised by using a preset loss function, and the parameters of the image description generation model are iteratively optimized based on the loss value to obtain a trained image description generation model.
2. The method according to claim 1, characterized in that The image description generation model is used to perform the first stage of decoding processing on the visual feature samples of the image samples and the label sequence with teacher guidance information to generate the initial prediction sequence with self-learning information, including: Using a pre-trained visual encoder to segment the image sample into multiple image block samples, and combining the global image representation labels of the image samples, mapping the multiple image block samples into image block embedding representations to obtain visual feature samples of the image samples; The first Transformer decoder of the image description generation model is used to perform a first-stage decoding process on the tag sequence and the visual feature sample in parallel to generate the initial prediction sequence.
3. The method according to claim 2, characterized in that Performing a first-stage decoding process on the tag sequence and the visual feature sample in parallel using a first Transformer decoder of the image description generation model to generate the initial prediction sequence includes: Using the first Transformer decoder to process the tag sequence based on a masked multi-head self-attention mechanism and layer normalization, to obtain a tag sequence matrix; According to the linear projection matrix, the first Transformer decoder is used to perform multi-head self-attention mechanism-based processing and layer normalization processing on the label sequence matrix and the visual feature samples to generate the initial prediction sequence.
4. The method according to claim 1, wherein Based on the visual clue guidance strategy, the multi-level visual perception plan sampling is performed on the label sequence and the initial prediction sequence in sequence using the teaching gate that changes dynamically with training in the image description generation model, and the mixed information sequence obtained includes: Using the teaching gating to perform sentence-level visual perception planning sampling on the labeled sequence and the initial prediction sequence to obtain an initial mixed information sequence with mixed guidance information, wherein the teaching gating is used to adjust the ratio between the teacher guidance information and the self-learning information; The preset parts of speech in the mixed guidance information are used as visual clues, and the teaching gate is used to perform word-level visual perception planning sampling on the initial mixed information sequence and the tag sequence to obtain the mixed information sequence.
5. The method according to claim 4, characterized in that The teaching gate is used to perform sentence-level visual perception planning sampling on the label sequence and the initial prediction sequence to obtain an initial mixed information sequence with mixed guidance information, including: Setting the teaching gate based on the planned sampling starting step number, the interval step number of linear increase of the teaching gate, the preset scaling factor and the preset maximum value of the teaching gate; The first screening ratio determined by the teaching gate is used to perform information screening between the teacher guidance information and the self-learning information to obtain the initial mixed information sequence with mixed guidance information, and the initial mixed information sequence is used as the input sequence of the second stage decoding processing, wherein the mixed guidance information integrates the teacher guidance information and the self-learning information.
6. The method according to claim 4, characterized in that The preset part of speech in the mixed guidance information is used as a visual clue, and the teaching gate is used to perform word-level visual perception planning sampling on the initial mixed information sequence and the tag sequence, so as to obtain the mixed information sequence including: The nouns, verbs and adjectives in the mixed guidance information are used as visual clues to perform part-of-speech screening on the initial mixed information sequence and the labeled sequence, and the initial mixed information sequence and the labeled sequence are screened in parallel using a second screening ratio determined by the teaching gate to obtain the mixed information sequence.
7. The method according to claim 1, characterized in that The image description generation model is used to perform a second-stage decoding process on the visual feature sample and the mixed guide sequence generated by the mixed information sequence and the tag sequence to obtain an image description generation sample of the image sample, including: By combining the visual clue guidance strategy and utilizing the teaching gating of the image description generation model, the mixed information sequence and the label sequence are mixed to generate the mixed guidance sequence; The mixed guide sequence and the visual feature sample are subjected to a second-stage decoding process by using a second Transformer decoder of the image description generation model to obtain an image description generation sample of the image sample, wherein the second Transformer decoder shares parameters of the first Transformer decoder of the image description generation model.
8. The method according to claim 7, characterized in that Performing a second-stage decoding process on the mixed guide sequence and the visual feature sample using a second Transformer decoder of the image description generation model to obtain an image description generation sample of the image sample includes: Using the second Transformer decoder to process the hybrid guide sequence based on a masked multi-head self-attention mechanism and layer normalization, to obtain a hybrid guide sequence matrix; Using the second Transformer decoder, the mixed guide sequence matrix and the visual feature sample are processed based on a multi-head self-attention mechanism and layer normalized to obtain an initial image description generation sample; The second Transformer decoder is used to perform activation processing on the initial image description generation sample to obtain an image description generation sample of the image sample.
9. The method according to claim 1, characterized in that The data processing process of the image description generation model is supervised by using a preset loss function, and the parameters of the image description generation model are iteratively optimized based on the loss value to obtain a trained image description generation model, including: Using a cross-entropy loss function to supervise the difference between the image description generation samples output by the image description generation model and the teacher guidance information, and iteratively optimizing the parameters of the image description generation model by minimizing the cross-entropy loss value; The semantic similarity between the image description generation samples output by the image description generation model and the teacher guidance information is supervised by using the self-critical sequence training loss function, and the parameters of the image description generation model are iteratively updated through the self-critical sequence training loss value to obtain the trained image description generation model.
10. A method for generating an image description, characterized in that: The method comprises: A pre-trained visual encoder is used to extract features of a target image to obtain visual features of the target image, and a trained image description generation model is used to perform multi-stage decoding processing on the visual features to generate image description text of the target image, wherein the trained image description generation model is trained according to the training method according to any one of claims 1 to 9.
11. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Two-stage image description generation method, system and device and storage medium
CN117576534A
Image description generation system, training method, generation method and electronic equipment
CN120219769A