Zero sample image text conversion method and device
By conducting entity characterization training and decoder model training in zero-sample image text conversion, combining the graphic-text alignment characterization model and the comparison learning loss function, the problem of low accuracy in generated text description in the prior art is solved, and a more accurate image text conversion effect is achieved.
Patent Information
- Application Number
- CN202510093791.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing zero-sample image text conversion method results in a low accuracy in text descriptions of generated input images, mainly because the model may bias when understanding the concepts, introducing concepts that are independent of the core content of the image.
By obtaining the text corpus, the entity representation training is performed using the preset graphic and text alignment representation model and the comparison learning loss function to generate the target entity representation matrix; then, the preset language processing model is used to extract the training embedding, and the decoder model is trained based on the cross entropy loss function; finally, the text description corresponding to the image is generated by the combination of the target embedding extraction and the decoder model.
Improve the accuracy of text description, and reduce the generation of irrelevant content by better capturing object information in the image.
Smart Images

Figure CN119942557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a zero-sample image-to-text conversion method and device. Background Art
[0002] Image-to-text conversion is a technique that converts image input into corresponding text descriptions. Its implementation relies on model training on a large number of image-text pairs. By leveraging cross-modal alignment information from pre-trained models, text representations are used to mimic image representations, while introducing noise to bridge the gap between modalities.
[0003] In the multimodal field, large-scale image-text alignment models such as Contrastive Language-Image Pre-training (CLIP) are emerging. These pre-trained models are trained on massive amounts of image-text pairs to achieve alignment across different modalities. These models have demonstrated excellent performance in discriminative tasks such as classification, segmentation, and detection. However, further research is needed to effectively transfer the perceptual capabilities of these models to zero-shot generation tasks, such as image-to-text conversion.
[0004] Most existing zero-shot image-to-text methods focus on integrating detected concepts (such as entities) into the input to generate a textual description of the image. However, the models used in these methods may have biased understanding of concepts and may introduce concepts that are irrelevant to the core content of the image. This can cause the generated textual description of the input image to contain more irrelevant content, resulting in lower accuracy of the generated textual description of the input image. Summary of the Invention
[0005] The present invention provides a zero-shot image-to-text conversion method and device, which are used to solve the technical problem that the existing zero-shot image-to-text conversion method results in low accuracy of the text description of the input image generated.
[0006] A first aspect of the present invention provides a zero-sample image-to-text conversion method, comprising:
[0007] Obtain a text corpus, perform entity representation training based on the text corpus using a preset image-text alignment representation model and a preset contrastive learning loss function, and generate a target entity representation matrix;
[0008] Using the preset image-text alignment representation model and the preset language processing model to perform training embedding extraction based on the model training text in the text corpus to generate hard and soft embeddings to be trained;
[0009] Based on a preset cross entropy loss function, the initial decoder model is trained according to the hard and soft embeddings to be trained to determine the target decoder model;
[0010] Performing target embedding extraction based on the image to be converted and the target entity representation matrix through the preset image-text alignment representation model and the preset language processing model to generate target hard and soft embeddings;
[0011] The target decoder model is used to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted.
[0012] Optionally, the adopting a preset image-text alignment representation model and a preset contrastive learning loss function to perform entity representation training based on the text corpus to generate a target entity representation matrix includes:
[0013] Counting the number of noun entities in the text corpus to determine the number of occurrences of each noun entity;
[0014] Taking any noun entity corresponding to a number of occurrences greater than a preset number threshold as a noun entity to be trained, and initializing the entity vector corresponding to each noun entity to be trained, and determining the initial entity vector corresponding to each noun entity to be trained;
[0015] Using a plurality of the initial entity vectors, constructing an initial entity representation matrix;
[0016] Using a preset image-text alignment representation model to simulate image representation of the matrix training text in the text corpus to generate a matrix training image representation;
[0017] Using a preset grammar parsing tool to construct positive and negative example entity sets based on the noun entity set, and determining the positive and negative example entity sets;
[0018] Performing similarity calculation on the matrix training image representation and the entity vectors corresponding to the positive and negative example entities in the positive and negative example entity sets to determine the positive and negative example similarity;
[0019] Substituting the positive and negative example similarities into a preset contrastive learning loss function and taking the derivative to determine the contrastive learning gradient;
[0020] Using the contrastive learning gradient to update the initial entity representation matrix, determine an intermediate entity representation matrix, and count the number of matrix updates in real time;
[0021] Determine whether the matrix update times reaches a preset first training times threshold;
[0022] If it is achieved, the intermediate entity representation matrix is used as the target entity representation matrix.
[0023] Optionally, the hard and soft embeddings to be trained include hard embeddings to be trained and soft embeddings to be trained; the preset language processing model includes a forward perceptron module and a grammatical parser; and the using of the preset image-text alignment representation model and the preset language processing model to perform training embedding extraction based on the model training text in the text corpus to generate the hard and soft embeddings to be trained includes:
[0024] Using a preset image-text alignment representation model to simulate image representation of the model training text in the text corpus to generate a model training image representation;
[0025] Performing vector projection on the model training image representation through a forward perceptron module to generate a soft embedding to be trained;
[0026] Multiple noun entities in the model training text are used as input to the grammatical parser, and hard embeddings to be trained are output.
[0027] Optionally, the performing model training on the initial decoder model based on the preset cross entropy loss function and the hard and soft embedding to be trained to determine the target decoder model includes:
[0028] splicing the hard embedding to be trained and the soft embedding to be trained to generate a spliced embedding to be trained;
[0029] Substituting the spliced embedding to be trained into a preset cross entropy loss function and taking the derivative to determine the cross entropy gradient;
[0030] Using the cross entropy gradient to update the model parameters of the initial decoder model, determine the intermediate decoder model, and count the number of model updates in real time;
[0031] Determine whether the model update times reaches a preset second training times threshold;
[0032] If achieved, the intermediate decoder model is used as the target decoder model.
[0033] Optionally, the target hard and soft embedding includes a target soft embedding and a target hard embedding; the target entity representation matrix includes target entity vectors corresponding to a plurality of noun entities to be trained; and the target hard and soft embeddings are generated by performing target embedding extraction based on the image to be converted and the target entity representation matrix using the preset image-text alignment representation model and the preset language processing model, including:
[0034] Inputting the image to be converted into the preset image-text alignment representation model for representation to generate a representation of the image to be converted;
[0035] Performing vector projection on the image representation to be converted using a forward perceptron module to generate a target soft embedding;
[0036] Calculate the similarity between the image representation to be converted and the target entity vector corresponding to each noun entity to be trained, and determine the target similarity corresponding to each noun entity to be trained;
[0037] Any noun entity to be trained corresponding to a target similarity greater than a preset similarity threshold is taken as a target noun entity;
[0038] The plurality of target noun entities are used as inputs of a grammatical parser, and a target hard embedding is output.
[0039] Optionally, performing text conversion according to the target hard and soft embedding using the target decoder model to generate a text description corresponding to the image to be converted includes:
[0040] splicing the target soft embedding and the target hard embedding to generate a target spliced embedding;
[0041] The target concatenated embedding is converted using the target decoder model to generate a text description corresponding to the image to be converted.
[0042] A second aspect of the present invention provides a zero-sample image-to-text conversion device, comprising:
[0043] An acquisition module is used to acquire a text corpus, perform entity representation training based on the text corpus using a preset image-text alignment representation model and a preset contrastive learning loss function, and generate a target entity representation matrix;
[0044] A first extraction module is configured to use the preset image-text alignment representation model and the preset language processing model to perform training embedding extraction based on the model training text in the text corpus to generate hard and soft embeddings to be trained;
[0045] A training module, configured to perform model training on an initial decoder model based on a preset cross entropy loss function and the hard and soft embeddings to be trained, and determine a target decoder model;
[0046] A second extraction module is configured to extract a target embedding based on the image to be converted and the target entity representation matrix using the preset image-text alignment representation model and the preset language processing model to generate a target hard and soft embedding;
[0047] A conversion module is used to use the target decoder model to perform text conversion according to the target hard and soft embedding to generate a text description corresponding to the image to be converted.
[0048] A third aspect of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the zero-sample image-to-text conversion method as described in any one of the above items.
[0049] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the above-mentioned zero-sample image-to-text conversion methods when executed.
[0050] A fifth aspect of the present invention provides a computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, wherein when the program instructions are executed by a computer, the computer is caused to perform the steps of the zero-sample image-to-text conversion method as described in any one of the above items.
[0051] It can be seen from the above technical solutions that the present invention has the following advantages:
[0052] The above technical solution of the present invention provides a zero-sample image-to-text conversion method, which first obtains a text corpus, adopts a preset image-text alignment representation model and a preset contrast learning loss function to perform entity representation training based on the text corpus, and generates a target entity representation matrix; then, adopts a preset image-text alignment representation model and a preset language processing model to perform training embedding extraction based on the model training text in the text corpus, and generates hard and soft embeddings to be trained; based on a preset cross-entropy loss function, the initial decoder model is trained according to the hard and soft embeddings to be trained, and the target decoder model is determined; the target embedding extraction is performed according to the image to be converted and the target entity representation matrix through the preset image-text alignment representation model and the preset language processing model. Take, generate target hard and soft embeddings; finally, use the target decoder model to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted; based on the above scheme, the preset image-text alignment representation model and the preset language processing model are used to perform training embedding extraction according to the model training text in the text corpus, and combined with the preset cross entropy loss function, the initial decoder model is trained to determine the target decoder model, and then the target decoder model is used to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted. The target decoder model trained with zero-sample images can better capture the object information in the image, thereby improving the accuracy of the text description. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0054] Figure 1A flowchart of the steps of a zero-sample image-to-text conversion method provided in Example 1 of the present invention;
[0055] Figure 2 A schematic diagram of a zero-sample image-to-text conversion framework provided in Example 1 of the present invention;
[0056] Figure 3 This is a structural block diagram of a zero-sample image-to-text conversion device provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0057] Embodiments of the present invention provide a zero-shot image-to-text conversion method and apparatus, which are used to solve the technical problem that the existing zero-shot image-to-text conversion method results in low accuracy of the text description of the input image generated.
[0058] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0059] See also Figure 1 , Figure 1 This is a flowchart of the steps of a zero-sample image-to-text conversion method provided in Example 1 of the present invention.
[0060] The present invention provides a zero-sample image-to-text conversion method, comprising:
[0061] Step 101: obtain a text corpus, use a preset image-text alignment representation model and a preset contrastive learning loss function to perform entity representation training based on the text corpus, and generate a target entity representation matrix.
[0062] The text corpus includes multiple model training texts for initial decoder model training and multiple matrix training texts for initial entity representation matrix training, and both the model training texts and the matrix training texts are composed of multiple noun entities.
[0063] Furthermore, the process of using a preset image-text alignment representation model and a preset contrastive learning loss function to perform entity representation training based on a text corpus and generate a target entity representation matrix can be achieved by executing the following steps S11 to S110:
[0064] Step S11: Count the number of noun entities in the text corpus to determine the number of occurrences of each noun entity;
[0065] Step S12: taking any noun entity corresponding to a number of occurrences greater than a preset threshold as a noun entity to be trained, and initializing the entity vector corresponding to each noun entity to be trained to determine the initial entity vector corresponding to each noun entity to be trained;
[0066] Step S13: using multiple initial entity vectors to construct an initial entity representation matrix;
[0067] It should be noted that, from all the matrix training texts and all the model training texts in the text corpus, all noun entities (noun entities to be trained) whose occurrence times exceed the preset threshold are extracted to construct a noun entity set, and the entity vectors corresponding to all the noun entities to be trained in the noun entity set are randomly initialized to obtain a trainable initial entity representation matrix; wherein, the preset threshold can be set as needed, and the present invention is not limited to this.
[0068] Step S14: using a preset image-text alignment representation model to perform image representation simulation on the matrix training text in the text corpus to generate a matrix training image representation;
[0069] It should be noted that for each matrix training text in the text corpus, the image-text alignment representation model CLIP (preset image-text alignment representation model) can be used for representation. Because the input required for inference is the image representation, and the image representation cannot be obtained during training under the zero-sample setting, the alignment characteristics of the preset image-text alignment representation model can be used to simulate the corresponding image representation by adding Gaussian noise to the text representation. Specifically, the matrix training text input to the preset image-text alignment representation model is first represented to obtain the corresponding text representation, and then Gaussian noise is added to the text representation to simulate the matrix training image representation. This process can be expressed as:
[0070] ;
[0071] in, Train image representation for the i-th matrix; Train text representation for the i-th matrix; is the selected Gaussian noise variance; is Gaussian noise; It is a preset image-text alignment representation model.
[0072] Step S15: Using a preset grammar parsing tool to construct positive and negative example entity sets based on the noun entity set, and determining the positive and negative example entity sets;
[0073] The positive and negative entity sets include positive entity sets and negative entity sets.
[0074] It should be noted that the grammatical parsing tool NLTK is used to extract the noun phrases contained in the noun entity set to form a positive entity set. Randomly select N' negative entities, exclude the entities in the positive entity set, and form a negative entity set Specifically, the grammar parsing tool NLTK is used to extract multiple noun phrases from the noun entity set to form a positive entity set. , then randomly select N' noun entities in the noun entity set, and exclude the noun entities in the positive entity set to form the negative entity set .
[0075] Step S16: performing similarity calculation on the matrix training image representation and the entity vectors corresponding to the positive and negative example entities in the positive and negative example entity sets to determine the positive and negative example similarity;
[0076] The positive and negative example similarity includes positive example similarity and negative example similarity.
[0077] It should be noted that based on the positive and negative entity sets obtained in the above steps, the image representation I can be calculated i The similarity between the matrix training image representation and the positive entity representation in the positive entity set and the negative entity representation in the negative entity set:
[0078] ;
[0079] ;
[0080] in, is the positive example similarity; Train image representation for the i-th model; is the entity vector corresponding to the positive entity; is the positive entity set; f is the entity in the positive entity set; is the negative example similarity; is the negative entity set; is the entity vector corresponding to the negative entity; m is the entity in the negative entity set.
[0081] It is worth mentioning that since an entity is given multiple types of representations, the similarity of the type with the highest similarity is used as the similarity between the image and the entity when searching, and the image representation I i Take the similarity between the matrix training image representation and the negative entity representation in the negative entity set as an example:
[0082] ;
[0083] in, Train image representation for the i-th model; is the entity representation of the j-th type; K is the number of entity types; is the entity vector corresponding to the negative entity.
[0084] Step S17: Substitute the similarity between positive and negative examples into a preset contrastive learning loss function and derive it to determine the contrastive learning gradient;
[0085] Step S18: Update the initial entity representation matrix using contrastive learning gradient, determine the intermediate entity representation matrix, and count the number of matrix updates in real time;
[0086] Step S19: determine whether the matrix update times reaches a preset first training times threshold;
[0087] Step S110: If the result is reached, the intermediate entity representation matrix is used as the target entity representation matrix.
[0088] It should be noted that if the number of matrix updates does not reach the preset first training number threshold, the intermediate entity representation matrix is used as the new initial entity representation matrix, and a new matrix training text is selected from the text corpus for image representation simulation to generate a new matrix training image representation, and the entity vectors corresponding to the new positive and negative entities are selected from the positive and negative entity sets to perform similarity calculation with the new matrix training image representation to determine the new positive and negative similarity, and then jump to execute step S17 until the number of matrix updates reaches the preset first training number threshold, and the intermediate entity representation matrix determined when the number of matrix updates reaches the preset first training number threshold is used as the target entity representation matrix; wherein, the preset first training number threshold can be set as needed, and the present invention is not limited to this.
[0089] Furthermore, the contrastive learning loss function is preset, specifically:
[0090] ;
[0091] Where L is the contrastive learning loss value; margin is the tolerable interval threshold during alignment learning; is the positive example similarity; is the negative example similarity.
[0092] It is worth mentioning that when executing matrix training, the parameters of the image-text alignment representation model CLIP (preset image-text alignment representation model) are kept frozen, and only the entity representation matrix training is performed.
[0093] Step 102: Use a preset image-text alignment representation model and a preset language processing model to perform training embedding extraction based on the model training text in the text corpus to generate hard and soft embeddings to be trained.
[0094] The hard and soft embeddings to be trained include hard embeddings to be trained and soft embeddings to be trained.
[0095] The preset language processing model includes a forward perceptron module (MLP, Multilayer Perceptron) and a grammar parser.
[0096] Specifically, step 102 may include the following sub-steps S21-S23:
[0097] Step S21: using a preset image-text alignment representation model to simulate image representation of the model training text in the text corpus to generate a model training image representation;
[0098] Step S22: Perform vector projection on the model training image representation through the forward perceptron module to generate a soft embedding to be trained;
[0099] Step S23: Use multiple noun entities in the model training text as input to the grammatical parser and output hard embeddings to be trained.
[0100] It should be noted that for the input text, i.e., the model training text in the text corpus, the alignment capabilities of the CLIP pre-trained model are utilized to simulate the image representation through the text representation to obtain the model training image representation. In subsequent practical applications, the trained decoder GPT-2 model (target decoder model, Decoder of the Generative Pretrained Transformer 2) is used as the decoder for description generation. Since the representation is a continuous vector and the input of GPT-2 is discrete text, the present invention uses a forward perceptron module (MLP) to project the continuous vector into the decoder's embedding layer to implement information input. This part of the input is defined as soft embedding:
[0101] ;
[0102] in, Soft embedding to be trained; It is the forward perceptron module; To preset the image-text alignment representation model; Train text representation for the i-th model.
[0103] Furthermore, to align image and text information in a more fine-grained manner, a grammatical parser is used to extract the noun entity set contained in the input text, and a hard embedding input with specific semantics is constructed using the following template:
[0104] ;
[0105] in, is the hard embedding to be trained.
[0106] Step 103: Based on a preset cross entropy loss function, the initial decoder model is trained according to the hard and soft embeddings to be trained to determine the target decoder model.
[0107] Specifically, step 103 may include the following sub-steps S31-S35:
[0108] Step S31: concatenate the hard embedding to be trained and the soft embedding to be trained to generate a concatenated embedding to be trained;
[0109] Step S32: Substitute the spliced embedding to be trained into a preset cross entropy loss function and derive it to determine the cross entropy gradient;
[0110] Step S33: Use cross entropy gradient to update the model parameters of the initial decoder model, determine the intermediate decoder model, and count the number of model updates in real time;
[0111] Step S34: determine whether the number of model updates reaches a preset second training number threshold;
[0112] Step S35: If reached, use the intermediate decoder model as the target decoder model.
[0113] It should be noted that soft embedding and hard embedding will be spliced in the embedding layer of the GPT-2 decoder as the input of the decoder, and the corresponding text will be used. (i.e., model training text) as the label of self-decoding supervision, and use the preset cross entropy loss function to train the generative model (initial decoder model); wherein, the preset cross entropy loss function is specifically:
[0114] ;
[0115] in, is the cross entropy loss value; Training text for the model; are the model parameters of the initial decoder model; For all words generated in the previous i moments, it represents the words generated in all historical moments; is the i-th generated word; Soft embedding to be trained; Hard embedding to be trained; For splicing; Concatenate embeddings to be trained.
[0116] Furthermore, if the number of model updates does not reach the preset second training number threshold, the intermediate decoder model is used as the new initial decoder model, and new model training text is selected from the text corpus to generate a new concatenated embedding to be trained, and then the process jumps to step S32 until the number of model updates reaches the preset second training number threshold, and the intermediate decoder model determined when the number of model updates reaches the preset second training number threshold is used as the target decoder model.
[0117] Step 104: Target embedding extraction is performed based on the image to be converted and the target entity representation matrix using a preset image-text alignment representation model and a preset language processing model to generate target hard and soft embeddings.
[0118] Target hard and soft embedding includes target soft embedding and target hard embedding.
[0119] The target entity representation matrix includes target entity vectors corresponding to multiple noun entities to be trained.
[0120] Specifically, step 104 may include the following sub-steps S41-S45:
[0121] Step S41: input the image to be converted into a preset image-text alignment representation model for representation, thereby generating a representation of the image to be converted;
[0122] It should be noted that the input image, that is, the image to be converted, is directly input into the preset image-text alignment representation model for representation to generate a representation of the image to be converted.
[0123] Step S42: Use a forward perceptron module to perform vector projection on the image representation to be converted to generate a target soft embedding;
[0124] Step S43: Calculate similarity between the image representation to be converted and the target entity vector corresponding to each noun entity to be trained, and determine the target similarity corresponding to each noun entity to be trained;
[0125] It should be noted that, based on the above steps, the target entity representation matrix is obtained by training the initial entity representation matrix. The initial entity representation matrix is composed of multiple initial entity vectors corresponding to the noun entities to be trained. When the initial entity representation matrix completes the training, each initial entity vector in the initial entity representation matrix serves as the corresponding target entity vector, thereby forming the target entity representation matrix. Therefore, the target entity representation matrix includes multiple target entity vectors corresponding to the noun entities to be trained.
[0126] Furthermore, the present invention uses cosine similarity to calculate the similarity between the input image (image representation to be converted) and each entity (noun entity to be trained) during inference:
[0127] ;
[0128] in, is the target similarity; is the representation of the i-th image to be converted; is the target entity vector corresponding to the noun entity to be trained; is the temperature coefficient, which is used to control randomness. The higher the temperature, the greater the randomness of entity retrieval; exp is the exponential function; j is the j-th entity.
[0129] Step S44: taking any noun entity to be trained corresponding to a target similarity greater than a preset similarity threshold as a target noun entity;
[0130] Furthermore, several entities matching the input image are selected from the entity set by threshold screening, that is, multiple target noun entities are obtained. This process can be expressed as:
[0131] ;
[0132] in, is the entity set consisting of all target noun entities; is the target similarity; is a given screening probability, i.e., a preset similarity threshold; e is an entity; It is a set of the first k entities with the greatest similarity; k is a given number selected from the range of the highest similarity.
[0133] Step S45: take the multiple target noun entities as input to the grammatical parser and output the target hard embedding.
[0134] Step 105: Use the target decoder model to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted.
[0135] Specifically, step 105 may include the following sub-steps S51-S52:
[0136] Step S51: splicing the target soft embedding and the target hard embedding to generate a target spliced embedding;
[0137] Step S52: Use the target decoder model to convert the target spliced embedding to generate a text description corresponding to the image to be converted.
[0138] It should be noted that the trained generative model (target decoder model) and the retrieved entity set (the set of all target noun entities) are comprehensively utilized to fill in the final generated template and generate the final text description (the text description corresponding to the image to be converted). This process can be expressed as:
[0139]
[0140] ;
[0141] ;
[0142] Where C is the text description corresponding to the image to be converted; Hard embedding for the target; Soft embedding for the target; For splicing; Concatenate embeddings for the target; is the target decoder model; is the representation of the i-th image to be converted; is the entity set consisting of all target noun entities.
[0143] As a comparison of technical effects, we can refer to existing technologies. Image-to-text conversion refers to converting an input image into a corresponding text description. Traditional image-to-text conversion methods typically require a large amount of image-to-text data for training, which poses challenges in data acquisition. In recent years, zero-shot image-to-text conversion methods have attracted attention. These methods use text representations to simulate image representations by leveraging cross-modal alignment information from pre-trained models and introduce noise to bridge the modality gap. However, existing zero-shot image-to-text conversion methods mainly focus on integrating detected concepts (such as entities) into the input, with less research on how to train the matching relationship between images and entities in a zero-shot setting through techniques such as contrastive learning. This results in the generated text containing more irrelevant content and lower accuracy.
[0144] For the above questions, please refer to Figure 2 This paper proposes a zero-shot image-to-text conversion method, which can be divided into three parts: entity representation training, text generation model training (decoder model training), and inference. Specifically, a domain-specific entity dictionary is first constructed and a multi-type learnable representation matrix is assigned to each entity. This matrix is used to learn the relationship between entity representations and image representations, enabling the direct retrieval of multiple noun entities associated with an image based on the similarity between the representations. This involves constructing an entity set and training the representation of each entity. Next, the present invention uses the text representation of the image-text alignment model and Gaussian noise to simulate the image representation. A parser is used to extract the noun entities contained in the text and further align them. The text is then trained on a text generation model for reconstruction tasks. During inference, the training text representation is replaced by the aligned image representation, and the extracted entity concepts are populated with the retrieved entities, achieving image-independent zero-shot image-to-text generation training. The parser extracts ideal entities, which are then concatenated with soft embeddings, and the decoder is trained with autoregression. Finally, during the inference phase, the trained representation matrix is used to retrieve the image-aligned entities. Combined with the feature vectors, these entities are input into the decoder to obtain a textual description of the image.
[0145] In summary, the present invention aims to simultaneously improve image features and entity information to enhance the understanding and generation capabilities of language models. Through a multi-type entity representation framework, multiple types of representation vectors are explicitly assigned to entities corresponding to the image, thereby effectively improving the semantic representation completeness of the entity, so as to retrieve more entities that match the input image more accurately, align image and text information at multiple granularities, and thus generate more text descriptions. This embodiment enables the image-to-text conversion model trained by the zero-shot method to better capture object information in the image, and can be used in image description tasks. It also has certain reference significance for the same problems existing in other multimodal zero-shot tasks.
[0146] In an embodiment of the present invention, the present invention provides a zero-sample image-to-text conversion method, which first obtains a text corpus, uses a preset image-to-text alignment representation model and a preset contrast learning loss function to perform entity representation training based on the text corpus, and generates a target entity representation matrix; then, uses a preset image-to-text alignment representation model and a preset language processing model to perform training embedding extraction based on the model training text in the text corpus, and generates hard and soft embeddings to be trained; based on a preset cross-entropy loss function, the initial decoder model is model trained according to the hard and soft embeddings to be trained to determine the target decoder model; and the target embedding is performed according to the image to be converted and the target entity representation matrix through the preset image-to-text alignment representation model and the preset language processing model. Extract and generate target hard and soft embeddings; finally, use the target decoder model to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted; based on the above scheme, a preset image-text alignment representation model and a preset language processing model are used to perform training embedding extraction according to the model training text in the text corpus, and combined with a preset cross-entropy loss function, the initial decoder model is trained to determine the target decoder model, and then the target decoder model is used to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted. The target decoder model trained with zero-sample images can better capture the object information in the image, thereby improving the accuracy of the text description.
[0147] See also Figure 3 , Figure 3 This is a structural block diagram of a zero-sample image-to-text conversion device provided in Example 2 of the present invention.
[0148] The present invention provides a zero-sample image-to-text conversion device, comprising:
[0149] An acquisition module 301 is used to acquire a text corpus, perform entity representation training based on the text corpus using a preset image-text alignment representation model and a preset contrastive learning loss function, and generate a target entity representation matrix;
[0150] A first extraction module 302 is configured to extract training embeddings based on model training texts in a text corpus using a preset image-text alignment representation model and a preset language processing model to generate hard and soft embeddings to be trained;
[0151] A training module 303 is configured to perform model training on the initial decoder model based on a preset cross entropy loss function and the hard and soft embeddings to be trained to determine a target decoder model;
[0152] A second extraction module 304 is configured to extract target embeddings based on the image to be converted and the target entity representation matrix using a preset image-text alignment representation model and a preset language processing model to generate target hard and soft embeddings;
[0153] The conversion module 305 is used to perform text conversion according to the target hard and soft embedding using the target decoder model to generate a text description corresponding to the image to be converted.
[0154] Furthermore, the acquisition module 301 is specifically configured to:
[0155] Count the number of noun entities in the text corpus and determine the number of occurrences of each noun entity;
[0156] Any noun entity corresponding to a number of occurrences greater than a preset threshold is taken as a noun entity to be trained, and the entity vector corresponding to each noun entity to be trained is initialized to determine the initial entity vector corresponding to each noun entity to be trained;
[0157] Using multiple initial entity vectors, construct an initial entity representation matrix;
[0158] Use the preset image-text alignment representation model to simulate the image representation of the matrix training text in the text corpus to generate the matrix training image representation;
[0159] Use the preset grammar parsing tool to construct positive and negative entity sets based on the noun entity set to determine the positive and negative entity sets;
[0160] Calculate the similarity between the matrix training image representation and the entity vectors corresponding to the positive and negative entities in the positive and negative entity sets to determine the positive and negative similarity;
[0161] Substitute the similarity between positive and negative examples into the preset contrastive learning loss function and take the derivative to determine the contrastive learning gradient;
[0162] Use contrastive learning gradient to update the initial entity representation matrix, determine the intermediate entity representation matrix, and count the number of matrix updates in real time;
[0163] Determine whether the number of matrix updates reaches a preset first training number threshold;
[0164] If achieved, the intermediate entity representation matrix is used as the target entity representation matrix.
[0165] Furthermore, the hard and soft embeddings to be trained include hard embeddings to be trained and soft embeddings to be trained; the preset language processing model includes a forward perceptron module and a grammar parser; the first extraction module 302 is specifically used to:
[0166] Use the preset image-text alignment representation model to simulate the image representation of the model training text in the text corpus to generate the model training image representation;
[0167] Perform vector projection on the model training image representation through the forward perceptron module to generate the soft embedding to be trained;
[0168] Multiple noun entities in the model training text are used as input to the grammatical parser, and the output is the hard embedding to be trained.
[0169] Furthermore, the training module 303 is specifically configured to:
[0170] Concatenate the hard embedding to be trained and the soft embedding to be trained to generate the concatenated embedding to be trained;
[0171] Substitute the spliced embedding to be trained into the preset cross entropy loss function and take the derivative to determine the cross entropy gradient;
[0172] Use cross-entropy gradient to update the model parameters of the initial decoder model, determine the intermediate decoder model, and count the number of model updates in real time;
[0173] Determine whether the number of model updates reaches a preset second training number threshold;
[0174] If achieved, the intermediate decoder model is used as the target decoder model.
[0175] Furthermore, the target hard and soft embedding includes target soft embedding and target hard embedding; the second extraction module 304 is specifically configured to:
[0176] Input the image to be converted into a preset image-text alignment representation model for representation, and generate a representation of the image to be converted;
[0177] A forward perceptron module is used to perform vector projection on the image representation to be transformed and generate a target soft embedding.
[0178] Calculate the similarity between the image representation to be converted and the target entity vector corresponding to each noun entity to be trained, and determine the target similarity corresponding to each noun entity to be trained;
[0179] Any noun entity to be trained corresponding to a target similarity greater than a preset similarity threshold is taken as a target noun entity;
[0180] Multiple target noun entities are taken as input to the grammatical parser, and the target hard embedding is output.
[0181] Furthermore, the conversion module 305 is specifically configured to:
[0182] Concatenate the target soft embedding and the target hard embedding to generate the target concatenated embedding;
[0183] The target decoder model is used to transform the target concatenated embedding to generate a text description corresponding to the image to be transformed.
[0184] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0185] An embodiment of the present invention further provides a computer device including a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the zero-sample image-to-text conversion method of the first embodiment.
[0186] An embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the zero-sample image-to-text conversion method of the first embodiment are implemented.
[0187] An embodiment of the present invention further provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the steps of the zero-sample image-to-text conversion method of the first embodiment are implemented.
[0188] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0189] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0190] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A zero-sample image-to-text conversion method, characterized in that: include: Acquire a text corpus, use a preset image-text alignment representation model and a preset contrastive learning loss function to perform entity representation training based on the text corpus, and generate a target entity representation matrix; Using the preset image-text alignment representation model and the preset language processing model to perform training embedding extraction based on the model training text in the text corpus to generate hard and soft embeddings to be trained; Based on a preset cross entropy loss function, the initial decoder model is trained according to the hard and soft embeddings to be trained to determine a target decoder model; Performing target embedding extraction according to the image to be converted and the target entity representation matrix through the preset image-text alignment representation model and the preset language processing model to generate target hard and soft embedding; The target decoder model is used to perform text conversion according to the target hard and soft embeddings to generate a text description corresponding to the image to be converted.
2. The zero-sample image-to-text conversion method according to claim 1, characterized in that: The method of using a preset image-text alignment representation model and a preset contrastive learning loss function to perform entity representation training according to the text corpus to generate a target entity representation matrix includes: Performing quantitative statistics on multiple noun entities in the text corpus to determine the number of occurrences of each noun entity; Taking any noun entity corresponding to a number of occurrences greater than a preset number threshold as a noun entity to be trained, and initializing the entity vector corresponding to each noun entity to be trained, and determining the initial entity vector corresponding to each noun entity to be trained; Using a plurality of the initial entity vectors, constructing an initial entity representation matrix; Using a preset image-text alignment representation model to simulate the image representation of the matrix training text in the text corpus to generate a matrix training image representation; Using a preset grammar parsing tool to construct positive and negative example entity sets according to the noun entity set, and determining the positive and negative example entity sets; Performing similarity calculation on the matrix training image representation and the entity vectors corresponding to the positive and negative example entities in the positive and negative example entity sets to determine the positive and negative example similarity; Substituting the positive and negative example similarities into a preset contrastive learning loss function and taking the derivative to determine a contrastive learning gradient; The initial entity representation matrix is updated using the contrastive learning gradient, an intermediate entity representation matrix is determined, and the number of matrix updates is counted in real time; Determine whether the matrix update times reaches a preset first training times threshold; If it is achieved, the intermediate entity representation matrix is used as the target entity representation matrix.
3. The zero-sample image-to-text conversion method according to claim 1, characterized in that: The hard and soft embeddings to be trained include hard embeddings to be trained and soft embeddings to be trained; the preset language processing model includes a forward perceptron module and a grammar parser; the using the preset image-text alignment representation model and the preset language processing model to perform training embedding extraction according to the model training text in the text corpus to generate the hard and soft embeddings to be trained includes: Using a preset image-text alignment representation model to simulate image representation of the model training text in the text corpus to generate a model training image representation; Performing vector projection on the model training image representation through a forward perceptron module to generate a soft embedding to be trained; A plurality of noun entities in the model training text are used as inputs of a grammatical parser, and hard embeddings to be trained are output.
4. The zero-sample image-to-text conversion method according to claim 3, characterized in that: The method of performing model training on the initial decoder model based on the preset cross entropy loss function and the hard and soft embedding to be trained to determine the target decoder model includes: Concatenating the hard embedding to be trained and the soft embedding to be trained to generate a concatenated embedding to be trained; Substituting the spliced embedding to be trained into a preset cross entropy loss function and taking the derivative to determine the cross entropy gradient; The cross entropy gradient is used to update the model parameters of the initial decoder model, determine the intermediate decoder model, and count the number of model updates in real time; Determine whether the model update times reaches a preset second training times threshold; If achieved, the intermediate decoder model is used as the target decoder model.
5. The zero-sample image-to-text conversion method according to claim 3, characterized in that: The target hard and soft embedding includes a target soft embedding and a target hard embedding; the target entity representation matrix includes a plurality of target entity vectors corresponding to the noun entities to be trained; the target embedding extraction is performed according to the image to be converted and the target entity representation matrix by the preset image-text alignment representation model and the preset language processing model to generate the target hard and soft embedding, including: Inputting the image to be converted into the preset image-text alignment representation model for representation, and generating a representation of the image to be converted; Using a forward perceptron module to perform vector projection on the image representation to be converted to generate a target soft embedding; Calculate the similarity between the image representation to be converted and the target entity vector corresponding to each noun entity to be trained, and determine the target similarity corresponding to each noun entity to be trained; Taking any noun entity to be trained corresponding to a target similarity greater than a preset similarity threshold as a target noun entity; The plurality of target noun entities are used as inputs of a grammatical parser, and a target hard embedding is output.
6. The zero-sample image-to-text conversion method according to claim 1, characterized in that: The using the target decoder model to perform text conversion according to the target hard and soft embedding to generate a text description corresponding to the image to be converted includes: splicing the target soft embedding and the target hard embedding to generate a target spliced embedding; The target decoder model is used to transform the target concatenated embedding to generate a text description corresponding to the image to be transformed.
7. A zero-sample image-to-text conversion device, characterized in that: include: An acquisition module is used to acquire a text corpus, perform entity representation training based on the text corpus using a preset image-text alignment representation model and a preset contrastive learning loss function, and generate a target entity representation matrix; A first extraction module, configured to use the preset image-text alignment representation model and the preset language processing model to perform training embedding extraction according to the model training text in the text corpus to generate hard and soft embeddings to be trained; A training module, used for performing model training on an initial decoder model based on a preset cross entropy loss function and the hard and soft embeddings to be trained to determine a target decoder model; A second extraction module, configured to perform target embedding extraction according to the image to be converted and the target entity representation matrix through the preset image-text alignment representation model and the preset language processing model, and generate target hard and soft embedding; A conversion module is used to use the target decoder model to perform text conversion according to the target hard and soft embedding to generate a text description corresponding to the image to be converted.
8. A computer device, characterized in that: The invention comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the zero-sample image-to-text conversion method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the zero-sample image-to-text conversion method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, wherein when the program instructions are executed by a computer, the computer is caused to execute the zero-sample image-to-text conversion method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text generation image model training method and text generation image method and device
CN118015637A
Lightweight multi-modal image description generation method based on CLIP encoder
CN118069877A
Systems and methods for vision-and-language representation learning
US20220391755A1