Image description generation model training method, image description generation method, equipment, medium and product
By inserting Chinese vocabulary with high semantic similarity into the vocabulary of the image description generation model, an extended vocabulary was formed, and the model was trained based on this table, the problem of the model's performance degradation in Chinese scenarios was solved, and the ability to directly process Chinese was realized, and performance performance and generalization were improved.
Patent Information
- Application Number
- CN202510179565.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-27
AI Technical Summary
The performance of the existing image description generation model has significantly decreased in Chinese scenarios, mainly because the model only supports English vocabulary and cannot directly process Chinese text, so it is necessary to translate Chinese into English through a translation network.
By obtaining the image description generation model, original vocabulary and new vocabulary library to be trained, find the target original vocabulary with semantic similarity greater than the preset threshold, insert the added Chinese vocabulary to a position adjacent to the target original vocabulary, form an extended vocabulary, and train the image description generation model based on the vocabulary.
This enables the image description generation model to directly process Chinese vocabulary, avoids intermediate steps in the translation network, improves the model's performance in Chinese scenarios, and enhances the model's semantic understanding and generalization capabilities.
Smart Images

Figure CN120220147A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to an image description generation model training method, an image description generation method, a device, a medium and a product. Background Art
[0002] Image description, or image annotation, is a technology that converts image content into natural language descriptions. Image description is currently widely used in various scenarios, such as large-scale question answering, visual search, automatic image annotation, content recommendation, and assistance for people with visual impairments. The accuracy of image description is directly related to the effectiveness of these applications and user experience, and has important practical significance.
[0003] The generation of image descriptions is mostly achieved by using an image description generation model based on a deep learning framework. That is, the basic features of the image are extracted through the image description generation model, and corresponding description sentences are generated based on the basic features to obtain the image description text. The training of the image description generation model often requires a large amount of training corpus, including training images and corresponding description sentences. However, the current models used for image description generation are usually trained based on English vocabulary, and the models only support English vocabulary. In Chinese scenarios, a translation network is needed to translate Chinese into English, and Chinese text cannot be directly processed, resulting in a significant decrease in the performance of the image description generation model in Chinese scenarios.
[0004] Therefore, how to improve the performance of image description generation models in Chinese scenarios is a technical problem that needs to be solved urgently in this technical field. Summary of the invention
[0005] The main purpose of this application is to provide an image description generation model training method, image description generation method, device, medium and product, aiming to solve the technical problem of how to improve the performance of the image description generation model in Chinese scenarios.
[0006] To achieve the above object, the present application provides an image description generation model training method, the image description generation model training method comprising the following steps:
[0007] Obtaining an image description generation model to be trained, an original vocabulary table and a newly added vocabulary library, wherein the newly added vocabulary library includes at least one newly added Chinese vocabulary, and the original vocabulary table includes at least one original vocabulary;
[0008] For each newly added Chinese word in the newly added vocabulary library, searching the original vocabulary in the original vocabulary for a target word whose semantic similarity with the newly added Chinese word is greater than a preset threshold, and inserting the newly added Chinese word into a position adjacent to the target original word;
[0009] until all the newly added Chinese words in the newly added vocabulary library are inserted into the original vocabulary list to obtain an extended vocabulary list;
[0010] Train the image description generation model based on the extended vocabulary list to obtain the trained image description generation model.
[0011] In one embodiment, before the step of training the image description generation model based on the extended vocabulary list, the method further includes:
[0012] Initialize the embedding vectors of each original word in the extended vocabulary list;
[0013] For each newly added Chinese word in the extended vocabulary list, determine the embedding vector of the target original word adjacent to the newly added Chinese word as the target embedding vector, and initialize the embedding vector of the newly added Chinese word with the target embedding vector.
[0014] In one embodiment, the step of finding the target original word in the original vocabulary list whose semantic similarity with the newly added Chinese word is greater than a preset threshold includes:
[0015] Input the newly added Chinese word into a pre-trained translation network to obtain a translated word;
[0016] Compare the semantic similarity between each original word in the original vocabulary list and the translated word, and find the target original word whose semantic similarity is greater than the preset threshold from each original word.
[0017] In addition, to achieve the above object, the present application also provides an image description generation method, and the image description generation method includes the following steps:
[0018] Obtain the original image to be generated with an image description and the target image description generation model, and input the original image into the target image description generation model to obtain a description generation result;
[0019] Wherein, the target image description generation model is an image description generation model trained by using the above-mentioned image description generation model training method.
[0020] In one embodiment, after the step of inputting the original image into the target image description generation model to obtain a description generation result, the method further includes:
[0021] Retrieve the target data matching the description generation result in a preset text retrieval database.
[0022] In one embodiment, the description generation result is the output data of the output layer of the target image description generation model. The step of retrieving target data matching the description generation result in the preset text retrieval database includes:
[0023] Perform vectorization processing on the description generation result to obtain a description vector, and retrieve a target index vector matching the description vector in the preset text retrieval database;
[0024] Determine that the stored data associated with the target index vector in the preset text retrieval database is the retrieved target data.
[0025] In one embodiment, the description generation result is the input data of the output layer of the target image description generation model. The step of retrieving target data matching the description generation result in the preset text retrieval database includes:
[0026] Retrieve a target index vector matching the description generation result in the preset text retrieval database;
[0027] Determine that the stored data associated with the target index vector in the preset text retrieval database is the retrieved target data.
[0028] In addition, to achieve the above object, the present application further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the above-mentioned image description generation model training method and / or image description generation method.
[0029] In addition, to achieve the above object, the present application further provides a readable storage medium, which is a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and the computer program is executed by a processor to implement the steps of the above-mentioned image description generation model training method and / or image description generation method.
[0030] The present application further provides a computer program product, including a computer program, and the computer program implements the steps of the above-mentioned image description generation model training method and / or image description generation method when executed by a processor.
[0031] One or more technical solutions proposed by the present application have at least the following technical effects:
[0032] Obtain an image caption generation model to be trained, an original vocabulary, and a new vocabulary library. The new vocabulary library includes at least one new Chinese word, and the original vocabulary includes at least one original word. For each new Chinese word in the new vocabulary library, find the target original word in the original vocabulary whose semantic similarity to the new Chinese word is greater than a preset threshold, and insert the new Chinese word at a position adjacent to the target original word. Until all new Chinese words in the new vocabulary library are inserted into the original vocabulary to obtain an extended vocabulary. Train the image caption generation model based on the extended vocabulary to obtain the trained image caption generation model. In this way, in the embodiment of the present application, at least one new Chinese word is inserted into the original vocabulary of the image caption generation model, that is, the image caption generation model is trained by adding new Chinese words to the vocabulary, so that the model can directly process Chinese words without the need to use a translation network to translate Chinese into English, thereby improving the performance of the image caption generation model in the Chinese scenario. And further, in the embodiment of the present application, the new Chinese word is inserted at a position adjacent to the target original word with similar semantics. It can be understood that in the model, the embedding vector of a word is learned through context information. Placing words with similar semantics together can make these words more likely to be close to each other in the embedding space, which helps the model better capture the semantic relationship between words during training, improve the semantic understanding ability of the model, and further improve the performance of the image caption generation model in the Chinese scenario. At the same time, this structure can reduce the model's dependence on specific words, enable the model to better handle polysemous words and word variants during training, enhance its adaptability in different contexts, and thus improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0034] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 Schematic diagram of the traditional training process for training a BLIP model applicable to the Chinese scenario;
[0036] Figure 2 Schematic diagram of the traditional inference process for using the BLIP model in the Chinese scenario;
[0037] Figure 3 It is a schematic flowchart of the first embodiment of the method for training an image description generation model of the present application;
[0038] Figure 4 It is a schematic flowchart of the brief process of the method for training an image description generation model of the present application;
[0039] Figure 5 It is a schematic diagram of the device structure of the device for training an image description generation model of the present application;
[0040] Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the device for the method of training an image description generation model in an embodiment of the present application.
[0041] The realization of the purpose, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0042] To make the above objects, features and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] Traditional image description generation models, such as the BLIP model (Bootstrapping Language-Image Pre-training), are trained on English datasets, which may result in the situation that many out-of-vocabulary words cannot be recognized in Chinese tasks, and thus the performance of the model in Chinese scenarios is poor.
[0044] In the traditional method, in a Chinese scenario, a translation network is needed to translate Chinese into English so that the model can indirectly process Chinese. For example, as shown in Figure 1 During the training stage, a translation network is needed to translate Chinese training labels into English, and then the translated English is used as the training label of the BLIP model to train the BLIP model; as shown in Figure 2 During the inference stage, the output of the BLIP model needs to be converted back into Chinese using a translation network. This indirect way of processing Chinese has at least two problems: First, adding a translation network to the model will make the model larger and the inference speed slower. Second, during the mutual translation between English and Chinese, some semantics will be lost, thus weakening the ability of the model and significantly reducing the performance of the model in Chinese scenarios.
[0045] Based on this, the main solution of this application is as follows: Obtain an image caption generation model to be trained, an original vocabulary, and a new vocabulary library, where the new vocabulary library includes at least one new Chinese word, and the original vocabulary includes at least one original word; for each new Chinese word in the new vocabulary library, find a target original word in the original vocabulary whose semantic similarity to the new Chinese word is greater than a preset threshold, and insert the new Chinese word at a position adjacent to the target original word; until all the new Chinese words in the new vocabulary library are inserted into the original vocabulary to obtain an extended vocabulary; train the image caption generation model based on the extended vocabulary to obtain the trained image caption generation model. In this way, in the embodiments of this application, by inserting at least one new Chinese word into the original vocabulary of the image caption generation model, that is, training the image caption generation model by adding new Chinese words to the vocabulary, the model can directly process Chinese words without the need to use a translation network to translate Chinese into English, thereby improving the performance of the image caption generation model in the Chinese scenario. And further, this application inserts the new Chinese word at a position adjacent to the target original word with similar semantics. It can be understood that in the model, the embedding vector of a word is learned through context information. Placing words with similar semantics together can make these words more likely to be close to each other in the embedding space, which helps the model better capture the semantic relationship between words during training, improve the semantic understanding ability of the model, and further improve the performance of the image caption generation model in the Chinese scenario. At the same time, this structure can reduce the model's dependence on specific words, enable the model to better handle polysemous words and word variants during training, enhance its adaptability in different contexts, and thus improve the generalization ability of the model.
[0046] It should be noted that the execution subject of each embodiment of the image caption generation model training method of this application can be a computing service device with data processing, model communication, and program running functions, such as a server, a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions, such as a robot, etc. Each embodiment of the image caption generation model training method of this application does not make specific limitations on this.
[0047] Based on this, this application proposes an image caption generation model training method according to the first embodiment. Please refer to Figure 3 as shown, the image caption generation model training method includes the following steps S10 to S40:
[0048] Step S10: Obtain an image description generation model to be trained, an original vocabulary, and a new vocabulary library. Among them, the new vocabulary library includes at least one new Chinese word, and the original vocabulary includes at least one original word;
[0049] The image description generation model is a model for generating image descriptions, that is, a model that converts images into text descriptions. Specifically, it can be a model constructed based on a deep learning architecture, such as a BLIP model, a CLIP (Contrastive Language-Image Pre-training) model, etc. This embodiment does not make specific limitations on this. Exemplarily, the BLIP model is used as the image description generation model to elaborate and illustrate the embodiments of the present application.
[0050] The original vocabulary is the vocabulary currently used to train the BLIP model, such as the WordPiece vocabulary, which contains the vocabulary set used by the model in the training stage. Each original word in the original vocabulary is usually a common English word, sub-word, and some special tokens, etc. For example, the original vocabulary may include common words such as "man", "woman", "tree", "car", and sub-word units such as "##ing", "##ed", etc.
[0051] The new vocabulary library refers to the newly added vocabulary library, which contains at least one new Chinese word. The new Chinese word is a word in the Chinese language. The new Chinese words in the new vocabulary library can be sourced from a Chinese vocabulary, such as the vocabulary used to train the Chinese BERT (Bidirectional Encoder Representations from Transformers) model, a professional term library, a user-defined vocabulary list (such as in a game application scenario, "Sun Wukong", "Chang'e", "Peach Orchard", "Dragon Palace", etc.), newly emerging words dynamically crawled, etc. This embodiment does not make specific limitations on this.
[0052] Step S20: For each new Chinese word in the new vocabulary library, find the target original word in the original vocabulary whose semantic similarity to the new Chinese word is greater than a preset threshold, and insert the new Chinese word at a position adjacent to the target original word;
[0053] It should be noted that if the target original word is not found in the original vocabulary, that is, there is no target original word in the original vocabulary with a semantic similarity greater than the preset threshold among the newly added Chinese words, then the newly added Chinese word can be inserted into any position in the original vocabulary, such as inserting the newly added Chinese word at the end of the original vocabulary. If there are multiple original words in the original vocabulary with a semantic similarity greater than the preset threshold to the newly added Chinese word, the original word with the highest semantic similarity can be selected as the target original word, so as to place the words with higher semantic similarity together, which helps the model better handle semantic similarity tasks, such as word similarity calculation and analogical reasoning.
[0054] After finding the target original word in the original vocabulary, insert the newly added Chinese word at a position adjacent to the target original word, that is, insert the newly added Chinese word into the original vocabulary at a position adjacent to the target original word. The position adjacent to the target original word can specifically be a position before or after the target original word.
[0055] Step S30, until all the newly added Chinese words in the newly added vocabulary library are inserted into the original vocabulary to obtain an extended vocabulary.
[0056] After inserting all the newly added Chinese words in the newly added vocabulary library into the original vocabulary according to the semantic order, an extended vocabulary is obtained. That is, the extended vocabulary is the vocabulary obtained by inserting all the newly added Chinese words into the original vocabulary.
[0057] Step S40, train the image description generation model based on the extended vocabulary to obtain the trained image description generation model.
[0058] After obtaining the extended vocabulary, train the BLIP model based on the extended vocabulary. During the training process, if the preset training end condition is met, the trained BLIP model is obtained. If the preset training end condition is not met, the model parameters of the BLIP model are iteratively optimized until the preset training end condition is met.
[0059] The training end condition can be a condition set in advance, such as reaching a predetermined number of iterations, the loss function value being lower than a predetermined threshold, the computing resources being exhausted, reaching the time limit, the accuracy rate reaching a predetermined threshold, etc. This embodiment does not make specific limitations on this.
[0060] It can be understood that when training the BLIP model, images are also required as inputs, and the corresponding text descriptions are used as labels to train the BLIP model. Based on this, the BLIP model can be specifically trained using an extended vocabulary and a preset image-text pair dataset. The image-text pair dataset includes multiple images and the text descriptions corresponding to each image. Specifically, a publicly available dataset such as the COCO (Common Objects in Context) dataset can be selected, or it can also be a dataset prepared in advance by relevant personnel based on actual needs. This embodiment does not make specific restrictions on this.
[0061] In this embodiment, an image description generation model to be trained, an original vocabulary, and a new vocabulary library are obtained. The new vocabulary library includes at least one new Chinese word, and the original vocabulary includes at least one original word. For each new Chinese word in the new vocabulary library, a target original word whose semantic similarity to the new Chinese word in the original vocabulary is greater than a preset threshold is found, and the new Chinese word is inserted at a position adjacent to the target original word. Until all the new Chinese words in the new vocabulary library are inserted into the original vocabulary to obtain an extended vocabulary. The image description generation model is trained based on the extended vocabulary to obtain the trained image description generation model. In this way, in this embodiment, at least one new Chinese word is inserted into the original vocabulary of the image description generation model, that is, the image description generation model is trained by adding new Chinese words to the vocabulary, so that the model can directly process Chinese words without the need to use a translation network to translate Chinese into English, thereby improving the performance of the image description generation model in the Chinese scenario. And further, in this embodiment, the new Chinese word is inserted at a position adjacent to the target original word with similar semantics. It can be understood that in the model, the embedding vector of a word is learned through context information. Placing words with similar semantics together can make these words more likely to be close to each other in the embedding space, which helps the model to better capture the semantic relationship between words during training, improve the semantic understanding ability of the model, and further improve the performance of the image description generation model in the Chinese scenario. At the same time, this structure can reduce the model's dependence on specific words, enable the model to better process polysemous words and word variants during training, enhance its adaptability in different contexts, and thus improve the generalization ability of the model.
[0062] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, before the step of training the image description generation model based on the extended vocabulary, the method further includes:
[0063] Step A10, initialize the embedding vectors of each original word in the extended vocabulary;
[0064] It should be noted that the embedding vector refers to a high-dimensional space vector to which a word is mapped in the Embedding layer of the BLIP model. For the original words in the extended vocabulary, that is, the words initially in the original vocabulary, they can be initialized in the current way of initializing the embedding vectors of the BLIP model. For example, before training the BLIP model, a pre-trained text encoder (such as BERT or its variant) will be loaded, and the BLIP model will use the word embedding layer in the pre-trained text encoder to initialize the embedding vectors of each original word in the original vocabulary.
[0065] Step A20, for each newly added Chinese word in the extended vocabulary, determine the embedding vector of the target original word adjacent to the newly added Chinese word as the target embedding vector, and initialize the embedding vector of the newly added Chinese word with the target embedding vector.
[0066] For the newly added Chinese words that are newly added to the original vocabulary, the embedding vectors of these newly added Chinese words are vectorized with the embedding vectors of the target original words adjacent to them (that is, the target embedding vectors). Specifically, for each newly added Chinese word, the distance between its embedding vector and the target embedding vector is less than a certain threshold. For example, the embedding vector of the newly added Chinese word can be initialized with the weighted average or nearby value of the target embedding vector, so that the two words are closer in the embedding space, thereby reducing the number of iterations during model training, accelerating model convergence, and improving model training efficiency.
[0067] Furthermore, for the dimension of the output vector of the last output layer of the BLIP model, it is correspondingly modified to the table dimension of the extended vocabulary, where the table dimension of the extended vocabulary refers to the total number of words in the extended vocabulary. For example, assuming that there are m words in the extended vocabulary in total, the dimension of the output vector of the last output layer of the BLIP model is correspondingly modified to m dimensions.
[0068] Furthermore, for the initialization of the weights of the output vector of the last output layer of the BLIP model, the mean filling method can be adopted, that is, the mean value of the initial weight values of all original words in the original vocabulary in the output vector is used to fill the corresponding weight initial values of the newly added Chinese words in the output vector, so as to make the model parameter distribution of the BLIP model uniform, reduce the number of iterations during model training, and improve the training efficiency of the model.
[0069] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the content that is the same as or similar to the above-mentioned first embodiment and second embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, the step of searching for the target original word in the original vocabulary whose semantic similarity with the newly added Chinese word is greater than a preset threshold includes:
[0070] Step B10, input the newly added Chinese word into a pre-trained translation network to obtain a translated word;
[0071] It should be noted that this translation network is used to translate the newly added Chinese word into a translated word in the same language as the original words in the original vocabulary. If the original vocabulary contains original words belonging to multiple languages, the translation network will translate the newly added Chinese word into the translated word in the language with the highest proportion, where the language with the highest proportion is the language with the highest proportion among the languages to which the original words in the original vocabulary belong. For example, assuming that all or more than half of the original words in the original vocabulary belong to the English language, the translation network will translate the newly added Chinese word into English.
[0072] Step B20, compare the semantic similarity between each original word in the original vocabulary and the translated word, and search for the target original word whose semantic similarity is greater than the preset threshold from each of the original words.
[0073] It can be understood that the calculation and acquisition of semantic similarity between the same language are more convenient and simple than those between different languages. Based on this, in this embodiment, the translation network is used to translate the target original word to obtain a translated word, so that the semantic similarity comparison is performed based on the translated word in the same language as most of the original words, and then the target original word can be found, which can improve the search efficiency of the target original word.
[0074] Exemplarily, in order to help understand the technical concept or technical principle of the image description generation model training method after combining this embodiment with the first embodiment and the second embodiment, a specific embodiment is now listed. In this specific embodiment, please refer to Figure 4 As shown, modifications are made to the text encoder and decoder parts of the BLIP model. First, the original vocabulary of the English BLIP model for text to token (word element) is modified ( Figure 4 The original token word list shown in it). Specifically, Chinese words are added to the original vocabulary to obtain an extended vocabulary ( Figure 4 The extended vocabulary shown in it, adding Chinese). These Chinese words are taken from the vocabulary of the Chinese BERT model and cover all Chinese words.
[0075] After adjusting the original vocabulary, modify the weights of the corresponding Embedding layer in the model. First, modify the Embedding dimension in the original model to the dimension of the extended vocabulary. Second, initialize the weights of this part of the extended dimension. The initialization strategy in this specific implementation uses the form of mean filling, and the mean value comes from the mean value of the Embedding weights in the same layer. Finally, modify the dimension of the last output layer of the model's text decoder, also to the dimension of the extended vocabulary, and the initialization of the weights also uses the form of mean filling.
[0076] After initializing the weights of the Embedding layer ( Figure 4 the initial layer shown in) and the weights of the output layer in the BLIP model, train the BLIP model with the extended vocabulary to obtain the trained BLIP model.
[0077] It should be noted that the above examples are only used to assist in understanding this application and do not constitute a limitation on the method for training the image description generation model of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0078] In addition, an embodiment of this application also proposes an image description generation method, and the image description generation method includes the following step S100:
[0079] Step S100, obtain the original image to be generated with an image description and the target image description generation model, and input the original image into the target image description generation model to obtain a description generation result;
[0080] Among them, the target image description generation model is an image description generation model trained by using the method for training the image description generation model described in any one of the above embodiments.
[0081] The execution subject of each embodiment of the image description generation method can be a computing service device with data processing, model communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions, such as AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, AR helmets, VR helmets, smart glasses, earphones, etc. This embodiment does not make specific limitations on this. Further, the execution subject of the image description generation method and the method for training the image description generation model in this application can be the same or different, and no specific limitations are made here.
[0082] It is easy to understand that after obtaining the original image for which the image description is to be generated and inputting the original image into the target image description generation model, the description generation result can be obtained, and the description generation result can specifically be the output of the target image description generation model.
[0083] In a possible implementation manner, after the step of inputting the original image into the target image description generation model to obtain the description generation result, the method further includes:
[0084] Step S200, retrieving target data that matches the description generation result in a preset text retrieval database.
[0085] The preset text retrieval database can specifically be a database pre-set by relevant personnel, and can specifically be a RAG (Retrieval-Augmented Generation) database. The RAG database is a vector database dedicated to storing and querying high-dimensional vector data. It enables efficient retrieval of vectors closest to the query vector by converting text data into vectors (embeddings) and storing these vectors, thereby finding relevant text content.
[0086] Considering that traditional RAG for processing images uses OCR or large model technologies to parse key information in the image and then perform text encoding. However, when there is no text information in the image or the content of the image cannot be understood by the large model, RAG for the image will fail at this time. Based on this, in a preferred implementation manner, the preset text retrieval database is a RAG database, so that the image is converted into a corresponding text description through a fine-tuned BLIP model, and subsequent RAG retrieval is performed based on the text, realizing the form of encoding the image into RAG and the low-cost retrieval of the image to the RAG database.
[0087] Considering the rapid growth of the types of intelligent AR+AI (Artificial Intelligence) glasses, it has become increasingly easy to capture images by intelligent devices. At the same time, with the emergence of vision large models, it has become very important how to enable the large model to better understand the content in the image. Generally speaking, the vision large model cannot achieve a detailed understanding of images in specific scenarios (such as game scenarios). To achieve a detailed understanding, it is necessary to train and fine-tune the vision large model, but the cost of fine-tuning the large model is very high. Based on this, in this embodiment, the target image description generation model is used to generate the text description of the image, and then the retrieval is performed in the text retrieval database based on the text description, so that the image can be retrieved in text-based data, realizing an end-to-end large model + text retrieval solution, thereby improving the visual ability boundary of the large model in a low-cost manner and better adapting to specific scenarios.
[0088] In a possible implementation, the description generation result is the output data of the output layer of the target image description generation model. The step of retrieving target data matching the description generation result in the preset text retrieval database includes:
[0089] Step S201: Perform vectorization processing on the description generation result to obtain a description vector, and retrieve a target index vector matching the description vector in the preset text retrieval database;
[0090] It should be noted that the output layer specifically refers to the last layer network of the target image description generation model, which aims to convert the internal feature representation of the model into the final prediction result, so that the model can directly output the interpretation or prediction of the input data. The description generation result is the output data of the output layer of the target image description generation model, that is, the final output of the target image description generation model.
[0091] If the description generation result is the output data of the output layer of the target image description generation model, and the indexes of the preset text retrieval database are stored in vector form, then perform vectorization processing on the description generation result, so as to obtain a target index vector by means of vector matching in the preset retrieval database.
[0092] Vectorization refers to the process of converting data into a numerical vector. The specific processing method of vectorization processing can be preset. For example, the vectorization processing of data can be completed by means of embedding (such as Embedding networks like Text2Vec). Embedding Processing usually refers to converting data from its original form into a low-dimensional and continuous vector representation, which can capture the internal features and structures of the data.
[0093] Step S202: Determine the stored data associated with the target index vector in the preset text retrieval database as the retrieved target data.
[0094] The target data can specifically be the original text fragment stored in the preset text retrieval database corresponding to the target index vector.
[0095] In a possible implementation, the description generation result is the input data of the output layer of the target image description generation model. The step of retrieving target data matching the description generation result in the preset text retrieval database includes:
[0096] Step S203: Retrieve a target index vector matching the description generation result in the preset text retrieval database;
[0097] It should be noted that the generated result of the target image description is the input data of the output layer of the target image description generation model, that is, the output of the second-to-last layer network of the target image description generation model, such as the output of the LM layer (Language Modeling Layer) of the BLIP model. At this time, it can be understood that the output of the second-to-last layer network of the model is essentially data in vector mode. Therefore, when the index of the preset text retrieval database is stored in vector form, there is no need to perform vectorization processing on the description generation result, and the target index vector can be directly obtained by vector matching based on the description generation result in the preset retrieval database.
[0098] Step S204, determine that the stored data associated with the target index vector in the preset text retrieval database is the retrieved target data.
[0099] Similarly, the target data can specifically be the original text fragment stored in the preset text retrieval database corresponding to the target index vector.
[0100] In addition, an embodiment of the present application also proposes an image description generation model training device. Referring to Figure 5 as shown, the image description generation model training device includes:
[0101] An acquisition module 10, configured to acquire an image description generation model to be trained, an original vocabulary, and a new vocabulary library, where the new vocabulary library includes at least one new Chinese vocabulary, and the original vocabulary includes at least one original vocabulary;
[0102] An insertion module 20, configured to, for each new Chinese vocabulary in the new vocabulary library, find a target original vocabulary in the original vocabulary whose semantic similarity to the new Chinese vocabulary is greater than a preset threshold, and insert the new Chinese vocabulary at a position adjacent to the target original vocabulary;
[0103] The insertion module 20 is further configured to obtain an extended vocabulary until all new Chinese vocabularies in the new vocabulary library are inserted into the original vocabulary;
[0104] A training module 30, configured to train the image description generation model based on the extended vocabulary to obtain the trained image description generation model.
[0105] In addition, an embodiment of the present application also proposes an electronic device, where the electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image description generation model training method and / or the image description generation method as described above.
[0106] Refer to Figure 6, which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may further include, but is not limited to, mobile terminals such as headphones, AR glasses, VR glasses, AR helmets, VR helmets, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0107] As Figure 6 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.
[0108] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the model through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0109] The electronic device provided in the embodiments of the present application adopts the image description generation model training method and / or the image description generation method in the above embodiments, and can solve the technical problem of how to improve the performance of the image description generation model in the Chinese scenario. Compared with the prior art, the beneficial effects of the electronic device provided in the present application are the same as those of the image description generation model training method and / or the image description generation method provided in the above embodiments, and other technical features in the electronic device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.
[0110] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0111] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0112] In addition, to achieve the above object, the embodiments of the present application further provide a readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image description generation model training method and / or the image description generation method in the above embodiments.
[0113] The computer-readable storage medium provided by the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0114] The above computer-readable storage medium may be included in an electronic device; or may exist separately without being assembled into the electronic device.
[0115] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device is caused to: obtain an image description generation model to be trained, an original vocabulary, and a new vocabulary library, where the new vocabulary library includes at least one new Chinese vocabulary, and the original vocabulary includes at least one original vocabulary; for each new Chinese vocabulary in the new vocabulary library, find a target original vocabulary in the original vocabulary whose semantic similarity with the new Chinese vocabulary is greater than a preset threshold, and insert the new Chinese vocabulary at a position adjacent to the target original vocabulary; until all the new Chinese vocabularies in the new vocabulary library are inserted into the original vocabulary to obtain an extended vocabulary; train the image description generation model based on the extended vocabulary to obtain the trained image description generation model.
[0116] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0118] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0119] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned image description generation model training method and / or image description generation method, and can solve the technical problem of how to improve the performance of the image description generation model in the Chinese scenario. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the image description generation model training method and / or image description generation method provided in the above embodiments, and will not be elaborated here.
[0120] In addition, an embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-described method for training an image description generation model and / or the method for generating an image description are implemented.
[0121] The specific implementation manner of the computer program product of the present application is basically the same as that of each of the above embodiments of the method for training an image description generation model and / or the method for generating an image description, and will not be elaborated herein.
[0122] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system including a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0123] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server or a model device, etc.) to execute the methods described in each embodiment of the present application.
[0125] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for training an image description generation model, characterized in that: The image description generation model training method comprises the following steps: Obtaining an image description generation model to be trained, an original vocabulary table and a newly added vocabulary library, wherein the newly added vocabulary library includes at least one newly added Chinese vocabulary, and the original vocabulary table includes at least one original vocabulary; For each newly added Chinese word in the newly added vocabulary library, searching the original vocabulary in the original vocabulary for a target word whose semantic similarity with the newly added Chinese word is greater than a preset threshold, and inserting the newly added Chinese word into a position adjacent to the target original word; until all the newly added Chinese words in the newly added vocabulary library are inserted into the original vocabulary list to obtain an expanded vocabulary list; The image description generation model is trained based on the expanded vocabulary to obtain the trained image description generation model.
2. The image description generation model training method according to claim 1, characterized in that: Before the step of training the image description generation model based on the expanded vocabulary, the method further includes: Initializing the embedding vector of each original word in the expanded vocabulary; For each newly added Chinese word in the expanded vocabulary, an embedding vector of a target original word adjacent to the newly added Chinese word is determined as a target embedding vector, and the embedding vector of the newly added Chinese word is initialized with the target embedding vector.
3. The image description generation model training method according to claim 1, characterized in that: The step of searching for a target original vocabulary in the original vocabulary table whose semantic similarity with the newly added Chinese vocabulary is greater than a preset threshold comprises: Inputting the newly added Chinese vocabulary into a pre-trained translation network to obtain a translated vocabulary; The semantic similarity between each original word in the original vocabulary and the translated word is compared, and a target original word whose semantic similarity is greater than a preset threshold is found from each original word.
4. A method for generating an image description, characterized in that: The image description generating method comprises the following steps: Acquire an original image to be described and a target image description generation model, and input the original image into the target image description generation model to obtain a description generation result; The target image description generation model is an image description generation model trained by the image description generation model training method according to any one of claims 1 to 3.
5. The image description generating method according to claim 4, characterized in that: After the step of inputting the original image into the target image description generation model to obtain a description generation result, the method further includes: The target data matching the description generation result is retrieved in a preset text retrieval database.
6. The image description generating method according to claim 5, characterized in that: The description generation result is output data of an output layer of the target image description generation model, and the step of searching a preset text retrieval database for target data matching the description generation result comprises: Vectorizing the description generation result to obtain a description vector, and searching a preset text retrieval database for a target index vector that matches the description vector; The stored data associated with the target index vector in the preset text retrieval database is determined as the retrieved target data.
7. The image description generating method according to claim 5, characterized in that: The description generation result is input data of the output layer of the target image description generation model, and the step of searching the preset text retrieval database for target data matching the description generation result comprises: Retrieving a target index vector matching the description generation result in a preset text retrieval database; The stored data associated with the target index vector in the preset text retrieval database is determined as the retrieved target data.
8. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that: The readable storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, which implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.