Image text description generation method and device and storage medium
By performing object detection and feature extraction on images and generating image text descriptions, the problem of low accuracy in image text description in the prior art is solved, and accurate description of fine-grained information of image elements is achieved.
Patent Information
- Application Number
- CN202510142871.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
AI Technical Summary
When generating image text descriptions, it is difficult to accurately reflect the fine-grained information of each image element in the image, resulting in the generated text description being low.
By performing object detection on the image to be described, image elements and their detection categories are determined, image description features and element description features are generated, and input them into a generation model to generate text descriptions.
Through fine-grained image description features and element description features, accurate semantic information is provided for the generation model, and the generated text description is more accurate, including detailed descriptions of each element in the image.
Smart Images

Figure CN120107659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, device and storage medium for generating image text description. Background Art
[0002] The generation of image text descriptions is widely used in computer vision, artificial intelligence and other fields. For example, it can automatically generate image tags and summaries, and automatically generate advertising copy for artwork images.
[0003] In existing methods for determining text descriptions of images, image features obtained by feature encoding of the image are usually input into a generative model, and the encoded features are decoded by the generative model to obtain a text description. Image features tend to reflect the global semantic information of the image, and image features often cannot accurately reflect the fine-grained information of each image element contained in the image. The text description obtained based on such image features may lack the description of individual image elements in the image, resulting in low accuracy of the generated text description.
[0004] Therefore, the present specification provides a method for generating a text description of an image. Summary of the invention
[0005] The present specification provides a method, device, storage medium and electronic device for generating image text description, so as to at least partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This specification provides a method for generating an image text description, including:
[0008] Acquire an image to be described, perform target detection on the image to be described, and determine image elements contained in the image to be described and detection categories corresponding to the image elements;
[0009] Determine the image description feature of the image to be described according to the image element; determine the element description feature of the image to be described according to the detection category;
[0010] The image description features and the element description features are input into a generation model to determine a text description of the image to be described.
[0011] Optionally, determining the image description feature of the image to be described according to the image element specifically includes:
[0012] Determine an image area of the image element in the image to be described;
[0013] Feature encoding is performed on the image region to obtain image description features of the image to be described.
[0014] Optionally, the image to be described consists of multiple target images;
[0015] Determining image description features of the image to be described according to the image elements specifically includes:
[0016] Determining candidate description features of each target image according to image elements contained in each target image;
[0017] Determining the similarity between the candidate descriptive features of each target image;
[0018] Determine the nodes corresponding to each target image, determine the weights of the edges between the nodes according to the similarities, and obtain target graph data;
[0019] Determine the importance score of each node in the target graph data by using a graph ranking algorithm;
[0020] Select a specified number of nodes according to the importance scores of the nodes;
[0021] The candidate description features of the target image corresponding to the specified number of nodes are used as image description features of the image to be described.
[0022] Optionally, determining the element description feature of the image to be described according to the detection category specifically includes:
[0023] The detection categories corresponding to the image elements contained in the target image corresponding to the specified number of nodes are used as element description features of the image to be described.
[0024] Optionally, the image to be described consists of multiple target images;
[0025] Determining image description features of the image to be described according to the image elements; and determining element description features of the image to be described according to the detection category, specifically including:
[0026] For each target image, according to the image elements contained in the target image, the image description feature of the target image is determined; according to the detection category corresponding to the image elements contained in the target image, the element description feature of the target image is determined;
[0027] Inputting the image description feature and the element description feature into a generation model to determine a text description of the image to be described specifically includes:
[0028] Inputting the image description features of the target image and the element description features of the target image into a generation model to determine a text description of the target image;
[0029] The text descriptions of each target image are semantically integrated to determine the text description of the image to be described.
[0030] Optionally, the image description feature and the element description feature are input into a generation model to determine a text description of the image to be described, specifically including:
[0031] The image description feature and the element description feature are integrated to determine a comprehensive description feature;
[0032] At the first time step, a preset start mark is input into the generation model to obtain a hidden layer vector of the first time step; the hidden layer vector and the comprehensive description feature are input into the generation model to obtain a generated word of the first time step;
[0033] For each time step after the first time step, determine the word sequence composed of the generated words of each time step before the time step; input the word sequence into the generation model to obtain the hidden layer vector of the time step; input the hidden layer vector of the time step and the comprehensive description feature into the generation model to obtain the generated word of the time step; when the generated word of the time step is a preset end mark, stop generating, use the word sequence composed of the generated words of each time step as the text description of the image to be described, and when the generated word of the time step is not a preset end mark, continue the generation process of the next time step.
[0034] Optionally, the image description feature and the element description feature are input into a generation model to determine a text description of the image to be described, specifically including:
[0035] The image description feature and the element description feature are integrated to determine a comprehensive description feature; and the number of generated paths is determined;
[0036] At the first time step, a preset start mark is input into the hidden layer of the generative model to obtain the hidden layer vector of the first time step; the hidden layer vector and the comprehensive description feature are input into the regression layer of the generative model to determine the matching probability of each candidate word in the vocabulary and the image to be described at the first time step;
[0037] Determine the candidate words of the number of generation paths according to the order of the matching probabilities at the first time step, and use them as the generation words of each generation path at the first time step respectively;
[0038] Based on the generated words of each generation path at the first time step, for each time step after the first time step, determine the word sequence composed of the generated words of each time step before the time step on the generation path; input the word sequence into the hidden layer of the generation model to obtain the hidden layer vector of the time step; input the hidden layer vector of the time step and the comprehensive description feature into the regression layer of the generation model to predict the matching probability of each candidate word in the vocabulary and the image to be described at the time step;
[0039] The candidate word corresponding to the maximum matching probability among the matching probabilities is used as the generated word on the generation path at this time step;
[0040] When the generated word of the time step on the generated path is a preset end mark, stop generating, and use the word sequence composed of the generated words of each time step of the generated path as the text description corresponding to the generated path; when the generated word of the time step is not a preset end mark, continue the generation process of the next time step;
[0041] Determine the joint probability of the generated path according to the matching probability determined at each time step in the text description corresponding to the generated path;
[0042] According to the joint probability of each generation path, a text description of the image to be described is determined in the text descriptions corresponding to each generation path.
[0043] This specification provides a device for generating an image text description, the device comprising:
[0044] A detection module, which obtains an image to be described, performs target detection on the image to be described, and determines image elements contained in the image to be described and detection categories corresponding to the image elements;
[0045] A feature determination module determines the image description features of the image to be described according to the image elements; and determines the element description features of the image to be described according to the detection category;
[0046] The generation module inputs the image description features and the element description features into a generation model to determine a text description of the image to be described.
[0047] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method for generating the above-mentioned image text description is implemented.
[0048] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for generating image text description when executing the program.
[0049] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0050] In the method for generating text description of an image provided in the present specification, an image to be described is obtained, target detection is performed on the image to be described, and the image elements contained in the image to be described and the detection categories corresponding to the image elements are determined. Based on the image elements, the image description features of the image to be described are determined, and based on the detection categories, the element description features of the image to be described are determined. The image description features and the element description features are input into a generation model to determine the text description of the image to be described. In this method, the image elements of the image to be described are used to represent the detailed information of the image to be described. By determining the image description features of the visual modality and the element description features of the text modality, the generation model is provided with accurate semantics of the image details, thereby generating an accurate text description. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0052] Figure 1 A flowchart of a method for generating an image text description in this specification;
[0053] Figure 2 A schematic diagram of a device for generating text description of an image provided in this specification;
[0054] Figure 3 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0056] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0057] Figure 1 The following is a flow chart of a method for generating a text description of an image in this specification, which specifically includes the following steps:
[0058] S100: Acquire an image to be described, perform target detection on the image to be described, and determine image elements contained in the image to be described and detection categories corresponding to the image elements.
[0059] All steps in the method for generating text description of an image provided in this specification can be implemented by any electronic device with computing functions, such as a terminal, a server, etc. For the sake of ease of description, the method for generating text description of an image provided in this specification is described below with the server as the execution subject.
[0060] This specification does not limit the number of images to be described. When there is only one image to be described, the server determines the text description of the one image to be described. When there are multiple images to be described, the server can determine the text descriptions of the multiple images to be described respectively, or generate a summary text description of the multiple images to be described.
[0061] The following first takes the number of images to be described as one as an example to illustrate the method for generating image text description in this specification, and an embodiment of multiple images to be described will be provided later.
[0062] In order to determine the fine-grained information of the image to be described, the server performs object detection on the image to be described, determines the image elements contained in the image to be described, and the detection categories corresponding to the image elements. Here, the number of image elements determined and the specific detection categories of the image elements are related to the content of the image to be described.
[0063] The target detection in this step can be implemented by a detection model for multi-target detection, or by several detection models for single-target detection. This specification does not limit the specific structure of the detection model.
[0064] The server may obtain a detection result of the detection model, where the detection result includes a detection frame of an image element detected in the image to be described, and a detection category corresponding to the detection frame.
[0065] For example, if the image to be described is a landscape photo, the server determines three image objects in the image to be described, and the detection categories of the three image objects are mountains, rivers, and trees, respectively.
[0066] Each image element contained in the image to be described reflects the detailed information of the image to be described, and can provide fine-grained semantic information for the generation model, thereby generating a more accurate image description.
[0067] S102: Determine image description features of the image to be described according to the image elements; and determine element description features of the image to be described according to the detection category.
[0068] The image element is a target detected by the server in the image to be described, which is image data. The detection category of the image element is used to indicate the category of the detected image element, which is text data.
[0069] The server performs feature encoding on the image to be described according to the image elements to obtain image description features of the image modality of the image to be described. This specification does not limit the feature encoding method here, which can be performed through a convolutional neural network or determined based on histogram features, texture features, etc. of the image elements.
[0070] Specifically, the server determines an image region of the image element in the image to be described, performs feature encoding on the image region, and obtains image description features of the image to be described.
[0071] When there are multiple image elements, the server determines the image area of each image element in the image to be described. For each image element, feature encoding is performed on the image area corresponding to the image element to obtain the element feature of the image element. Then, the image description feature of the image to be described is composed of the element features of each image element.
[0072] The server performs feature encoding based on the detected categories of the image elements to obtain the element description features of the text modality of the image to be described. The feature encoding operation here can be implemented by any text embedding model.
[0073] The present specification does not limit the order of determining the image description features and the element description features.
[0074] The object detection operation in this specification is determined by the detection model, and the detection model recognizes the image elements of the predetermined multiple categories. Therefore, the detection category corresponding to the image element of the image to be described determined by the server is selected from the predetermined multiple categories. For an image element, if the image element is recognized from the image to be described in step S100, it means that the training data of the detection model contains sample images with labels corresponding to the detection category of the image element.
[0075] Therefore, the description of the detection category determined in this specification is relatively fixed, and the element description features determined based on the detection category are also relatively fixed. The element description features reflect what image elements are contained in the image to be described. The image description features provide rich visual semantic information for the generation model. Compared with the element description features, they contain visual information such as shape, state, color, etc., and have richer meanings.
[0076] S104: Input the image description features and the element description features into a generation model to determine a text description of the image to be described.
[0077] The image description features and element description features accurately reflect the semantics of the image to be described in the vector dimension. The server inputs the image description features and element description features into the generation model, and decodes the image description features and element description features through the generation model to obtain the text description of the image to be described.
[0078] This specification does not limit the specific network structure of the generation model, for example, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Generative Pre-Trained Transformer (GPT), etc.
[0079] The generation model in this specification can be obtained by training using each sample image and a text description of each sample image as training data.
[0080] The element description feature can ensure that the generated text description contains descriptions of all image elements. The image description feature is used to enable the generation model to determine more detailed semantic information of each image element, thereby obtaining an accurate text description.
[0081] In the method for generating text description of an image provided in the present specification, an image to be described is obtained, target detection is performed on the image to be described, and the image elements contained in the image to be described and the detection categories corresponding to the image elements are determined. Based on the image elements, the image description features of the image to be described are determined, and based on the detection categories, the element description features of the image to be described are determined. The image description features and the element description features are input into a generation model to determine the text description of the image to be described. In this method, the image elements of the image to be described are used to represent the detailed information of the image to be described. By determining the image description features of the visual modality and the element description features of the text modality, the generation model is provided with accurate semantics of the image details, thereby generating an accurate text description.
[0082] In the above step S102, the server may also determine the image description features according to the image to be described including the detection frame.
[0083] Specifically, the server determines the color of the detection frame contained in the image to be described according to the preset correspondence between each detection category and the color of the detection frame. Then, the image to be described marked with detection frames of different colors is feature encoded to obtain image description features.
[0084] In this embodiment, the detection frame colors corresponding to image elements of different detection categories are different, and the color, size, position and other data of the detection frame contained in the image to be described reflect the detailed information in the image to be described. Therefore, the server can directly perform feature encoding on the entire image to be described.
[0085] The above is a method for determining text description when the number of images to be described is one, and the following is a method for determining text description when the number of images to be described is multiple. In this embodiment, the image to be described is composed of multiple target images.
[0086] When there are multiple images to be described, the server may determine the text description of each target image included in the images to be described through the method of steps S100 to S104 as the text description of the images to be described.
[0087] In this embodiment, because the text descriptions of each target image are generated independently, there is no contextual association between the text descriptions of each target image, resulting in a text description of the image to be described that is not smooth in context and a poor reading experience for users.
[0088] Therefore, in one or more embodiments of the present specification, when the image to be described is composed of multiple target images, in order to determine a text description of the image to be described with smooth context, the server semantically integrates the text descriptions of each target image to obtain a text description of the image to be described.
[0089] Furthermore, considering that the number of target images contained in the image to be described is relatively large, the length of the text description of the image to be described determined by the above method will also be relatively long. Therefore, when the image to be described is composed of multiple target images, this specification provides another method for determining the text description of the image to be described.
[0090] The server applies the method described in step S100 to respectively determine each image element contained in the image to be described and the detection category corresponding to each image element.
[0091] In the above step S102, the server determines candidate description features of each target image according to the image elements contained in each target image, and determines the image description features of the image to be described by screening each candidate description feature.
[0092] In one embodiment, the server may also adopt the following screening method:
[0093] The server may cluster the candidate description features of each target image according to the similarity of the candidate description features of each target image, and use the candidate description features corresponding to the cluster centers obtained by clustering as the image description features of the image to be described.
[0094] In another embodiment, the server may also adopt the following screening method:
[0095] First, the server combines each target image in pairs to determine each image pair, and calculates the similarity between the candidate description features of the target images contained in each image pair.
[0096] Secondly, the server constructs nodes corresponding to each target image. For each image pair, the weights of the edges between the nodes corresponding to the target images contained in the target image pair are determined based on the similarities between the candidate descriptive features of the target images contained in the image pair. The target graph data is constructed based on the determined nodes and the weights of the edges between the nodes.
[0097] Then, the importance score of each node is determined through the graph ranking algorithm. The higher the importance score of a node, the more representative the target image corresponding to the node is in the image to be described.
[0098] This specification does not restrict the specific application of graph ranking algorithms. Graph ranking algorithms such as PageRank (PP), Hyperlink-Induced Topic Search (HITS), and Stochastic Approach for Link-Structure Analysis (SALSA) can be selected as needed.
[0099] The server sorts the nodes according to the importance scores of the nodes, and selects a specified number of nodes from the nodes according to the sorting results. The specific value of the specified number can be set according to the requirements.
[0100] Finally, the server selects the candidate description features of the target image corresponding to the specified number of nodes as the image description features of the image to be described.
[0101] The image description features of the image to be described determined by this embodiment are selected according to the importance score and can represent the overall semantics of the image to be described. Therefore, even if a small number of image description features of the image to be described are screened in this embodiment, an accurate and complete text description of the image to be described can be generated.
[0102] In the above step S102, the server determines the element description features of the image to be described according to the detection categories corresponding to the image elements of each target image.
[0103] Finally, in the above step S104, the server inputs the image description features of the image to be described and the element description features of the image to be described determined in this embodiment into the generation model to obtain a text description of the image to be described.
[0104] Because the image description features of the image to be described determined in this embodiment are not candidate description features of all target images, but the image ranking algorithm is applied to determine a specified number of representative images in each target image. Based on the candidate description features of these representative images, the image description features of the image to be described are determined.
[0105] The image description features determined in this embodiment are determined based on the candidate description features of these representative images, which reduces the data volume of the image description features while ensuring the representativeness of the semantics of the image description features. Through the image description features, a description text of each specific representative target image can be determined.
[0106] The description text of the image to be described in this embodiment is generated at one time and has contextual semantic coherence.
[0107] In another embodiment of the present specification, the server may further screen the detection categories of the image elements based on the screening and determination of the image description features, so that the server can determine a more concise description text corresponding to each target image by generating a model.
[0108] Specifically, the server uses the detection categories corresponding to the image elements contained in the target image corresponding to the specified number of nodes as element description features of the image to be described.
[0109] In the above step S104, when the generation model adopts a cyclic network structure, the server determines the text description of the image to be described according to the following steps.
[0110] The server fuses the image description features and the element description features to determine the comprehensive description features. The comprehensive description features reflect the fine-grained semantics of the image to be described. This specification does not limit the specific means of fusion, which can be neural network fusion, feature splicing, feature weighting, etc.
[0111] The server then generates the words in the description text one by one at each time step.
[0112] At the first time step, the server inputs the preset start tag into the generative model to obtain the hidden layer vector of the first time step. The hidden layer vector of the first time step and the comprehensive description feature are input into the generative model to obtain the generated word of the first time step.
[0113] For each time step after the first time step, determine the word sequence composed of the generated words of each time step before the time step. Input the word sequence into the generation model to obtain the hidden layer vector of the time step. Input the hidden layer vector and the comprehensive description feature of the time step into the generation model to obtain the generated word of the time step. When the generated word of the time step is a preset end mark, stop generating, and use the word sequence composed of the generated words of each time step as the text description of the image to be described. When the generated word of the time step is not a preset end mark, continue the generation process of the next time step.
[0114] In this embodiment, the generation model may include a hidden layer and a regression layer, the hidden layer is used to output the hidden layer vector, and the regression layer is used to receive the hidden layer vector and the comprehensive description feature and output the generated word.
[0115] In one or more embodiments of the present specification, in order to further improve the accuracy of the determined description text, in the above step S104, the server may also generate multiple description texts at the same time, and determine a better text description among the multiple text descriptions.
[0116] First, the server fuses the image description features and the element description features to determine the comprehensive description features and the number of generated paths.
[0117] The number of generated paths is the number of description texts generated by the server at the same time, and the specific value can be set as needed. The larger the value of the generated path, the more description texts the server determines, and the more likely it is to determine a more accurate description text among the description texts.
[0118] At each time step, the server generates the words in the description text one by one.
[0119] At the first time step, the preset start marker is input into the hidden layer of the generative model to obtain the hidden layer vector of the first time step. The hidden layer vector of the first time step and the comprehensive description feature are input into the regression layer of the generative model to determine the matching probability of each candidate word in the vocabulary and the image to be described at the first time step.
[0120] Secondly, the server determines the candidate words of the number of generated paths according to the order of the matching probabilities in the first time step, and uses them as the generated words of each generated path in the first time step.
[0121] Next, based on the generated words of each generation path at the first time step, the server determines the word sequence composed of the generated words of each time step before the time step on the generation path for each time step after the first time step. The word sequence is input into the hidden layer of the generation model to obtain the hidden layer vector of the time step. The hidden layer vector of the time step and the comprehensive description feature are input into the regression layer of the generation model to predict the matching probability of each candidate word in the vocabulary and the image to be described at the time step.
[0122] Then, the server uses the candidate word corresponding to the maximum matching probability among the matching probabilities as the generated word on the generation path at the time step. When the generated word on the generation path at the time step is a preset end mark, the generation is stopped, and the word sequence composed of the generated words of each time step of the generation path is used as the text description corresponding to the generation path. When the generated word at the time step is not a preset end mark, the generation process of the next time step is continued.
[0123] Next, the server determines the joint probability of the generated path according to the matching probability determined at each time step in the text description corresponding to the generated path.
[0124] Specifically, the server may determine the joint probability of the generated path according to the product of the matching probabilities determined at each time step in the text description corresponding to the generated path.
[0125] Finally, the server determines the text description of the image to be described from the text descriptions corresponding to the generation paths according to the joint probability of the generation paths. Specifically, the server may use the description text corresponding to the generation path with the largest joint probability as the text description of the image to be described. Alternatively, the server may randomly select a text description from the text descriptions corresponding to the joint paths as the text description of the image to be described.
[0126] It should be noted that the vocabulary in the above two embodiments includes all words that may be used in the text description generation process. Any existing vocabulary can be selected, or common words can be counted and constructed according to the text description generation scenario.
[0127] In addition, this specification does not limit the specific forms of the preset start mark and end mark used in the above two embodiments, and they can be set to any character or character combination as needed.
[0128] The above is the method for generating image text description provided in this specification. Based on the same idea, this specification also provides a corresponding device for generating image text description, such as Figure 2 shown.
[0129] Figure 2 A schematic diagram of a device for generating an image text description provided in this specification specifically includes:
[0130] The detection module 200 is used to obtain an image to be described, perform target detection on the image to be described, and determine image elements contained in the image to be described and detection categories corresponding to the image elements;
[0131] A feature determination module 202 is used to determine the image description feature of the image to be described according to the image element; and determine the element description feature of the image to be described according to the detection category;
[0132] The generation module 204 is used to input the image description features and the element description features into a generation model to determine a text description of the image to be described.
[0133] Optionally, the feature determination module 202 is specifically configured to determine an image region of the image element in the image to be described, perform feature encoding on the image region, and obtain image description features of the image to be described.
[0134] Optionally, the image to be described is composed of multiple target images, and the feature determination module 202 is specifically used to determine the candidate description features of each target image according to the image elements contained in each target image, determine the similarities between the candidate description features of each target image, determine the nodes corresponding to each target image, determine the weights of the edges between the nodes according to the similarities, obtain target graph data, determine the importance scores of each node in the target graph data through a graph ranking algorithm, select a specified number of nodes according to the importance scores of each node, and use the candidate description features of the target images corresponding to the specified number of nodes as the image description features of the image to be described.
[0135] Optionally, the feature determination module 202 is specifically configured to use the detection categories corresponding to the image elements contained in the target image corresponding to the specified number of nodes as element description features of the image to be described.
[0136] Optionally, the image to be described is composed of multiple target images, and the feature determination module 202 is specifically used to determine the image description feature of each target image according to the image elements contained in the target image; and determine the element description feature of the target image according to the detection category corresponding to the image elements contained in the target image. The generation module 204 is specifically used to input the image description feature of the target image and the element description feature of the target image into the generation model, determine the text description of the target image, perform semantic integration on the text descriptions of each target image, and determine the text description of the image to be described.
[0137] Optionally, the generation module 204 is specifically used to fuse the image description features and the element description features to determine a comprehensive description feature, and at the first time step, input a preset start tag into the generation model to obtain a hidden layer vector of the first time step; input the hidden layer vector and the comprehensive description feature into the generation model to obtain a generated word of the first time step, and determine a word sequence composed of generated words of each time step before the first time step for each time step after the first time step; input the word sequence into the generation model to obtain a hidden layer vector of the time step, input the hidden layer vector of the time step and the comprehensive description feature into the generation model to obtain a generated word of the time step, and when the generated word of the time step is a preset end tag, stop generating, and use the word sequence composed of the generated words of each time step as the text description of the image to be described, and when the generated word of the time step is not a preset end tag, continue the generation process of the next time step.
[0138] Optionally, the generation module 204 is specifically used to fuse the image description features and the element description features, determine the comprehensive description features, determine the number of generation paths, and at the first time step, input the preset start mark into the hidden layer of the generation model to obtain the hidden layer vector of the first time step; input the hidden layer vector and the comprehensive description features into the regression layer of the generation model to determine the matching probability of each candidate word in the vocabulary and the image to be described at the first time step, determine the candidate words of the number of generation paths in the order of each matching probability at the first time step, and use them as the generated words of each generation path at the first time step respectively; based on the generated words of each generation path at the first time step, determine the word sequence composed of the generated words of each time step before the time step on the generation path for each time step after the first time step; input the word sequence into the generation model to obtain the hidden layer vector of the time step; the hidden layer vector of the time step and the comprehensive description feature are input into the regression layer of the generation model to predict the matching probability of each candidate word in the vocabulary and the image to be described at the time step, and the candidate word corresponding to the largest matching probability among the matching probabilities is used as the generated word on the generation path at the time step; when the generated word at the time step on the generation path is a preset end mark, the generation is stopped, and the word sequence composed of the generated words at each time step of the generation path is used as the text description corresponding to the generation path; when the generated word at the time step is not a preset end mark, the generation process of the next time step is continued, and the joint probability of the generation path is determined according to the matching probability determined at each time step in the text description corresponding to the generation path, and the text description of the image to be described is determined in the text descriptions corresponding to the generation paths according to the joint probabilities of each generation path.
[0139] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Provides a method for generating text descriptions of images.
[0140] This manual also provides Figure 3 The schematic structure diagram of the electronic device shown in FIG. Figure 3 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to the software implementation, this specification does not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0141] For the improvement of a technology, it can be clearly distinguished whether it is a hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0142] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0143] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0144] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0145] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0146] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0147] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0149] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0150] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0151] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0152] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0153] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0155] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0156] The above description is only an embodiment of this specification and is not intended to limit this specification. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this application.
Claims
1. A method for generating a text description of an image, characterized in that: include: Acquire an image to be described, perform target detection on the image to be described, and determine image elements contained in the image to be described and detection categories corresponding to the image elements; Determining image description features of the image to be described according to the image elements; Determining, according to the detection category, element description features of the image to be described; The image description features and the element description features are input into a generation model to determine a text description of the image to be described.
2. The method according to claim 1, characterized in that Determining image description features of the image to be described according to the image elements specifically includes: Determine an image area of the image element in the image to be described; Feature encoding is performed on the image region to obtain image description features of the image to be described.
3. The method according to claim 1, characterized in that The image to be described is composed of multiple target images; Determining image description features of the image to be described according to the image elements specifically includes: Determining candidate description features of each target image according to image elements contained in each target image; Determining the similarity between the candidate description features of each target image; Determine the nodes corresponding to each target image, determine the weights of the edges between the nodes according to the similarities, and obtain target graph data; Determine the importance score of each node in the target graph data by using a graph ranking algorithm; Select a specified number of nodes according to the importance scores of the nodes; The candidate description features of the target image corresponding to the specified number of nodes are used as image description features of the image to be described.
4. The method according to claim 3, characterized in that Determining the element description features of the image to be described according to the detection category specifically includes: The detection categories corresponding to the image elements contained in the target image corresponding to the specified number of nodes are used as element description features of the image to be described.
5. The method according to claim 1, characterized in that The image to be described is composed of multiple target images; Determining image description features of the image to be described according to the image elements; Determining the element description features of the image to be described according to the detection category specifically includes: For each target image, according to the image elements contained in the target image, the image description feature of the target image is determined; according to the detection category corresponding to the image elements contained in the target image, the element description feature of the target image is determined; Inputting the image description feature and the element description feature into a generation model to determine a text description of the image to be described specifically includes: Inputting the image description features of the target image and the element description features of the target image into a generation model to determine a text description of the target image; The text descriptions of each target image are semantically integrated to determine the text description of the image to be described.
6. The method according to claim 1, characterized in that Inputting the image description feature and the element description feature into a generation model to determine a text description of the image to be described specifically includes: The image description feature and the element description feature are integrated to determine a comprehensive description feature; At the first time step, a preset start mark is input into the generation model to obtain a hidden layer vector of the first time step; the hidden layer vector and the comprehensive description feature are input into the generation model to obtain a generated word of the first time step; For each time step after the first time step, determine the word sequence composed of the generated words of each time step before the time step; input the word sequence into the generation model to obtain the hidden layer vector of the time step; input the hidden layer vector of the time step and the comprehensive description feature into the generation model to obtain the generated word of the time step; when the generated word of the time step is a preset end mark, stop generating, use the word sequence composed of the generated words of each time step as the text description of the image to be described, and when the generated word of the time step is not a preset end mark, continue the generation process of the next time step.
7. The method according to claim 1, characterized in that Inputting the image description feature and the element description feature into a generation model to determine a text description of the image to be described specifically includes: The image description feature and the element description feature are integrated to determine a comprehensive description feature; and the number of generated paths is determined; At the first time step, a preset start mark is input into the hidden layer of the generative model to obtain the hidden layer vector of the first time step; the hidden layer vector and the comprehensive description feature are input into the regression layer of the generative model to determine the matching probability of each candidate word in the vocabulary and the image to be described at the first time step; Determine the candidate words of the number of generation paths according to the order of the matching probabilities at the first time step, and use them as the generation words of each generation path at the first time step respectively; Based on the generated words of each generation path at the first time step, for each time step after the first time step, determine the word sequence composed of the generated words of each time step before the time step on the generation path; input the word sequence into the hidden layer of the generation model to obtain the hidden layer vector of the time step; input the hidden layer vector of the time step and the comprehensive description feature into the regression layer of the generation model to predict the matching probability of each candidate word in the vocabulary and the image to be described at the time step; The candidate word corresponding to the maximum matching probability among the matching probabilities is used as the generated word on the generation path at this time step; When the generated word of the time step on the generated path is a preset end mark, stop generating, and use the word sequence composed of the generated words of each time step of the generated path as the text description corresponding to the generated path; when the generated word of the time step is not a preset end mark, continue the generation process of the next time step; Determine the joint probability of the generated path according to the matching probability determined at each time step in the text description corresponding to the generated path; According to the joint probability of each generation path, a text description of the image to be described is determined in the text descriptions corresponding to each generation path.
8. A device for generating text description of an image, characterized in that: include: A detection module, which obtains an image to be described, performs target detection on the image to be described, and determines image elements contained in the image to be described and detection categories corresponding to the image elements; A feature determination module, which determines the image description feature of the image to be described according to the image elements; Determining, according to the detection category, element description features of the image to be described; The generation module inputs the image description features and the element description features into a generation model to determine a text description of the image to be described.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.