An Image Caption Generation Method Based on External Triples and Abstract Relationships
By extracting triples in image description and building an external relationship library to generate abstract relationships, the problem of the generation description of the existing image description generation model is solved, and a richer and more accurate image description generation effect is achieved.
Patent Information
- Application Number
- CN202111638065.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The description generated by the existing image description generation model is too simple to effectively utilize context knowledge and high-level abstract symbols, resulting in a worse generation effect than object detection of images.
By extracting triples from image description, building an external relationship library, and querying similar relationships based on image target categories, generating abstract relationships, and improving the accuracy of model prediction.
The generated image description is richer and more accurate, and can effectively utilize context knowledge and high-level abstract symbols, improving the effect of image description generation.
Smart Images

Figure CN114332519B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image description generation method, specifically an image description generation method based on external triples and abstract relationships, belonging to the field of image description generation. Background Art
[0002] Image description generation is a comprehensive task that combines computer vision and natural language processing, and is extremely challenging. Inspired by the encoder-decoder, attention mechanism, and reinforcement learning-based training objectives in the field of natural language processing, modern image description generation models have made amazing progress, and researchers' attention to the field of image description generation has also been increasing. In some evaluation metrics, they even outperform humans.
[0003] Although the technology of image description generation methods is constantly evolving, there is a problem that has never been solved but cannot be ignored, that is, existing models only provide simple descriptions of the prominent objects in the image, and the generated results are even inferior to a series of object detections of the image. In the process of context reasoning, people will use the knowledge learned before to help us better complete the reasoning. In addition, some studies have shown that vision-based language generation is not end-to-end, but related to high-level abstract symbols. If the visual scene is abstracted into symbols, the generation process will become clear. Inspired by this, this paper extracts triples from image descriptions, constructs an external relationship library, queries similar relationships according to the target category of the image, and provides prior knowledge for the model. At the same time, the triples are abstractly clustered to generate abstract relationships, improving the accuracy of model prediction. Summary of the Invention
[0004] The purpose of the present invention is to provide an image description generation method based on external triples and abstract relationships for the deficiencies of the prior art, so as to solve the problem that the descriptions generated by traditional image description generation methods are too simple, and improve the prediction accuracy on the original basis.
[0005] The beneficial effects of the present invention are as follows:
[0006] The present invention extracts triples from image descriptions, constructs an external relationship library, and integrates the similar relationships related to the image into the model, making the expressions of the descriptions generated by the model richer.
[0007] The present invention clusters the triples according to text similarity, generates abstract relationships and integrates them into the model, making the descriptions generated by the model more accurate. Brief Description of the Drawings
[0008] Figure 1 is the overall implementation flowchart of the present invention
[0009] Figure 2 is the schematic diagram of constructing external triples and abstract relationships of the present invention
[0010] Figure 3 It is a schematic diagram for generating a scene graph of the present invention
[0011] Figure 4 It is a schematic diagram for generating an image description of the present invention
[0012] Figure 5 It is a schematic diagram of the overall structure of the present invention Detailed implementation manners
[0013] The present invention will be further described below with reference to the accompanying drawings
[0014] Refer to Figure 1 and 5 As shown, it is a flowchart of the overall implementation scheme of the present invention
[0015] To solve these problems, the present invention constructs an external relationship library, queries similar relationships and abstract relationships from the library according to the image target category, and fuses them with the scene graph features. Specifically, first, an open-domain knowledge extraction tool is used to extract triples from the image description text, construct an external relationship library, and perform feature encoding on the triples. According to the text similarity of the relationships in the triples, the triples with high similarity are clustered into one category, which is called an abstract relationship. At the same time, the model performs object detection on the image to obtain object visual features and semantic labels. The model queries the triples in the external relationship library whose subject or object is similar to the semantic label according to the text similarity. Then, the model uses the object visual features to predict the objects, attributes, and relationships of the image respectively to generate a scene graph, and uses a multi-modal graph convolutional neural network to fuse visual features and text features to perform feature encoding on the objects, attributes, and relationships. Finally, the encoded features of the scene graph objects, attributes, and relationships are fused with the encoded features of the similar relationships and abstract relationships and input into a two-layer LSTM sequence generation model to obtain the final image description
[0016] Refer to Figure 1 and 5 As shown, an image description generation method based on external triples and abstract relationships includes the following steps
[0017] An image description generation method based on external triples and abstract relationships includes the following steps
[0018] Step (1) Use an open-domain knowledge extraction tool to extract triples from the image description text, construct an external relationship library, and perform feature encoding on the triples
[0019] Step (2) According to the text similarity of the relationship rel in the triples, cluster the triples with text similarity higher than the set threshold into one category, which is called an abstract relationship R abs ;
[0020] In step (3), object detection is performed on the image to obtain a set of target visual features V and a set of target categories W; according to the text similarity, triples in the external relation library where the subject or object (i.e., the target obj) is similar to the target category are queried, which is called the similarity relation R sim ;
[0021] In step (4), the target visual features V are used to predict the target obj, attribute attr, and relation rel of the image respectively to generate a scene graph; and the multi-modal graph convolutional neural network MGCN is used to fuse the target visual features and the word vectors of the target category W to perform feature encoding on the target obj, attribute attr, and relation rel;
[0022] In step (5), the image description generation model is used to fuse the encoded features of the scene graph and the encoded features of the relations to obtain the fused features; the encoded features of the relations include the encoded features of the similarity relations and the encoded features of the abstract relations; the fused features are input into the double-layer LSTM decoder of the image description generation model for training, and the optimal training model is selected; the image is input into the trained image description generation model to output the corresponding image description.
[0023] Further, as Figure 2 shown, the specific implementation process of step (1) is as follows:
[0024] 1-1 Use the image text descriptions in the MSCOCO and Visual Genome datasets, and use the open-domain knowledge extraction tool OpenIE to extract the triples R = {subject, predicate, object} in the image text descriptions to construct an external relation library;
[0025] 1-2 Use the pre-trained language model BERT to encode the image text descriptions to obtain the feature encodings of each word in all image text descriptions; assume that the image text description consists of K words, then the feature vector of this segment of image text description is {e0, e1, e2,..., e k ,..., e K}, where e k represents the feature encoding of the kth word, which is a 768-dimensional feature vector;
[0026] 1-3 Since the extracted triples are words that have appeared in the image text descriptions, assume that the positions of the three words in the image text description are i, j, k, then the encoded feature d of the triple is the average value of the feature encodings at the corresponding positions of the triple in the description, as shown in formula (1);
[0027]
[0028] Further, the specific implementation process in step (2) is as follows:
[0029] 2-1 Calculate the text similarity, using cosine similarity as the calculation function. Assume that the encoded features of two triples are d i′ , d j′ , then the similarity of the two triples is shown in formula (2);
[0030]
[0031] where i′, j′ represent the i′-th and j′-th triples, and the value ranges from 1 to N t , N t represents the number of triples;
[0032] 2-2 Use an unsupervised text clustering algorithm to cluster the triples with text similarity greater than a set threshold into one category, which is called the abstract relationship R abs ;
[0033] 2-3 Perform feature representation on the abstract relationship R abs . Assume that there are K1 triples in the abstract relationship R abs , then the abstract relationship, that is, the triple set Then the feature encoding of this type of abstract relationship R abs is shown in formula (3);
[0034]
[0035] where d′ k′ represents the encoded feature corresponding to the triple r′ k′ .
[0036] Further, the specific implementation process in step (3) is as follows:
[0037] 3-1 Use Faster RCNN pre-trained on the Visual Genome dataset to perform object detection on the image. Faster RCNN can obtain the object category W and the corresponding region and features of the object in the image; for the image I, take the final output of Faster RCNN and obtain the object category set W = {w1, w2,..., w s}, w s ∈R d and the object visual feature set V = {v1, v2,..., v s}, v s ∈R d , as shown in formula (4);
[0038] W, V = Faster RCNN(I) #(4)
[0039] 3-2 According to the target category set W, calculate the text similarity according to formula (2), and query the triples similar to the target category in the external relationship library, which is called the similar relationship R sim ;
[0040] 3-3 Similar to the abstract relationship, for the similar relationship R sim perform feature representation. Assume that there are K2 triples in the similar relationship, then the similar relationship is the triple set Then the feature encoding of this type of similar relationship R sim is shown in formula (5);
[0041]
[0042] Among them, d″ k″ represents the encoding feature corresponding to the triple d″ k″ ;
[0043] Furthermore, as Figure 3 shown, the specific implementation process described in step (4) is as follows:
[0044] 4-1 Use the target visual feature V to predict the target obj, attribute attr, and relationship rel of the image respectively to generate a scene graph; for the target, use Faster RCNN for target detection; for the attribute, use a pre-trained attribute classifier for attribute prediction; for the relationship, use the MOTIFS scene graph generation model for relationship detection; finally, obtain the category word vectors e o , e a , e r and their corresponding visual features v o , v a , v r ;
[0045] 4-2 In order to obtain better node features, fuse the corresponding category word vectors and visual features, and obtain the new fused node features u o , u a , u r , where W1 and W2 are fusion parameters;
[0046] u = ReLU(W1e + W2v) - (W1e - W2v) 2 #(6)
[0047] 4-3 The fused node features u o , u a , u rIt is input into the multi-modal graph convolutional neural network MGCN for encoding to obtain the encoded features of the scene graph as shown in Formulas (7) to (9);
[0048]
[0049]
[0050]
[0051] where f r , f a , f o are networks with independent parameters, and this network is composed of a fully connected layer and a ReLU layer; o x is the x-th target node, r x,y is the relationship node between the x-th target and the y-th target, o y is the target node of the y-th target; a x,l is the l-th attribute node of the x-th target node; sbj(o x ) is the set of subject nodes connected to the x-th target node, and o p is the subject target among them; obj(o x ) is the set of object nodes of the x-th target node, and o q is the object target among them; Na x , Nr x are respectively the number of attribute nodes and the number of relationship nodes of the x-th target; u is the fused node feature.
[0052] Furthermore, as Figure 4 shown, the specific implementation process of step (5) is as follows:
[0053] 5-1 Incorporate the inductive bias into the image description generation model, and the model fuses the encoded features of the scene graph and the encoded features of the relationships to obtain the final fused feature V^, as shown in Formula (10);
[0054] V^ = Dα = D·softmax(D T V`)#(10)
[0055] where D is the concatenation of the similarity relationship encoded feature D sim and the abstract relationship encoded feature D abs , and V′ is the concatenation of the encoded features of the scene graph ;
[0056] 5-2 Conduct end-to-end training on the MSCOCO dataset with the number of epochs set to 20, the learning rate to 0.00001, the batch size to 16, and use the Adam optimizer to gradually adjust the learning rate; use beam search during inference with a beam size of 5; train the model using the standard cross-entropy loss as shown in formula (11);
[0057]
[0058] where T is the length of the input sequence, y t is the word generated after inputting the t-th feature, y 1:t is the first to t-th words of the true description, and θ is the model parameter;
[0059] 5-3 Input the test image into the model to obtain the image description.
[0060] Compare the image description generation method based on the present patent invention with existing benchmark models and image description generation models based on prior knowledge. The comparison results are shown in Table (1):
[0061] model B@1 B@4 M R C S Up-Down 79.8 36.3 27.7 59.6 120.1 21.4 SGAE 81.0 39.0 28.4 58.9 129.1 22.2 this patent 81.5 39.7 28.9 60.1 130.2 24.1
[0062] Among them, Up-Down is an existing benchmark model, and SGAE is an image description generation model based on prior knowledge; B@N represents BLEU@N (N = 1, 4), M represents METOR, R represents ROUGE-L, C represents CIDEr-D, and S represents SPICE, all of which are evaluation metrics for image description models. The higher the evaluation metric, the more accurate the generated description. It can be seen from the table that the present patent has a relatively high improvement in the above evaluation metrics compared with other models, indicating that the image description generation method based on external triples and abstract relationships is effective in improving image description generation.
Claims
1. An image description generation method based on external triples and abstract relationships, characterized in that It includes the following steps: Step (1): Use an open-domain knowledge extraction tool to extract triples from the image description text, construct an external relationship library, and perform feature encoding on the triples; Step (2) clusters the triples with a text similarity of the relationship rel in the triples higher than a set threshold into one category, which is called the abstract relationship R abs ; Step (3): Perform object detection on the image to obtain a set of target visual features V and a set of target categories W; Query the triples similar to the target obj and the target category in the external relationship library according to text similarity, which is called the similarity relationship R sim ; Step (4): Use the target visual features V to predict the target obj, attribute attr, and relationship rel of the image respectively to generate a scene graph; and use a multi-modal graph convolutional neural network MGCN to fuse the target visual features with the word vectors of the target category W to perform feature encoding on the target obj, attribute attr, and relationship rel; Step (5): An image description generation model is used to fuse the scene graph encoding features and the relationship encoding features to obtain fused features; The relationship encoding features include the encoding features of similarity relationships and the encoding features of abstract relationships; The fused features are input into a two-layer LSTM decoder of the image description generation model for training to select the optimal training model; the image is input into the trained image description generation model to output the corresponding image description; The specific implementation process of Step (4) is as follows: Using the target visual feature V, the target obj, attribute attr, and relationship rel of the image are predicted respectively to generate a scene graph. For the target, Faster RCNN is used for target detection. For the attribute, a pre-trained attribute classifier is used for attribute prediction. For the relationship, the MOTIFS scene graph generation model is used for relationship detection. Finally, the category word vectors e o , e a , e r and their corresponding visual features v o , v a , v r ; 4-2 To obtain better node features, the corresponding category word vectors are fused with visual features, and the new fused node feature u is obtained through formula (6). o , u a , u r , where W1 and W2 are fusion parameters; u = ReLU(W1e + W2v) - (W1e - W2v) 2 (6) The fused fusion node features u o , u a , u r are input into the multi-modal graph convolutional neural network MGCN for encoding to obtain the scene graph encoding features as shown in Formulas (7) to (9); Among them, f r , f a , f o is a network with independent parameters, which is composed of a fully connected layer and a ReLU layer; o x is the x-th target node, r x,y is the relationship node between the x-th target and the y-th target, o y is the target node of the y-th target; a x,l is the l-th attribute node of the x-th target node; sbj(o x ) is the set of subject nodes connected to the x-th target node, o p is the subject target among them; obj(o x ) is the set of object nodes of the x-th target node, o q is the object target among them; Na x , Nr x are respectively the number of attribute nodes and the number of relationship nodes of the x-th target; u is the fused node feature.
2. The image description generation method based on external triples and abstract relationships according to claim 1, wherein As described in Step (1), the specific implementation process is as follows: 1-1: Use the image text descriptions in the MSCOCO and Visual Genome datasets, and use the open-domain knowledge extraction tool OpenIE to extract the triples R = {subject, predicate, object} in the image text description to construct an external relationship library; 1-2 Use the pre-trained language model BERT to encode the image text description to obtain the feature encoding of each word in all image text descriptions; assuming that the image text description consists of K words, the feature vector of this segment of image text description is {e0, e1, e2, …, e k , …, e K}, where e k represents the feature encoding of the k-th word, which is a 768-dimensional feature vector; 1-3: Since the extracted triples are words that have appeared in the image text description, assuming the positions of the three words in the image text description are i, j, k, the encoding feature d of the triple is the average of the feature encodings at the corresponding positions of the description of the triple, as shown in formula (1); 3. The method for generating an image description based on external triples and abstract relationships according to claim 2, wherein As described in Step (2), the specific implementation process is as follows: 2-1 Calculate the text similarity, using cosine similarity as the calculation function. Assume that the encoded features of two triples are d i' , d j' , then the similarity of the two triples is shown in formula (2); where i' and j' represent the i'-th and j'-th triples, and the value ranges from 1 to N t , N t denotes the number of triples; Using an unsupervised text clustering algorithm, group triples with text similarity greater than a set threshold into one class, which is called the abstract relationship R abs ; 2 - 3 pairs of abstract relationships R abs Perform feature representation. Assume the abstract relationship R abs There are K1 triples, then the abstract relationship is the triple set Then this type of abstract relationship R abs The feature encoding is as shown in formula (3); Among them, d' k' represents the encoded feature corresponding to the triple r' k' 4. The method for generating an image description based on external triples and abstract relationships according to claim 3, wherein As described in Step (3), the specific implementation process is as follows: 3-1 Use Faster R-CNN pre-trained on the Visual Genome dataset to perform object detection on images. Faster R-CNN can obtain the object category W, as well as the region and features of the corresponding object in the image. For image I, take the final output of Faster R-CNN and obtain the set of object categories W = {w1, w2, …, w s}, w s ∈R d and the set of object visual features V = {v1, v2, …, v s}, v s ∈R d , as shown in formula (4); W, V = Faster RCNN(I) (4) 3-2 Calculate the text similarity according to the target category set W, and query the triples similar to the target category in the external relationship library, which is called the similar relationship R sim ; 3-3 Similar to the abstract relationship, for the similarity relationship R sim perform feature representation. Assume that there are K2 triples in the similarity relationship, then the similarity relationship is the triple set Then this type of similarity relationship R sim The feature encoding is shown in formula (5); Among them, d” k” represents the triple d” k” corresponding coding feature.
5. The method for generating an image description based on external triples and abstract relationships according to claim 4, wherein As described in Step (5), the specific implementation process is as follows: 5-1: Incorporate the inductive bias into the image description generation model, and the model fuses the scene graph encoding features and the relationship encoding features to obtain the final fused feature V^, as shown in formula (10); V^ = Dα = D·softmax(D T V`) (10) Among them, D is the concatenation of the similarity relationship coding feature D sim and the abstract relationship coding feature D abs , and V` is the concatenation of the scene graph coding feature ; 5-2: Perform end-to-end training on the MSCOCO dataset, set the epoch to 20, the learning rate to 0.00001, the batchsize to 16, and use the Adam optimizer to gradually adjust the learning rate; use beam search during the inference process, with the beam size being 5; use the standard cross-entropy loss to train the model, as shown in formula (11); where T is the length of the input sequence, y t is the word generated after inputting the t-th feature, y 1:t is the first to t-th words of the true description, and θ is the model parameter; 5-3: Input the test image into the model to obtain the image description.
Citation Information
Patent Citations
Image description generation method based on relation between external knowledge and targets
CN113609326A
Image description method and system based on two-way feature encoder
CN113642630A