Image Caption Generation Method Based on Adaptive Attention Mechanism and Knowledge Graph
By combining adaptive attention mechanism and knowledge graph, image description is generated using MSCOCO and Visual Genome datasets, the problem of inaccurate image description in the prior art is solved, object detection and relationship recognition are optimized, and image descriptions that are more in line with human description are generated.
Patent Information
- Application Number
- CN202211579831.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-09
AI Technical Summary
The prior art is difficult to accurately combine the environment and target behavior in the image to generate vivid descriptions in image description, and lacks the ability to process common sense in real world.
Combining the adaptive attention mechanism and knowledge graph, a knowledge graph is generated through MSCOCO and Visual Genome datasets, a vectorized representation is used using the TransR model, and an image description is generated based on the LSTM model.
It realizes the use of knowledge graph information and visual information more accurately in image description, optimizes object detection and relationship recognition, and generates image descriptions that are more in line with human description.
Smart Images

Figure CN115964508B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image description generation method, specifically an image description generation method based on an adaptive attention mechanism and a knowledge graph, belonging to the field of image description generation. Background Art
[0002] Accurate and vivid image descriptions have different manifestation forms in different environments. Therefore, we must combine the information such as the environment where the image is located and the possible behaviors of the objects in the image to generate better descriptions. The knowledge graph has natural processing advantages for objects and the environments where they are located. It models the entities, attributes in the image and their relationships, and uses the basic data structure of "graph" to achieve the purpose of accurately expressing various relationships in the real world using a general language.
[0003] On the other hand, in image description, we need to know the semantic information sequence generated at the current time step and provide a network with an attention mechanism to generate corresponding pixel weights, thereby determining the subsequent description situation. The attention mechanism can consider the semantic information sequence generated so far and focus on the part of the image that needs to be described next. This enables the adaptive attention mechanism to achieve a good effect on the repeated utilization of the information generated by the knowledge graph.
[0004] Inspired by the strong semanticity of the knowledge graph and the adaptability of the attention mechanism, the present invention proposes an image description generation method based on an adaptive attention mechanism and a knowledge graph, realizing the graph processing of common sense knowledge in the real world, and accurately judging when to rely on the language model and knowledge graph information and when to rely on visual information to optimize object detection and object relationship recognition in image description. Summary of the Invention
[0005] The purpose of the present invention is to provide an image description generation method based on an adaptive attention mechanism and a knowledge graph, so as to use the knowledge graph containing rich semantic knowledge as prior knowledge and realize the advantages in the object detection and relationship prediction links during the image description process.
[0006] To implement this solution, on the basis of the traditional image description generation method, the present invention combines the knowledge graph generated by using the COCO dataset and the Visual Genome dataset and an adaptive attention mechanism model to invent an image description generation method based on an adaptive attention mechanism and a knowledge graph.
[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0008] An image description generation method based on an adaptive attention mechanism and a knowledge graph, characterized by including the following steps:
[0009] Step (1): Obtain the MSCOCO dataset and the Visual Genome dataset, and convert the descriptions of images in the MSCOCO dataset and the Visual Genome dataset into a knowledge graph.
[0010] Step (2): Use the TransR model to vectorize and represent the knowledge graph to obtain word embedding vectors.
[0011] Step (3): Generate an adaptive attention mechanism model through the word embedding vectors and visual sentinels; based on the LSTM model, combine the adaptive attention mechanism model and the knowledge graph to obtain an image description model based on the adaptive attention mechanism and the knowledge graph.
[0012] Step (4): Input the picture to be described into the image description model based on the adaptive attention mechanism and the knowledge graph obtained in Step (3) to obtain the image description of the picture to be described.
[0013] Preferably, Step (1) specifically includes the following steps:
[0014] Step 1-1: Save the descriptions of images in the MSCOCO dataset and the Visual Genome dataset in txt format, and use the information extractor OPENIE to convert the descriptions of images into triple information Triplets. The triple information Triplets include a subject, an object, and the relationship between the subject and the object.
[0015] Step 1-2: Represent the subject and the object in the following way:
[0016] (entity: ID, name:, LABEL), where entity: ID is the entity ID of the subject or the object, used to indicate which entities are in the picture, name is the entity attribute of the subject or the object, used to describe what type of thing the entity belongs to, and LABEL is the label, used to describe the type of the entity attribute.
[0017] Represent the relationship between the subject and the object in the following way:
[0018] (:START_ID, :END_ID, :TYPE), where :START_ID is the entity: ID of the subject, used to refer to the subject in the relationship between the subject and the object, :END_ID is the entity: ID of the object, used to refer to the object in the relationship between the subject and the object, and :TYPE is used to represent the relationship between the subject and the object.
[0019] Steps 1-3 store the subject, object, and the relationship between the subject and the object represented in the manner of Step 1-2 in a CSV format file, and import the CSV format file into the NEO4J graph database to obtain a knowledge graph.
[0020] Preferably, the step (2) specifically includes the following steps:
[0021] Step 2-1: Obtain image I, resize image I from any size of P*Q to a fixed size of M*N,
[0022] (M, N) = Re(P, Q, Scale)
[0023] where Scale is the scaling factor and Re is the resizing function,
[0024] and use the pre-trained Faster R-CNN model to perform object detection on image I, thereby obtaining candidate regions. The candidate regions include a set of candidate boxes B = {b_i|i = 1,..., n} and global feature V, as shown in the following formula:
[0025] (B, V) = FasterRCNN(M, N, I)
[0026] Input the detected object into the ResNet network to extract object features, obtaining object feature X, as shown in the following formula:
[0027] X = ResNet(B, V)
[0028] 2-2: Process the object feature X using the SoftMax model as shown in the following formula to obtain the category L = {l_i|i = 1,..., n} of each object, where l_i ∈ Z^d;
[0029] L = SoftMax(X)
[0030] where l_i represents the finally predicted category and Z^d represents the predicted categories,
[0031] Then, process the triple information Triplets obtained in Step 1 through the TransR model to obtain a vector group T = [V_h, V_r, V_t], where V_h, V_r, and V_t respectively represent the vector of the subject in a triple, the vector of the relationship between the subject and the object, and the vector of the object, as shown in the following formula:
[0032] T = TransR(Triplets)
[0033] Finally, use Algorithm 1 to obtain the feature K optimized by the knowledge graph, as shown in the following formula:
[0034] K = Algorithm1(Maxnum, L, V, T)
[0035] Among them, Maxnum is the maximum number of triples in the search results;
[0036] The running process of the Algorithm1 algorithm includes the following steps:
[0037] a. Input the global feature V, the target feature X, the vector group T, and the maximum number of triples Maxnum;
[0038] b. Query the knowledge graph Maxnum times. The query condition is whether X is equal to T, and save the query result as Save l , where l = 1 to Maxnum;
[0039] c. Update the global feature
[0040] 2 - 3 Fuse the feature K from the knowledge graph, the target feature X, the category feature L, and the global feature V, and use the sum of the above features as the input F of the decoder of the LSTM:
[0041] F = SoftMax(f(K, X, L, V))
[0042] Among them, f = W v V + W x X + W l L + W k K
[0043] Among them, W v , W x , W l , W k are the corresponding weight values.
[0044] Preferably, step (3) specifically includes the following steps:
[0045] 3 - 1 Create an adaptive attention mechanism model,
[0046] Input the image I into the adaptive attention mechanism model, and the output of the adaptive attention mechanism model is Vc = [V1,..., V M
[0047] Among them, V1,..., V M are the features of M regions in the image,
[0048] Calculate the attention information of M regions in the image according to the method shown in the following formula: a t
[0049]
[0050] Calculate the content vector c by the method shown in the following formula t ,
[0051]
[0052] where Sigmoid is the activation function, W v , W g , are the corresponding weight values, and h t is the state of the hidden layer at time t; θ is a k*1 vector with all elements being 1, used to generate a k*k matrix;
[0053] In 3-2, introduce a visual sentinel to control visual information, and combine the knowledge graph and the already generated target feature X through weights to obtain s t ,
[0054] s t =σ(W x X + W k K + W h h t-1 )⊙Sigmoid(x t )
[0055] Finally, introduce s t into the adaptive attention mechanism model, and generate a new content vector through the adaptive attention mechanism model in the following way
[0056]
[0057] where W x , W k , W h are the corresponding weight values, t represents the time step, and β t is a parameter. When β t is 1, the text at the current time step depends on the prior knowledge of the knowledge graph and text information, and when β t is 0, it only depends on visual information;
[0058] Integrate the LSTM model, the adaptive attention mechanism model, and the knowledge graph through the above steps to obtain an image description model based on the adaptive attention mechanism and the knowledge graph. The input of the model is the picture to be described, and the output is the image description.
[0059] Preferably, in the step (2), the pre-trained Faster R-CNN model is trained by using the MSCOCO dataset and the Visual Genome dataset.
[0060] Preferably, in the step (2), the maximum number of triples in the search results is 8 groups.
[0061] The beneficial effects of the present invention are as follows:
[0062] The present invention creates a set of triple information that combines real-world common sense.
[0063] The present invention combines a knowledge graph to create an adaptive attention mechanism model, which can relatively accurately determine when to rely on knowledge graph information and language models, and when to rely on visual information, and further optimizes the detection of target relationships in images.
[0064] The present invention utilizes the rich semantic relationships contained in the knowledge graph to optimize object detection and object relationship recognition in the process of scene graph generation. The problems of image object recognition and relationship detection are further optimized.
[0065] The present invention tests the model based on the adaptive attention mechanism and the knowledge graph, and achieves relatively good results on both MSCOCO and VisualGenome. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is the overall architecture diagram of the present invention
[0067] Figure 2 is the knowledge graph embedding framework of the present invention
[0068] Figure 3 is the adaptive attention mechanism model of the present invention
[0069] Figure 4 is the explanation of the application of the knowledge graph in the present invention
[0070] Figure 5 is the step flow of the specific implementation 2 of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] First, the nouns involved in the embodiments of the present application are briefly introduced:
[0072] Knowledge graph: It is a semantic knowledge network that stores objective fact concepts and their mutual relationships, and has the characteristics of high structurality.
[0073] MSCOCO dataset: It is a large-scale image dataset developed and maintained by Microsoft, storing more than 330,000 images, and more than 200,000 images are annotated.
[0074] Visual Genome dataset: It is a large-scale dataset for picture semantic understanding released by the Fei-Fei Li group at Stanford University in 2016.
[0075] Information extractor OPENIE: It is an ontology-free information extraction paradigm, and the extraction form it generates is (subject; relation; object).
[0076] Triple: It is the basic unit of knowledge representation in a knowledge graph, used to represent the relationship between entities or what the attribute value of a certain attribute of an entity is.
[0077] Subject and object: They are the two major entities of a triple.
[0078] TransR model: It is a model used to complete entity and relationship embedding of a knowledge graph.
[0079] Word embedding vector: It is a mathematical embedding data from the one-dimensional space of a word to a continuous vector space with a lower dimension.
[0080] Faster R-CNN: It is a recurrent neural network model.
[0081] Visual sentinel: It is an additional implicit representation stored by the codec, providing a fallback option for the codec.
[0082] Non-visual words: They are some words that rely more on semantic information rather than visual information.
[0083] NEO4J graph database: It is a high-performance NoSQL graph database that stores structured data on a network rather than in a table.
[0084] LSTM: It is a special recurrent neural network in deep learning.
[0085] CNN: It is a convolutional neural network in deep learning.
[0086] The present invention will be further described below with reference to the accompanying drawings.
[0087] Embodiment 1
[0088] Refer to Figure 1The figure shows the overall architecture diagram of the present invention. The present invention is used for the task of generating image descriptions. First, the present invention uses the information extractor OPENIE to obtain triples from the given MSCOCO and Visual Genome datasets. As shown in the figure, the following three sentences are obtained: a small old style car driving on the road. (A small old-fashioned car driving on the road), a person traveling in a peddle cab down a city street. (A person traveling in a pedicab along a city street), a yellow vehicle with no doors driving down the street. (A yellow vehicle without doors driving on the street). At the same time, the present invention provides triple results for the optimization process of the algorithm according to the predicted target category labels, and provides triple features for the model. In order to fuse the information of the knowledge graph with the image description generation model, the present invention first performs object detection processing on the input image to generate new visual features and the information of the detected candidate boxes. Subsequently, in order to generate accurate image descriptions, the present invention uses two encoders and decoders to process the image and triple information. Among them, TransR realizes the encoding work of the triple features. After processing, the features are input into the image description generation module. And CNN realizes the encoding work of the image, inputs the features into the decoder, and the bidirectional LSTM is responsible for the decoding work of the features, and finally generates accurate image description information expressing the meaning. As shown in the figure, that is: a person driving a small yellow car through the streets of the city (A person driving a small yellow car through the streets of the city). Finally, in order to verify the effectiveness of the model, the present invention is tested on the MSCOCO and Visual Genome datasets. The results show that the method proposed by the present invention has achieved excellent performance in various evaluation indicators.
[0089] An image description generation method based on an adaptive attention mechanism and a knowledge graph, comprising the following steps:
[0090] Step (1) Convert the descriptions of images in the MSCOCO dataset and the Visual Genome dataset into a knowledge graph;
[0091] Step (2) Use the TransR model to vectorize and represent the knowledge graph to obtain word embedding vectors;
[0092] Step (3) generates an adaptive attention mechanism model through word embedding vectors and visual sentinels; integrates the LSTM model, the adaptive attention mechanism model, and the knowledge graph to obtain an image description model based on the adaptive attention mechanism and the knowledge graph. The input of the image description model based on the adaptive attention mechanism and the knowledge graph is the picture to be described, and the output is the image description.
[0093] Step (4) inputs the picture to be described into the image description model based on the adaptive attention mechanism and the knowledge graph obtained in step (3) to obtain the image description of the picture to be described.
[0094] Among them, the specific implementation process of step (1) is as follows:
[0095] 1-1 In the present invention, the descriptions of images in the most common MSCOCO dataset and Visual Genome dataset in image processing are saved in txt format, and the information extractor OPENIE is used to convert the data processing into triple information.
[0096] 1-2 The triple information includes a subject, an object, and the relationship between the subject and the object.
[0097] The subject and the object are represented in the following way: (entity: ID, name:, LABEL), where entity: ID is the entity ID, used to indicate which entities are in the picture, name is the entity attribute, used to describe what type of thing the entity belongs to, and LABEL is the label, used to describe the type of the entity attribute.
[0098] The relationship between the subject and the object is represented in the following way:
[0099] (:START_ID, :END_ID, :TYPE), where :START_ID is the entity: ID of the subject, used to refer to the subject in the relationship between the subject and the object, :END_ID is the entity: ID of the object, used to refer to the object in the relationship between the subject and the object, and :TYPE is used to represent the relationship between the subject and the object.
[0100] 1-3 Store the triple information including the subject, the object, and the relationship between the subject and the object in a file in csv format, and import the csv format file into the NEO4J graph database to obtain the knowledge graph, as Figure 4 shown, to achieve visualization and facilitate searching.
[0101] Furthermore, the specific implementation process of step (2) is as Figure 2 shown:
[0102] The specific implementation of step (2) is as follows:
[0103] 2-1 First, resize the input image I from an arbitrary size of P*Q to a fixed size of M*N, and use the pre-trained Faster R-CNN model on the MSCOCO dataset and the Visual Genome dataset to perform object detection on the image. As a result, candidate regions can be obtained. The candidate regions include a set of candidate boxes B = {b_i|i = 1,..., n} and global features V. Input the detected objects into the ResNet network to extract object features, obtaining object features X, as shown in formulas (1), (2), and (3):
[0104] (M, N) = Re(P, Q, Scale) (1)
[0105] (B, V) = Faster RCNN(M, N, I) (2)
[0106] X = ResNet(B, V) (3)
[0107] Among them, Scale is the scaling factor, and Re is the resizing function.
[0108] 2-2 Use the SoftMax model to process the object features X to obtain the category L = {l_i|i = 1,..., n} of each object, where l_i ∈ Z^d. Among them, l_i represents the finally predicted category, and Z^d represents the predicted categories.
[0109] Then, process the triple information Triplets obtained in step 1 through the TransR model to obtain a vector group T = [V_h, V_r, V_t], where V_h, V_r, and V_t represent the vector of the subject in a triple, the vector of the relationship between the subject and the object, and the vector of the object, respectively, as shown in formulas (4) and (5):
[0110] L = SoftMax(X) (4)
[0111] T = TransR(Triplets) (5)
[0112] Finally, use algorithm Algorithml to obtain the features K optimized by the knowledge graph, as shown in formula (6):
[0113] K = Algorithm1(Maxnum, L, V, T) (6)
[0114] Among them, Maxnum is the maximum number of triples in the search results, and the default is 8 groups.
[0115] The running process of the Algorithm1 algorithm includes the following steps:
[0116] a. Input the global feature V, the target feature X, the vector group T, and the maximum number of triples Maxnum;
[0117] b. Query the knowledge graph Maxnum times. The query condition is whether X is equal to T, and save the query result as Save l , where l = 1 to Maxnum;
[0118] c. Update the global feature
[0119] 2-3 This invention integrates branches from different sources, including the feature K from the knowledge graph, the target feature X, the category feature L, and the global feature V. And during model training, the loss function is set as the traditional cross-entropy loss. The sum of the above features is used as the input of the decoder of the LSTM, thereby generating the knowledge graph, relationship, and target combination result. As shown in formulas (7)(8):
[0120] F = SoftMax(f(K, X, L, V)) (7)
[0121] f = W v V + W x X + W l L + W k K (8)
[0122] Among them, W v , W x , W l , W k are the corresponding weight values.
[0123] Furthermore, the specific implementation process of step (3) is as Figure 3 shown:
[0124] The specific implementation process of step (3) is as follows:
[0125] 3-1 Create an adaptive attention mechanism model. The adaptive attention mechanism model is used to let the adaptive attention mechanism model itself judge "where to look at the picture" and "when to look at the picture" when generating word embedding vectors in step 2.
[0126] First, input the image I into the adaptive attention mechanism model. The adaptive attention mechanism model outputs Vc = [V1,..., V M
[0127] Among them, V1,..., V M are the features of M regions in the image, obtaining the attention information of M regions in the image: a t , and finally obtaining the content vector ct, as shown in formulas (9)(10)(11):
[0128]
[0129] a t = softmax(z t ) (10)
[0130]
[0131] where Sigmoid is the activation function, W v , W g , are the corresponding weight values, h t is the state of the hidden layer at time t, θ is a k*1 vector with all elements being 1, and the ultimate goal is to generate a matrix of size k*k.
[0132] In 3-2, a visual sentinel is introduced to control visual information, and the knowledge graph and the already generated target feature X are combined by weights into s t , and finally s t is introduced into the adaptive attention mechanism model, and the adaptive attention mechanism model generates a new content vector as shown in formulas (12), (13), and (14):
[0133] g t = σ(W x X + W k K + W h h t-1 ) (12)
[0134] s t = gt⊙Sigmoid(x t ) (13)
[0135]
[0136] where W x , W k , W h are the corresponding weight values, t represents the time step, β t is a parameter. When β t is 1, the current time step text depends on the prior knowledge of the knowledge graph and text information, and when β t is 0, it only depends on visual information.
[0137] By integrating the LSTM model, the adaptive attention mechanism model, and the knowledge graph through the above steps, an image description model based on the adaptive attention mechanism and the knowledge graph is obtained. The input of the model is the picture to be described, and the output is the image description.
[0138] Step (4) inputs the picture to be described into the image description model based on the adaptive attention mechanism and knowledge graph obtained in step (3), and obtains the image description of the picture to be described. Specific Embodiment 2
[0140] Using the image description method of this patent invention, Figure 5 as the input image, first through the operation of step (1), a knowledge graph containing a large amount of common sense knowledge is generated.
[0141] (entity: ID, name:, LABEL), entity: ID is the entity ID, used to indicate which entities are in the picture. In this example, there are three entities: woman (female), tennis (tennis ball), and man (male). name is the entity attribute, used to describe what type of thing the entity belongs to. In this example, the entity attribute corresponding to the woman (female) entity is female. LABEL is the label, used to describe the type of the entity attribute. In this example, the label of the entity attribute of the woman (female) entity is person (person).
[0142] (:START_ID, :END_ID, :TYPE); :START_ID is the entity: ID of the subject, used to refer to the subject in the relationship between the subject and the object. In this example, the entity: ID of woman (female) is 0; :END_ID is the entity: ID of the object, used to refer to the object in the relationship between the subject and the object. In this example, the entity: ID of tennis (tennis ball) is 1; :TYPE is used to represent the relationship between the subject and the object. The relationship in this example is in front of (in front).
[0143] Subsequently, operate according to step (2) to resize the image into a 3*3 format, and use the model to obtain candidate regions A, B, and C through object detection. Predict the categories of the candidate regions as A: woman, B: tennis, C: man, and predict the relationships between pairwise candidate boxes: between A and B: woman-in front of tennis, between B and C: man-next to-tennis, between A and C: man-stand opposite-woman. Then, operate according to step (4) to vectorize the object detection and the predicted relationships between objects, and put A, B, C and their mutual relationships into the knowledge graph to find more appropriate relationships, obtaining new relationships that are more in line with common sense: between A and B: woman-standing on-tennis, between B and C: man-is facing-tennis, between A and C: man-consult with each other-woman. Finally, integrate the relationships generated by the knowledge graph into the word embedding vectors. As shown in step (5), during the training of the model, rely on the adaptive attention mechanism to determine "where to look at the picture" and "when to look at the picture", and finally generate a better description through bidirectional LSTM: a young woman standing on a tennis court facing an man.
[0144] Based on the image description generation method of the present patent invention, a comparative experiment is carried out with the currently popular image description generation methods. The comparison results are shown in Table 1:
[0145]
[0146] As can be seen from Table 1, in the MSCOCO dataset experiment, our method achieved scores of 39.2, 29.3, 60.1, 133.1, and 23.2 on BLEU@4, METEOR, ROUGE-L, CIDE, and SPICE respectively. Generally speaking, compared with other models (such as SCST.SG, SGAE), our method achieved more excellent results, which shows the effectiveness of adding the knowledge graph as prior knowledge and using the attention mechanism in image description for improving the model performance. However, due to the limitations of the encoder-decoder structure, the metrics of this model are slightly lower than those of the method based on the Transformer model (such as HIP) on METEOR and SPICE.
Claims
1. An image description generation method based on an adaptive attention mechanism and a knowledge graph, characterized in that, It includes the following steps: Step (1): Obtain the MSCOCO dataset and the Visual Genome dataset, and convert the descriptions of the images in the MSCOCO dataset and the Visual Genome dataset into a knowledge graph; Step (2): Use the TransR model to vectorize and represent the knowledge graph to obtain word embedding vectors; Step (3): Generate an adaptive attention mechanism model through the word embedding vectors and visual sentinels; Based on the LSTM model, combine the adaptive attention mechanism model and the knowledge graph to obtain an image description model based on the adaptive attention mechanism and the knowledge graph; Step (4): Input the picture to be described into the image description model based on the adaptive attention mechanism and the knowledge graph obtained in Step (3) to obtain the image description of the picture to be described; The specific steps of Step (2) include the following steps: Step 2-1: Obtain image I, resize image I from an arbitrary size of P*Q to a fixed size of M*N, (M, N) = Re(P, Q, Scale) where Scale is the scaling factor and Re is the resizing function, and use the pre-trained Faster R-CNN model to perform object detection on image I, thereby obtaining candidate regions. The candidate regions include a set of candidate boxes B = {b_i│i = 1,…,n} and global features V, as shown in the following formula: (B, V) = Faster RCNN(M, N, I) Input the detected objects into the ResNet network to extract object features to obtain object features X, as shown in the following formula: X = ResNet(B, V) 2-2: Process the object features X using the SoftMax model as shown in the following formula to obtain the category L = {l_i│i = 1,…,n} of each object, l_i ∈ Z^d; L = SoftMax(X) where l_i represents the finally predicted category and Z^d represents the predicted categories, Then, the triple information Triplets obtained in Step 1 is processed by the TransR model to obtain a vector group T = [V_h, V_r, V_t], where V_h, V_r, and V_t represent the vector of the subject in a triple, the vector of the relationship between the subject and the object, and the vector of the object, respectively, as shown in the following formula: T = TransR(Triplets) Finally, use the Algorithm1 algorithm to obtain the features K optimized by the knowledge graph, as shown in the following formula: K = Algorithm1(Maxnum, L, V, T) where Maxnum is the maximum number of triples in the search results; The running process of the Algorithm1 algorithm includes the following steps: a. Input the global feature V, object feature X, vector group T, and the maximum number of triples Maxnum; b. Query the knowledge graph Maxnum times with the query condition of whether X is equal to T, and save the query result as Save l , where l = 1 to Maxnum; c. Update the global feature 2-3: Fuse the features K from the knowledge graph, object feature X, category feature L, and global feature V, and use the sum of the above features as the input F of the decoder of the LSTM: F = SoftMax(f(K, X, L, V)) where f = W v V + W x X + W l L + W k K Among them, W v , W x , W l , W k are the corresponding weight values.
2. The method for generating an image description based on an adaptive attention mechanism and a knowledge graph according to claim 1, wherein The specific steps of step (1) include the following steps: Step 1-1: Save the descriptions of the images in the MSCOCO dataset and the Visual Genome dataset in txt format, and use the information extractor OPENIE to convert the descriptions of the images into triplet information Triplets. The triplet information Triplets includes the subject, the object, and the relationship between the subject and the object. Step 1-2: Represent the subject and the object in the following way: (entity:ID, name:, LABEL), where entity:ID is the entity ID of the subject or the object, used to indicate which entities are in the picture, name is the entity attribute of the subject or the object, used to describe what type of thing the entity belongs to, and LABEL is the label, used to describe the type of the entity attribute. Represent the relationship between the subject and the object in the following way: (:START_ID, :END_ID, :TYPE), where :START_ID is the entity:ID of the subject, used to refer to the subject in the relationship between the subject and the object, :END_ID is the entity:ID of the object, used to refer to the object in the relationship between the subject and the object, and :TYPE is used to represent the relationship between the subject and the object. Step 1-3: Store the subject, the object, and the relationship between the subject and the object represented in the way of step 1-2 in a csv format file, and import the csv format file into the NEO4J graph database to obtain a knowledge graph.
3. The method for generating an image description based on an adaptive attention mechanism and a knowledge graph according to claim 2, wherein The specific steps of step (3) include the following steps: 3-1: Create an adaptive attention mechanism model. Input the image I into the adaptive attention mechanism model, and the output of the adaptive attention mechanism model is Vc = [V1, …, V M where V1, …, V M are the features of M regions in the image, Calculate the attention information a of M regions in the image according to the method shown in the following formula t : Calculate the content vector c according to the method shown in the following formula t , Among them, Sigmoid is the activation function, W v , W g , are the corresponding weight values, h t is the state of the hidden layer at time t; θ is a k*1 vector with all elements being 1, used to generate a k*k matrix; 3-2 Introduce a visual sentinel to control visual information, and combine the knowledge graph and the already generated target feature X by weights into s t , s t = σ(W x X + W k K + W h h t-1 ) ⊙ Sigmoid(x t ) Finally, introduce s t into the adaptive attention mechanism model, and generate a new content vector through the adaptive attention mechanism model in the following formula Among them, W x , W k , W h are the corresponding weight values, t represents the time step, and β t is a parameter. When β t is 1, the current time-step text depends on the prior knowledge of the knowledge graph and text information, while when β t is 0, it only depends on visual information; Integrate the LSTM model, the adaptive attention mechanism model, and the knowledge graph through the above steps to obtain an image description model based on the adaptive attention mechanism and the knowledge graph. The input of the model is the picture to be described, and the output is the image description.
4. The method for generating image descriptions based on an adaptive attention mechanism and a knowledge graph according to claim 2, wherein In step (2), the pre-trained Faster R-CNN model is trained through the MSCOCO dataset and the Visual Genome dataset.
5. The method for generating image descriptions based on an adaptive attention mechanism and a knowledge graph according to claim 2, characterized in that In step (2), the maximum number of triplets in the search results is 8 groups.
Citation Information
Patent Citations
Image scene graph generation method based on knowledge graph
CN114329010A
Knowledge graph construction method and device based on document relationship extraction
CN115269857A