A Video Description Generation Method Based on Transitive Visual Relationship Detection

Through the transitive visual relationship detection module and graph convolutional neural network, the problem of visual relationship detection and feature representation in the video description generation algorithm is solved, and a more accurate natural language description is generated.

CN114037936BActive Publication Date: 2025-07-29HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111314705.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-08
Publication Date
2025-07-29
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

Existing video description generation algorithms have difficulties in detecting visual relationships between multiple entities in video and refine video feature representations, especially without the help of external knowledge bases or pre-trained models, model training is difficult and noise is introduced, affecting algorithm performance.

Method used

The transitive visual relationship detection module is adopted, including the shallow relationship detection of action-guided and the transitive deep relationship inference module. Combined with the graph convolutional neural network to refine the video feature representation, through the shallow relationship detection of action-guided and deep relationship inference, accurate natural language descriptions are generated using video and text features.

Benefits of technology

It realizes that the visual relationship between key entities and actions in the video is effectively detected without relying on external knowledge bases or pre-trained models, and the video feature representation is refined, generating a more accurate and smooth natural language description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114037936B_ABST
    Figure CN114037936B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating video descriptions based on transitive visual relationship detection. In particular, it relates to a modeling method for detecting shallow connections between visual entities - actions, transmitting and constructing a deep visual entity relationship graph, and refining video features relying on the visual entity relationship graph. The present invention includes the following steps: 1. Data preprocessing, extracting features from the video and constructing a dictionary for the text description. 2. A shallow relationship detection module guided by actions for generating a shallow relationship graph. 3. A transitive deep relationship reasoning module and a decoder module for reasoning about the deep relationship graph. 4. Model training, using the backpropagation algorithm to train the neural network parameters. The present invention proposes a modeling method for detecting shallow connections between video visual entities - actions, transmitting and constructing a deep visual entity relationship graph, and refining video features relying on the visual entity relationship graph, and has achieved the best results in the field of video description generation at present.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a deep neural network for video captioning (VC), and particularly to a modeling method for detecting shallow connections between visual entities and actions, transmitting and constructing a deep visual entity relationship graph, and refining video features relying on the visual entity relationship graph. Background Art

[0002] "Cross-media" unified representation is an intersection direction between the fields of computer vision and natural language processing research, aiming to bridge the "semantic gap" between different media (such as video and text) and establish a unified semantic representation. Based on the theoretical methods of cross-media unified representation, some current popular research directions have emerged, such as video captioning, video-text cross-media retrieval, visual relation detection, and video question answering (VideoQA), etc. The goal of video captioning is to use one or a few natural language sentences to summarize the content of a given video; video-text cross-media retrieval aims to find the most matching text description from a database given a video, or find the most matching video given a text description; the goal of visual relation detection is to input a video, and the model detects the connections between key entities in the video; the goal of video question answering is to input a video and a question described in natural language, and the algorithm automatically outputs an answer described in natural language.

[0003] With the rapid development of deep learning in recent years, using deep neural network frameworks, such as the encoder-decoder framework, for end-to-end modeling has become the mainstream research direction in the fields of computer vision and natural language processing. In video captioning algorithms, how to make the feature representation contain deeper semantic content under the condition of deep modeling of the video feature representation only in the encoder part without relying on an external knowledge base or a large-scale pre-trained model, so as to directly output accurate and fluent natural language descriptions when input into the decoder is a research question worthy of in-depth exploration.

[0004] In terms of practical applications, the video description generation algorithm has a very wide range of application scenarios. Image-based description generation systems have been widely used in practical scenarios such as automatic news writing, medical report generation, and picture review. With the rapid development of video platforms (such as Douyin and YouTube) and Internet technologies, currently, video-based text description generation has gradually become one of the popular research directions in computer vision. This technology can help us understand video content more simply and quickly, facilitating subsequent operations such as review and classification.

[0005] In summary, the video description generation algorithm based on the encoder-decoder structure is a direction worthy of in-depth research. This topic intends to start from several key difficult problems in this task, solve the problems existing in the current methods, and finally form a complete video description generation system.

[0006] Due to the complex video content, large number of entities, and complex relationships between entities in natural scenes, and the high degree of freedom in natural language description problems, the video description generation algorithm faces huge challenges. Specifically, there are mainly the following two aspects of difficulties:

[0007] (1) How to detect the visual relationships between multiple entities in a video: The problem of visual relationship detection is a classic and fundamental problem in computer vision. Its application in image-level data has been relatively mature, and common methods include object detection, action detection, or using scene graphs to represent, etc. In addition, models based on scene graph generation have achieved very good results in many computer vision fields, such as image description generation, image content question answering, and image retrieval. However, due to problems such as temporal entity redundancy and complex actions in video-level data, how to detect the visual relationships contained in a video and screen out key relationships is a great challenge. Therefore, visual relationship detection in video-level data is a direction worthy of in-depth research.

[0008] (2) How to refine the feature representation of a video using visual relationships: In video description generation algorithms, the input to the decoder only includes the video feature representation output by the encoder. To accurately describe video content, the visual features need to contain the global content of the video, as well as the feature representations of key objects and key actions. Currently, many video description generation algorithms jointly construct visual feature representations using external knowledge bases and pre-trained models. However, since external knowledge bases contain excessive common sense information and pre-trained models contain dataset information unrelated to this task, noise is introduced, making model training difficult. Therefore, how to enable the encoder to refine the video feature representation using only visual relationships without relying on external knowledge bases or pre-trained models is a difficult problem in video description generation algorithms and is also a crucial link affecting the performance of algorithm results. Summary of the Invention

[0009] The objective of the present invention is to provide a video description generation method based on transitive visual relationship detection in view of the deficiencies of the prior art.

[0010] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0011] Given a video v and its corresponding text description c to form a video-description pair v, c as the training set, a transitive visual relationship detection module is proposed. The transitive visual relationship detection module includes an action-guided shallow relationship detection module and a transitive deep relationship reasoning module.

[0012] Step (1), data preprocessing: Extract features from the video and construct a dictionary for the text description;

[0013] Preprocessing of video v: First, all videos are frame-extracted and each frame is scaled to a unified size. Then, different deep neural networks, namely 2D-CNN, C3D, and Faster-RCNN, are used to extract the video features V r 、V m and V o .

[0014] Preprocessing of text description c:

[0015] 1. Extract the verbs contained in the description and construct a verb embedding E a : First, use open-source tools to extract the subject-predicate-object triples in the description, select the predicate as the verb contained in the description, and construct a verb embedding based on all the verbs in the dataset.

[0016] 2. Construct a description embedding E c: Tokenize all the descriptions in the dataset and count the occurrences of each word. Discard words with occurrences less than the set threshold, and construct a descriptor embedding based on all the remaining words.

[0017] Step (2), the Action-guided Shallow Relationship Detection module;

[0018] Its process is as Figure 1 shown in part a) in []. Considering that the actions in the video are all initiated by entities, this module extracts the video action representation using the three-dimensional features and entity features of the video, and maps the action representation to a vector with the size of the verb dictionary using a non-linear mapping layer. Each value in the vector represents the probability of the corresponding verb. According to this vector, select twenty verbs a i (i = 1,..., 20) as the verbs included in this video. Based on the entities and verbs included in the video, construct an Object-Action relationship graph G oa .

[0019] Step (3), the Transitive Deep-level Relationship Reasoning module and the decoder module;

[0020] Its process is as Figure 1 shown in part b) in []. Perform a matrix multiplication operation on the Object-Action relationship graph and its transposed graph, so as to construct the relationship between entities transitively using the relationship between entities and actions. The output of the matrix multiplication is used as the deep Object-Object relationship graph G oo . After that, regard the Object-Object relationship graph as a graph, and the entity features as nodes. Use a graph convolutional neural network (GCN) to refine the feature representation of the nodes, and encode the relationship between entities into the entity feature representation. Finally, jointly construct the encoded feature V of the video based on the refined entity feature representation and the global feature enc , and use it as the input of the decoder. The decoder used in this article is LSTM. The input at each time step t is the hidden layer feature h t-1 of the previous time step, the word encoding E c y t-1 generated at the previous time step, the verb encoding V a , and the video encoded feature V enc . The output is the hidden layer feature h t of the current time step and the generated word probability distribution p θ (yt ) and finally generate the word at the current time step according to the probability distribution.

[0021] Step (4), Model Training

[0022] Calculate the negative log-likelihood loss based on the difference between the predicted verb and description and the actual verb and description of the video, and use the back-propagation algorithm (BP) to train the model parameters of the neural network defined above until the entire network model converges.

[0023] The data preprocessing described in step (1) is to extract features and construct word embeddings for the video:

[0024] 1-1. For video v, use existing deep neural networks to extract two-dimensional features three-dimensional features and entity features where d r , d m and d o are the sizes of two-dimensional, three-dimensional, and entity features respectively, and N is the number of entities in the video.

[0025] 1-2. For text description c, first use the open-source tool of nltk to extract the subject-predicate-object structure in the description, where the predicate, which is also the true verb a included in the video * is extracted to construct the verb embedding E a :

[0026]

[0027] where is the word embedding of the i-th verb, and i is the index value of the verb in the word embedding, and d w is the size of the word embedding.

[0028] Subsequently, all descriptions are tokenized and the number of occurrences of each word is counted. After discarding the words with less than 2 occurrences, it is used as the true description y of the video * , and all the remaining words are extracted to construct the description word embedding E c :

[0029]

[0030] where is the word embedding of the i-th word, and j is the index value of the word in the word embedding.

[0031] The Action-guided Shallow Relationship Detection module described in step (2), which includes an Action Detection Module and uses matrix multiplication of action and entity feature representations to obtain a shallow entity-action relationship graph. The specific process is as follows:

[0032] 2-1. In the action detection module, first set the parameter variables where is the mapping matrix, d is the output dimension, and the attention feature V att can be further calculated through the following formula:

[0033]

[0034] to obtain f att After that, use a one-layer Feed-Forward network to calculate the attention feature. The specific formula is as follows:

[0035] V att = LayerNorm(f att + W down (σ(W up f att ))) (Formula 4)

[0036] where and are the upsampling and downsampling mapping matrices respectively, d up is the sampling dimension, and σ is the ReLU activation function. Subsequently, the attention feature and the three-dimensional feature are mapped to the same space and then concatenated. The specific formula is as follows:

[0037] V a = [f map (V att ); f map (V m )] (Formula 5)

[0038] V a in Formula (5) is the action representation, and f map (x) is the mapping function, which is obtained by adding weights to a one-layer fully connected layer and then passing through a one-layer ReLU activation function. The specific formula is as follows

[0039] f map (x) = σ(W x x + b x ) (Formula 6)

[0040] where W xand b x are the mapping matrix and the weight vector. After obtaining the action representation, map it to a vector with the size of the action vocabulary dimension, convert each value in the vector into a probability value, and select the action of the video according to the size of the probability value. The specific formula is as follows:

[0041]

[0042] where is the mapping matrix, and l a is the size of the verb dictionary. Calculate the probability of each verb in the verb vocabulary according to formula (7) Select the 20 verbs with the highest probability Calculate the verb features using verb embeddings. The formula for the word features is as follows:

[0043]

[0044] where is the verb feature, and d a is the dimension of the verb feature.

[0045] 2-2. After obtaining the verb features, calculate the shallow entity-action relationship graph by combining the entity features. The formula is as follows:

[0046]

[0047] where is the shallow entity-action relationship graph. The values in this graph are all real numbers between 0 and 1, and through regularization, the sum of the probability values of each entity associated with the same verb is 1; W v and W l are the mapping matrices used to map entity and verb features to the same space.

[0048] The Transitive Deep-level Relationship Reasoning module and the decoder described in step (3) are as follows:

[0049] 3-1. In order to convert the connection between entities and verbs in the shallow entity-action relationship graph G oa into the connection between entities and entities, multiply the G oa generated in step (2) by its own transpose, and then obtain the deep entity-entity relationship graph after regularization. The formula is as follows:

[0050]

[0051] where It is a deep entity-entity relationship diagram, and each value in the diagram represents the degree of association between the entities at the corresponding positions. Subsequently, five most relevant entities are selected for each entity, and the association values are retained, while the association degrees of the remaining irrelevant entities are all set to 0, thus obtaining the filtered deep association diagram

[0052] 3-2. After obtaining the deep relationship diagram, use the graph convolutional neural network (GCN) to refine the entity features according to the values in the deep relationship diagram, enabling the features between associated entities to interact. The calculation formula is as follows:

[0053]

[0054] where is the refined entity feature, and W g is the learnable matrix in the graph convolutional neural network. After obtaining the refined features, we use formulas (3) and (4) to calculate the encoder feature V enc , where Q, K, and V in formula (3) are respectively W g′ V g , and are replaced, where V g is the global feature concatenated by V r and V m , and W g′ , and are all mapping matrices.

[0055] 3-3. After obtaining the encoder feature, use the LSTM network to generate the text description. At each time step t, the output of the LSTM is the hidden layer feature h t and the decoding feature o t , and then the word y t generated at the current time step is calculated according to the decoding feature. The input of the LSTM is calculated from the hidden layer feature h t-1 output at the previous time step t-1, the word y t-1 and the video feature. The calculation formula for the input part of the LSTM is as follows:

[0056]

[0057] where w t is the text feature calculated from the hidden layer feature, word, and verb feature obtained from the previous time step, where α1 and α2 are two values in α.

[0058] Then the text feature is concatenated with the global feature and the encoder feature to form the input of the LSTM. The specific calculation formula for the LSTM part is as follows:

[0059] {o t ,h t} = LSTM([w t ;V g ;V enc , h t-1 ) (Equation 13)

[0060] After obtaining the decoded features {o t , h t}, use a mapping matrix to map its dimensionality to the size of the descriptor embedding, and after normalizing it, obtain the probability distribution of the word generated at the current time step. The calculation formula is as follows:

[0061]

[0062] where is the probability distribution of the word generated at the current time step, is the mapping matrix. Finally, after obtaining the probability distribution, the word with the highest probability can be selected from it as the word output y t at the current time step. The calculation formula is as follows:

[0063]

[0064] The training model described in step (4) is as follows:

[0065] 4 - 1. Input the verb probability distribution generated in step (2) and the true verb a * extracted in step (1) into the negative log-likelihood loss function to obtain the loss value The specific formula is as follows:

[0066]

[0067] 4 - 2. Input the probability distribution generated in step (3) and the true description generated in step (1) into the negative log-likelihood loss function to obtain the loss value The specific formula is as follows:

[0068]

[0069] 4 - 3. Calculate the final loss based on the calculated loss values and . The calculation formula is as follows:

[0070]

[0071] Where λ is a preset hyperparameter. Finally, the backpropagation algorithm (BP) is used to adjust the parameters in the network.

[0072] The beneficial effects of the present invention are as follows:

[0073] The present invention proposes a deep neural network architecture for video caption generation tasks to solve the above two difficult problems. 1. A transitive visual relationship detection module is proposed to detect shallow entity-action relationships and deep entity-entity relationships; 2. A model based on graph convolutional neural network is proposed to refine video feature representations using visual relationships, so that the encoder features contain information about global content, key entities, and key actions at the same time. Brief Description of the Drawings

[0074] Figure 1 It is a flowchart of the present invention. Detailed Embodiments

[0075] The detailed parameters of the present invention will be further specifically described below.

[0076] As Figure 1 shown, the present invention provides a deep neural network framework for video captioning (VC).

[0077] The data preprocessing and feature extraction of the video in step (1) and the construction of a dictionary for the text description are as follows:

[0078] Here, the MSR-VTT and MSVD datasets are used as experimental data.

[0079] 1-1. For video data, an existing normalized deep residual network (InceptionNetV2) model is used here to extract video frame features, a three-dimensional convolutional neural network (C3D) model is used to extract video dynamic features, and a target recognition convolutional neural network (Faster-RCNN) model is used to extract entity features in video frames. Specifically, the video frames are cropped and scaled to a unified size, and are input into the normalized deep residual network frame by frame, and the obtained features are subjected to average pooling operation to obtain two-dimensional features Similarly, the video frames are input into the three-dimensional convolutional neural network together and subjected to average pooling operation to obtain three-dimensional features For entity features, 10 entities are detected on each video frame to extract features, a total of 10 video frames are detected, and all the entity features are concatenated to finally obtain entity features

[0080] 1-2. For the text description, first use the open-source tool of nltk to extract the subject-predicate-object structure in the description and extract the predicates as the verbs of the video. Subsequently, delete the verbs that appear less than 2 times, and use the remaining verbs to construct verb embeddings with GloVe. The length l of the verb embeddings in the MSR-VTT dataset a is 2615, and in MSVD it is 1338.

[0081] Subsequently, tokenize all the descriptions and count the number of occurrences of each word, delete the words that appear less than 2 times, and use the remaining words to construct description embeddings. The length l of the description embeddings in the MSR-VTT dataset c is 10536, and in MSVD it is 4064.

[0082] The Action-guided Shallow Relationship Detection module described in step (2) is as follows:

[0083] 2-1. For the input features and use the mapping matrices and to map them to 512-dimensional vectors respectively. In the feed-forward network, the sampling dimension d up is 2048. The mapping function f map (x) and are the mapping matrix and the weight vector. In the verb features, the verb feature dimension d a is 512.

[0084] 2-2. Use matrix multiplication operations to multiply the verb features and the entity features, where and are the mapping matrices used to map the entities and verb features to the same space. The size of the generated shallow relationship graph is 100×20.

[0085] The Transitive Deep-level Relationship Reasoning module and the decoder module described in step (3) are as follows:

[0086] 3-1. The size of the deep relationship graph generated based on the shallow relationship graph is 100×100.

[0087] 3-2. After obtaining the deep relationship graph, use the graph convolutional neural network (GCN) to refine the entity features according to the values in the deep relationship graph, where is a learnable matrix in the graph convolutional network. In the formula for calculating the encoder features, is the global feature, and are both mapping matrices.

[0088] 3-3. After obtaining the encoder features, use the LSTM network to generate text descriptions, where the hidden layer feature h t and the decoded feature o t have a feature dimension of 512 dimensions, and the dimension of the generated probability distribution vector is the same as the dimension of the corresponding word embedding. and are both mapping matrices. After obtaining the decoder features, calculate the word probability distribution, where l is the word embedding dimension of the corresponding dataset.

[0089] The training model described in step (4) is as follows:

[0090] For the predicted verb vector and description vector generated in steps (2) and (3), compare them with the correct answer to this question, calculate the difference between the predicted value and the actual correct value through the defined negative log-likelihood loss function to form a loss value, and then adjust the parameter values of the entire network according to this loss value using the backpropagation algorithm (BP) until the network converges, where λ is 0.2.

[0091] Table 1 shows the scores of four evaluation metrics of the method described in this article in the MSRVTT and MSVD datasets. Among them, BLEU@4, ROUGE-L, and METEOR represent the article accuracy, and CIDEr represents the article fluency. SAAT and GRU-EVE are both existing excellent methods. The method on the MSR-VTT dataset exceeds these two methods in all metrics. On the MSVD dataset, the GRU-EVE method is slightly higher than the method of the present invention in the METEOR metric, but other metrics are far lower than the method.

[0092]

[0093]

Claims

1. A video description generation method based on transitive visual relationship detection, characterized in that It includes the following steps: Given a video v and its corresponding text description c to form a video-description pair v, c as the training set, the proposed transitive visual relation detection module includes an action-guided shallow relation detection module and a transitive deep relation reasoning module; Step (1), data preprocessing: extract features from the video and construct a dictionary for the text description; Preprocessing of video v: First, all videos are frame-extracted and each frame is scaled to a unified size, and then different deep neural networks are used to extract the features of the video; Preprocessing of text description c: Extract the verbs contained in the text description c and construct the verb embedding Ε a : First, use an open-source tool to extract the subject-predicate-object triples in the description, select the predicates as the verbs contained in the description, and construct the verb embedding Ε based on all the verbs in the dataset a ; Construct the descriptor embedding Ε c : Tokenize all the descriptions in the dataset and count the occurrences of each word. Discard the words with occurrences less than the set threshold, and construct the descriptor embedding Ε based on all the remaining words c ; Step (2), construct an action-guided shallow relation detection module; Extract video action representations using the 3D features and entity features of the video, and use a non-linear mapping layer to map the action representations to a vector with the size of the verb dictionary dimension. Each value in the vector represents the probability of the corresponding verb. Select twenty verbs with the highest probabilities according to this vector, a i As the verbs included in the video, i = 1, …, 20; construct an entity-action relationship graph G based on the entities and verbs included in the video oa ; Step (3), construct a transitive deep relation reasoning module and a decoder; Multiply the entity-action relationship graph G oa by its transposed graph through matrix multiplication operation, so as to construct the relationship between entities transitively by using the relationship between entities and actions; the output of the matrix multiplication is used as the deep entity-entity relationship graph G oo . Then, regard the entity-entity relationship graph G oo as a graph, regard entity features as nodes, and use a graph convolutional neural network to refine the feature representation of the nodes, and encode the relationship between entities into the entity feature representation; finally, jointly construct the encoded feature V of the video according to the refined entity feature representation and the global feature enc , and use it as the input of the decoder; the decoder is an LSTM, and the input at each time step t is the hidden layer feature h t-1 of the previous time step, the word encoding Ε c y t-1 generated at the previous time step, the verb encoding V a , and the video encoded feature V enc . The output is the hidden layer feature h t of the current time step and the generated word probability distribution p θ (y t ). Finally, generate the word of the current time step according to the probability distribution; Step (4), model training Calculate the negative log-likelihood loss according to the difference between the predicted verb and description and the actual verb and description of the video, and use the backpropagation algorithm to train the model parameters of the neural network until the entire network model converges; The specific implementation of step (1) is as follows: 1-1. For video v, different deep neural networks, namely 2D-CNN, C3D, and Faster-RCNN, are used to extract the features V r , V m and V o ; two-dimensional features three-dimensional features and entity features where d r , d m and d o are the dimensions of two-dimensional, three-dimensional, and entity features respectively, and n is the number of entities in the video; 1-2. For the text description c, first use the open-source tool of nltk to extract the subject-predicate-object structure in the description, where the predicate is also the real verb a contained in this video * which is extracted to construct the verb embedding E a : wherein is the word embedding of the i-th verb, where i is the index value of the verb in the word embedding, and d w is the size of the word embedding; 1-3. Tokenize all descriptions and count the occurrences of each word. After discarding words with less than 2 occurrences, use the remaining words as the true description y of the video * and extract all the remaining words to construct the descriptor embedding Ε c : Among them, is the word embedding of the i-th word, and j is the index value of the word in the word embedding; The specific implementation of step (2) is as follows: The shallow relation detection module includes a narrative action detection module, and uses matrix multiplication with action and entity feature representations to obtain a shallow entity-action relation graph. The specific process is as follows: In the motion detection module, first set the parameter variable Q = W m V m , where is the mapping matrix, d is the output dimension, and the attention feature V att can be further calculated by the following formula: Obtain f att After that, a feed-forward network is used to calculate the attention features, and the specific formula is as follows: V att = LayerNorm(f att + W down (σ(W up f att ))) (Formula 4) where and are the upsampling and downsampling mapping matrices respectively, d up is the sampling dimension, and σ is the ReLU activation function; subsequently, the attention features and the three-dimensional features are mapped to the same space and then concatenated. The specific formula is as follows: V a = [f map (V att );f map (V m )] (Formula 5) V in formula (5) a is the action representation, and f map (x) is the mapping function, which is obtained by adding weights to a fully connected layer and then passing through a ReLU activation function. The specific formula is as follows f map (x) = σ(W x x + b x )(Equation 6) Among which W x and b x are the mapping matrix and the weight vector; after obtaining the action representation, map it to a vector with the dimension of the action vocabulary size, convert each value in the vector into a probability value, and select the action of the video according to the magnitude of the probability value. The specific formula is as follows: Among them is the mapping matrix, and l a is the size of the verb dictionary; calculate the probability of each verb in the verb vocabulary according to formula (7) Select the 20 verbs with the highest probabilities among them Calculate the verb features using verb embeddings; the formula for the word features is as follows: Among them is a verb feature, d a is the dimension of the verb feature; 2-2. After obtaining the verb features, calculate the shallow entity-action relation graph in combination with the entity features. The calculation formula is as follows: Among them, is a shallow entity-action relationship graph, and the values in this graph are all real numbers between 0 and 1, and through regularization, the probability values of each entity associated with the same verb are added up to 1; W v and W l are mapping matrices used to map entity and verb features to the same space.

2. The video description generation method based on transitive visual relationship detection according to claim 1, wherein The specific implementation of step (3) is as follows: 3-1. In order to transform the connection between entities and verbs in the shallow entity-action relationship graph G oa into the connection between entities and entities, multiply the G oa generated in step (2) by its own transpose, and then perform regularization to obtain the deep entity-entity relationship graph. The calculation formula is as follows: Among them is a deep entity-entity relationship diagram, and each value in this diagram represents the degree of association between the entities at the corresponding positions; then, for each entity, the five most relevant entities are selected and the association values are retained, while the association degrees of the remaining irrelevant entities are all set to 0, so as to obtain the filtered deep association diagram 3-2. After obtaining the deep relation graph, use the graph convolutional neural network to refine the entity features according to the values in the deep relation graph, so that the features between related entities interact. The calculation formula is as follows: Among them is the refined entity feature, W g is the learnable matrix in the graph convolutional neural network; after obtaining the refined feature, we calculate the encoder feature V using formulas (3) and (4) enc , where Q, K, and V in formula (3) are W g' V g , and are replaced respectively, where V g is the global feature concatenated by V r and V m , W g' , and are all mapping matrices; After obtaining the encoder features, use the LSTM network to generate text descriptions; at each time step t, the output of the LSTM is the hidden layer feature h t and the decoded feature o t , and then calculate the word y generated at the current time step according to the decoded feature t ; the input of the LSTM is calculated from the hidden layer feature h t-1 output at the previous time step t - 1 t-1 , the word y and the video features; the calculation formula for the input part of the LSTM is as follows: where w t is the text feature calculated from the hidden layer feature, word, and verb feature obtained from the previous time step, where α1 and α2 are two values in α; 3-4. Concatenate the text features with the global features and encoder features to form the input of the LSTM. The calculation formula of the LSTM part is as follows: {o t ,h t} = LSTM([w t ; V g ; V enc , h t-1 ) (Equation 13) Obtain the decoded feature {o t , h t}. After that, use a mapping matrix to map its dimensionality to the size of the descriptor embedding, and after normalizing it, obtain the probability distribution of the generated word at the current time step. The calculation formula is as follows: where is the probability distribution for generating a word at the current time step, is the mapping matrix; finally, after obtaining the probability distribution, the word with the highest probability is selected from it as the word output y at the current time step t , and the calculation formula is as follows:

3. The video description generation method based on transitive visual relationship detection according to claim 2, wherein The specific implementation of step (4) is as follows: 4-1. Input the verb probability distribution generated in step (2) and the true verbs extracted in step (1) into the negative log-likelihood loss function to obtain a loss value The specific formula is as follows: 4-2. Input the probability distribution generated in step (3) into the true description generated in step (1) and input it into the negative log-likelihood loss function to obtain the loss value The specific formula is as follows: 4-3. According to the calculated loss value and calculate the final loss, and the calculation formula is as follows: where λ is a hyperparameter set in advance; finally, use the backpropagation algorithm to adjust the parameters in the network.