Entity-aware sports video caption generation method enhanced by combining explicit and implicit knowledge

By combining explicit and implicit knowledge enhancement, the accuracy issues of player identity and scene category recognition in sports video subtitle generation are solved, achieving more accurate text description and scene understanding.

CN119653175BActive Publication Date: 2025-09-30BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411706995.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-09-30
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Existing technologies have difficulty in accurately identifying player identities and fine-grained scene categories in sports video subtitle generation, and ignore the implicit knowledge within visual data, resulting in scene confusion and inaccurate descriptions.

Method used

By mining the implicit scene knowledge and explicit game knowledge in training videos, combining the attention mechanism with spatiotemporal dynamic information modeling, a method for entity-aware sports video subtitle generation is designed that jointly enhances explicit and implicit knowledge, and uses learnable query vectors and decoders to generate accurate text descriptions.

Benefits of technology

It improves the accuracy and coherence of sports video subtitle generation, can better match scene-related entities, and enhances the model's ability to distinguish and understand visual scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653175B_ABST
    Figure CN119653175B_ABST
Patent Text Reader

Abstract

This method for generating entity-aware sports video subtitles by combining explicit and implicit knowledge enhancement belongs to the field of video analysis and understanding. Existing traditional video subtitle generation methods have difficulty directly generating text descriptions with player identities and specific scene categories based on the visual content of a video. This invention uses the list of players participating in the game as explicit knowledge to help the model acquire pre-match knowledge. It proposes using learnable vectors to adaptively capture scene-related video features under the guidance of knowledge. Based on the attention mechanism, the relationship between players and video content is modeled to generate entity-related video features. A decoder based on the principle of "scene first, entity later" is proposed to enhance the accuracy and coherence of generated subtitles. The spatiotemporal dynamic information of the video is modeled to capture the temporal and spatial relationships between video frames. The effectiveness of this invention is verified on the basketball datasets VC_NBA_2022 and NSVA and the football dataset Goal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on deep learning technology, a method for entity-aware sports video captioning is developed that combines explicit and implicit knowledge enhancement. First, based on match information, all players from each video clip corresponding to the match are selected as a candidate player list, i.e., explicit match knowledge. The visual features of the video data for each scene are averaged according to various scene labels to obtain implicit visual scene knowledge for each scene. Then, a learnable vector is used to adaptively capture scene-related video features guided by this knowledge. Based on an attention mechanism, the relationship between players and video content is modeled to generate entity-related video features. A decoder based on the "scene first, entity later" principle is then proposed to enhance the accuracy and coherence of the generated captions. Finally, the spatiotemporal dynamics of the video are modeled to capture the temporal and spatial relationships between video frames. This method belongs to the field of computer vision and specifically involves technologies such as deep learning, video analysis and understanding, and video captioning. Background Art

[0002] The explosive growth of video data has become a prominent feature of the current information age. To effectively understand and analyze this massive amount of video content and extract useful information, various video content description technologies have emerged. Video captioning refers to the process of automatically analyzing and processing video data using technologies such as computer vision and deep learning to extract key information from the video and then classify, identify, understand, and interpret it. This technology can help people better understand and analyze video content, improving its dissemination effectiveness and viewing experience. Furthermore, video content description tasks have broad applications in fields such as sports, education and training, healthcare, and human-computer interaction. While various video content understanding technologies have made significant progress and performance improvements, live text broadcasts of sports events require a deep semantic understanding of the video content. Commentators not only rely on pre-match information and professional knowledge to accurately identify players and scenes, but also need to incorporate background knowledge to provide viewers with a professional and engaging interpretation of the game. However, the model's inherent parameter knowledge is limited and lacks domain-specific knowledge, which limits its performance. Therefore, knowledge-enhanced video captioning is essential for applications in the sports field.

[0003] Sports video captioning is an important research topic in computer vision. Current research can be divided into traditional methods and methods based on external explicit knowledge. These methods each have their own limitations. Traditional methods aggregate a set of video frame features into a video representation through a visual encoder and then use a language decoder on the video representation to learn visual-text alignment to generate text descriptions. Traditional methods are unable to directly generate text descriptions with player identity information and fine-grained scene categories based on visual content. Methods based on external explicit knowledge enhance the visual content by leveraging visual details detected by external detectors or knowledge acquired from external knowledge bases to help the model generate text descriptions with player identity and fine-grained scene categories. However, the text descriptions produced by these methods are still not accurate. They focus solely on external explicit knowledge and ignore the complex underlying patterns and deeper relationships hidden within the visual data. In fact, different scenes with similar visual appearance are often mistaken for the same, which can lead to scene confusion in the model. Furthermore, different videos within the same scene category may share certain patterns or subtle connections within that scene. This implicit information cannot be obtained from the encoder or external explicit sources. Therefore, mining implicit knowledge from data and designing a framework to effectively integrate explicit and implicit knowledge are crucial for entity-aware sports video captioning. Summary of the Invention

[0004] To effectively address the limitations of existing traditional methods and those based on external explicit knowledge, we propose a deep learning-based entity-aware sports video captioning method that combines explicit and implicit knowledge enhancement. Taking sports video clips as input, we generate text descriptions with player identities and fine-grained scene categories based on explicit and implicit knowledge.

[0005] Based on the principle of analogical reasoning in cognitive psychology, people can compare the source video with other videos of similar visual scenes to achieve a more accurate understanding. Inspired by this, to more accurately generate descriptions of specific visual scenes, implicit visual scene knowledge is mined from training videos. All video features within the same scene category are averaged to obtain visual feature centers. Each center serves as a representative description for each scene category. It is a highly abstract, implicit representation of scene knowledge that defines the specific paradigm of a scene and demarcates its boundaries with other scenes. This knowledge can guide the model to distinguish different visual scenes. In addition, a list of players participating in each game is used as explicit game knowledge to assist the model in generating text descriptions with entity names.

[0006] To effectively leverage and aggregate both types of knowledge, a network that jointly enhances explicit and implicit knowledge is designed for entity-aware sports video captioning. First, a learnable query vector is used to adaptively capture scene-related video features guided by implicit scene knowledge, enhancing the model's depth of scene understanding. An attention mechanism is then employed to deeply model the interactions between players and the scene, generating entity-related video features. To fully leverage both scene-related and entity-related video features, a decoder employing a "scene-to-entity" approach is designed. This decoder first decodes scene-related features and then decodes subtitles containing player names within the scene. This decoder improves the accuracy and coherence of generated captions by providing contextual understanding, enabling them to more accurately match scene-related entities. Finally, spatiotemporal dynamic information is encoded by capturing the temporal and spatial relationships between video frames.

[0007] The combined use of explicit and implicit knowledge has been well applied in the task of generating entity-aware sports video subtitles, and has a good effect on improving the performance of subsequent tasks. The specific steps are as follows:

[0008] 1) Extraction of scene knowledge

[0009] When people watch a video, they usually compare it with other videos with the most similar scenes to gain a more accurate understanding, and use similar videos to compensate for visual details that may be missing in the source video. However, retrieving relevant information from external databases online to enhance video features increases computational costs. This method mines hidden scene knowledge in videos in a simple way. This implicit knowledge not only provides complementary visual information, but also enhances the model's ability to discriminate and understand visual scenes. Specifically, a visual encoder is used to extract global features of all video clips in the training set. N t represents the number of video clips in the training set, Represents the feature dimension. Video features are grouped according to the labels annotated for each video in the training set. The central features of each labeled scene are obtained by averaging the features of videos with the same label.

[0010]

[0011] Among them, V r g is a video collection with label r, and the global feature of the jth video in the collection is represented by v j K is the number of video features of label r. C rThe central feature of each label is obtained. Finally, the central feature set C = {C1, C2, C3, ..., C9} of the 9 labels is obtained. These 9 central features are defined as implicit scene knowledge, which helps the model distinguish different visual scenes and make up for the lack of information in the source video.

[0012] 2) Extraction of competition knowledge

[0013] In live sports broadcasts, commentators are given information about the game in advance, such as the teams and the identities of all players on each team. If the model has such clear game-related knowledge, it will be better able to generate text descriptions with player identities. Therefore, the game name, team, and all players corresponding to each video clip are obtained from the basketball knowledge graph proposed by Xi et al. All players in a game are aggregated to form a candidate player list, so that the model can obtain game-related knowledge. The visual encoder and text encoder are used to extract global image features of all players in the candidate list respectively. and global name text features Among them, N e Represents the number of players in the candidate player list. Add the image features and name text features according to the corresponding positions to obtain the multimodal player features E m .

[0014] E m =E p +E n W1, (2)

[0015] Among them, W1 is a randomly initialized learnable matrix that maps video features into text space.

[0016] 3) Spatiotemporal dynamic information modeling

[0017] Videos not only contain spatial information from static frames, but also embody behaviors and events that evolve over time, which are expressed through spatiotemporal features. Extracting these spatiotemporal dynamics is crucial for accurately recognizing objects, actions, and the overall meaning of scenes in videos.

[0018] First, the visual encoder encodes 18 frames of images into visual features Then convert it into visual features Along the spatial dimension, the video features are averaged to obtain spatial features Similarly, along the temporal dimension, the video features are averaged to obtain the temporal features Combine the spatial features and temporal features to obtain the spatiotemporal features V of a given video st .

[0019]

[0020] Among them, [,] represents the splicing function.

[0021] 4) Video-knowledge information interaction

[0022] Since there is a lot of redundant information in the video, directly interacting with the video knowledge will introduce a lot of noise information. Therefore, a learnable method is used to adaptively capture key information. Specifically, 18 learnable query vectors are first randomly initialized. The learnable query vector then interacts with the scene knowledge C and spatiotemporal video features in turn to obtain scene-related video features. The specific interaction process is shown in the following formula:

[0023]

[0024] in, and is a randomly initialized learnable matrix, and i is the number of attention heads. i ′ is the interactive feature obtained by the attention mechanism between the query vector and the scene knowledge. δ(·) represents the Softmax normalization function. The concatenated output features of i attention heads [V1′, V2′, …, V i ′] through the fully connected layer Perform integration to obtain the intermediate feature V′. i ″ is the feature obtained by the attention mechanism of the intermediate features and spatiotemporal features. The concatenated output features of i attention heads [V1″, V2″, …, V i ″]Through the fully connected layer Integrate to obtain scene-related features V i ″. FFN(·) represents the feed-forward layer. A total of 4 VKIM layers are stacked.

[0025] 5) Entity-Video Information Interaction (EVIM)

[0026] Learning entity-related video features based on attention mechanism The scaled dot product attention function is used to associate entities with the scene, and the residual connection operation is used to enhance the feature representation.

[0027]

[0028] in, and is a randomly initialized learnable matrix, V scene is the scene-related feature V″ obtained by formula (7). m is the multimodal player feature obtained by formula (2), and δ(·) represents the Softmax normalization function.

[0029] 6) Scene-Entity Decoder

[0030] The designed scene-to-entity decoder can first decode scene-related text features, and then decode text features with entity information based on this, making full use of scene and entity information. Scene cross attention is used to generate scene-related features

[0031]

[0032] in, Indicates layer normalization, l represents the current decoding layer, and l-1 represents the previous decoding layer. scene is the scene-related feature V″ obtained by formula (7). t is the current decoding time, 1:t is the time from the first decoding moment to the current decoding moment of the autoregressive decoding operation, is the previous self-attention layer M self (·) Sentence features output at decoding time t. M S-Att (·) represents the scene cross attention layer. The first layer embedding representation As shown in formula (11):

[0033]

[0034] Among them, token ω is the result obtained by the word segmenter after performing word segmentation on the input sentence, ω 0:t-1 represents the token at time t-1. ε(·) represents the token embedding representation extractor. τ pe (·) represents triangular position embedding, which enables the model to distinguish different positions in the sequence and thus better understand the sequence structure.

[0035] In order to obtain entity-aware text descriptions, we use entity cross attention M E-Att (·)Entity information V entity Scene-related features Interact so that entity information can be integrated into text features with scene context. Residual connection operation and feedforward network are used to enhance feature representation and output entity-related features.

[0036]

[0037] in, Denotes layer normalization. FFN(·) denotes the feedforward layer. The final decoder output is obtained by stacking L layers of decoding layers. The value of L is 3. The decoder combines the probability and the word dictionary to decode the corresponding word. The output probability O is obtained through the multi-layer perception layer MLP(·) and the Softmax function δ(·). p .

[0038]

[0039] Compared with the existing technology, it has the following advantages:

[0040] This approach uses a simple and effective method to mine implicit scene knowledge from training video data, helping the model distinguish between different scene types. A framework for entity-aware sports video captioning, augmented by both explicit and implicit knowledge, is designed. A learnable query vector is used to adaptively learn scene-related video features. An attention mechanism is used to associate entities with video content and learn entity-related video features. The feature representation of the video is enhanced by extracting spatiotemporal dynamic features. A "scene-first, entity-later" decoding approach is designed to provide contextual understanding to improve the accuracy and coherence of generated captions, enabling them to more accurately match scene-related entities. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flow chart. DETAILED DESCRIPTION

[0042] This paper proposes a deep learning-based entity-aware sports video subtitle generation method that combines explicit knowledge and implicit knowledge enhancement. The specific implementation steps of this invention are as follows:

[0043] Step 1: In a simple and effective way, the training video features are grouped according to the labels, and the central features of each group of features are averaged to obtain the central features of the scene type, thereby mining the visual scene knowledge hidden in the data. Specifically, the visual encoder is used to extract the global features of all video clips in the training set. N t represents the number of video clips in the training set, Represents the feature dimension. Video features are grouped according to the labels annotated for each video in the training set. The central features of each labeled scene are obtained by averaging the features of videos with the same label.

[0044]

[0045] Among them, V r g is a video collection with label r, and the global feature of the jth video in the collection is represented by v j K is the number of video features of label r. C r The central feature of each label is obtained. Finally, the central feature set C = {C1, C2, C3, ..., C9} of the 9 labels is obtained. These 9 central features are defined as implicit scene knowledge, which helps the model distinguish different visual scenes and make up for the lack of information in the source video.

[0046] Step 2: To mimic the pre-game knowledge of all players on both teams by a commentator, we retrieve the game name, team, and all player information corresponding to each video clip from the basketball knowledge graph proposed by Xi et al. All players in a game are aggregated into a candidate player list, allowing the model to acquire game-related knowledge. We then use a visual encoder and a text encoder to extract global image features for all players in the candidate list. and global name text features Among them, N e Represents the number of players in the candidate player list. Add the image features and name text features according to the corresponding positions to obtain the multimodal player features E m .

[0047] E m =E p +E n W1, (2)

[0048] Among them, W1 is a learnable matrix that maps video features into text space.

[0049] Step 3: Videos contain not only spatial information from static frames, but also behaviors and events that evolve over time, expressed as spatiotemporal features. Extracting these spatiotemporal dynamics is crucial for accurately identifying objects, actions, and the overall meaning of scenes in videos.

[0050] First, the visual encoder encodes 18 frames of images into visual features Then convert it into visual features Along the spatial dimension, the video features are averaged to obtain spatial features Similarly, along the temporal dimension, the video features are averaged to obtain the temporal features Combine the spatial features and temporal features to obtain the spatiotemporal features V of a given video st .

[0051]

[0052] Among them, [,] represents the splicing function.

[0053] Step 4: Since there is a lot of redundant information in the video, directly interacting with the video knowledge will introduce a lot of noise information. Therefore, a learnable method is used to adaptively capture key information. Specifically, first randomly initialize 18 learnable query vectors The learnable query vector then interacts with the scene knowledge C and spatiotemporal video features in turn to obtain scene-related video features. The specific interaction process is shown in the following formula:

[0054]

[0055] V″=FFN(F c 2 (V1″,V2″,…,V i ″)) (7)

[0056] in, and is a randomly initialized learnable matrix, and i is the number of attention heads. i ′ is the interactive feature obtained by the attention mechanism between the query vector and the scene knowledge. δ(·) represents the Softmax normalization function. The concatenated output features of i attention heads [V1′, V2′, …, V i ′] through the fully connected layer F c 1 (·) is integrated to obtain the intermediate feature V′. V i ″ is the feature obtained by the attention mechanism of the intermediate features and spatiotemporal features. The concatenated output features of i attention heads [V1″, V2″, …, V i ″]Through the fully connected layer Integrate to obtain scene-related features V i ″. FFN(·) represents the feed-forward layer. A total of 4 VKIM layers are stacked.

[0057] Step 5: Learn entity-related video features based on attention mechanism The scaled dot product attention function is used to associate entities with the scene, and the residual connection operation is used to enhance the feature representation.

[0058]

[0059] in, and is a randomly initialized learnable matrix, V scene is the scene-related feature V″ obtained by formula (7).

[0060] Step 6: The designed scene-to-entity decoder can first decode the scene-related text features, and then decode the text features with entity information based on this, making full use of the scene and entity information. Scene cross attention is used to generate scene-related features

[0061]

[0062] in, Represents layer normalization, l represents the current decoding layer, l-1 represents the previous decoding layer. t is the current decoding time, 1:t is the time from the first decoding moment to the current decoding moment of the autoregressive decoding operation, Vscene is the scene-related feature V″ obtained by formula (7). is the previous self-attention layer M self (·) Sentence features output at decoding time t. M S-Att (·) represents the scene cross attention layer. The first layer embedding representation As shown in formula (11):

[0063]

[0064] Among them, token ω is the result obtained by the word segmenter after performing word segmentation on the input sentence, ω 0:t-1 represents the token at time t-1, ε(·) represents the token embedding representation extractor, τ pe (·) represents triangular position embedding, which enables the model to distinguish different positions in the sequence and thus better understand the sequence structure.

[0065] In order to obtain entity-aware text descriptions, we use entity cross attention M E-Att (·)Entity information V entity Scene-related features Interact so that entity information can be integrated into text features with scene context. Residual connection operation and feedforward network are used to enhance feature representation and output entity-related features.

[0066]

[0067] in, Denotes layer normalization. FFN(·) denotes the feedforward layer. The final decoder output is obtained by stacking L layers of decoding layers. The value of L is 3. The decoder combines the probability and the word dictionary to decode the corresponding word. The output probability O is obtained through the multi-layer perception layer MLP(·) and the Softmax function δ(·). p .

[0068]

[0069] To verify the effectiveness of the proposed algorithm, performance comparison experiments were conducted on the entity-aware sports datasets VC_NBA_2022, Goal, and NSVA. The algorithm's performance was first tested on the VC_NBA_2022 dataset. As shown in Table 1, it achieved superior performance compared to the current best methods: the event knowledge-guided method "A Simple Yet Effective Knowledge-Guided Method for Entity-aware Video Captioning" proposed by Wu Lifang's team and the universal video understanding model "OmniViD: A Generative Framework for Universal Video Understanding" proposed by Wu Zuxuan's team. Universal model* and universal model** represent explicit knowledge augmentation and combined explicit and implicit knowledge augmentation, respectively. Knowledge-guided* indicates that implicit knowledge is added to the explicit knowledge augmentation.

[0070] Table 1: Performance comparison on the VC_NBA_2022 dataset

[0071] CIDEr METEOR Rouge-L BLEU-1 BLEU-2 BLEU-3 BLEU-4 General Model 71.2 27.5 50.7 49.7 42.4 36.5 29.2 General Model* 125.2 28.6 53.3 52.5 45.2 38.7 30.5 General Model 130.8 29.0 56.0 53.7 46.6 39.7 32.6 Knowledge guidance 138.5 28.0 54.9 53.1 46.4 38.8 32.4 Knowledge Guidance* 144.6 28.7 56.1 54.5 47.6 40.2 34.4 This method 140.7 29.5 56.8 57.1 50.3 43.1 36.7

[0072] Next, on the Goal football commentary dataset, the algorithm was experimentally compared with the best current approaches: the "A Simple Yet Effective Knowledge Guided Method for Entity-aware Video Captioning" method proposed by Wu Lifang's team, and the "OmniViD: A Generative Framework for Universal Video Understanding" model proposed by Wu Zuxuan's team. As shown in Table 2, the algorithm achieved superior performance. The "Universal Model*" and "Universal Model**" respectively represent explicit knowledge augmentation and combined explicit and implicit knowledge augmentation. Knowledge-Guided* refers to augmenting explicit knowledge with implicit knowledge.

[0073] Table 2: Performance comparison on the Goal dataset

[0074] CIDEr METEOR Rouge-L BLEU-1 General Model 3.0 5.9 9.1 10.7 General Model* 3.9 6.4 10.6 14.4 General Model 4.2 6.5 11.0 15.4 Knowledge guidance 3.7 6.4 10.5 14.9 Knowledge Guidance* 4.0 6.6 10.8 15.2 This method 4.1 6.6 11.2 15.3

[0075] Furthermore, using explicit knowledge and implicit knowledge as components, we conducted an entity-aware sports video caption generation task on the basketball dataset NSVA. The effectiveness of dual knowledge enhancement was verified by equipping the sports video analysis model "Sports Video Analysis on Large-Scale Data" proposed by Richard P. Wildes' team with explicit game knowledge and implicit visual scene knowledge. The analysis model has two modes: the first uses the S3D network as the feature extraction network, and the second uses the Timesformer (T) network as the feature extraction network. "(full)" indicates that it comprehensively utilizes video features, basketball features, basket features, player features, and venue features. "list" represents explicit knowledge, and "scene" represents implicit knowledge. "Analysis model*" indicates that the knowledge is equipped. As shown in Table 3, the proposed dual knowledge enhancement improves the performance of existing methods, proving the effectiveness of the proposed dual knowledge enhancement method.

[0076] Table 3: Performance comparison on the NSVA dataset

[0077]

Claims

1. An entity-aware sports video subtitle generation method enhanced by combining explicit knowledge and implicit knowledge, characterized by: Step (1) groups the training video features according to the labels and averages each group of features to obtain the central features of the scene, thereby mining the visual scene knowledge hidden in the data. Step (2) aggregates all players in a game into a candidate player list so that the model can acquire player knowledge related to the game; Step (3) enhances the feature representation of the video by extracting the spatiotemporal dynamic features of the video; Step (4) adaptively learning scene-related video features using the learnable query vector; Step (5) uses the attention mechanism to associate entities with video content and learn video features related to entities; In step (6), a scene-to-entity decoder is designed, which first decodes the scene-related text features and then decodes the text features with entity information based on this.

2. The method according to claim 1, characterized in that In step (1), the visual encoder is used to extract the global features of all video clips in the training set. N t represents the number of video clips in the training set, Represents feature dimension; The video features are grouped according to the labels of each video in the training set; the central features of each label scene are obtained by averaging the video features of the same label; in, is a video collection with label r, and the global feature of the jth video in the collection is represented by v j ; K is the number of video features of label r; C r is the central feature of each label; finally, the central feature set C = {C1, C2, C3, …, C9} of 9 labels is obtained. These 9 central features are defined as implicit scene knowledge.

3. The method according to claim 1, characterized in that In step (2), the game name, team, and all player information corresponding to each video clip are obtained from the basketball knowledge graph; all players in a game are aggregated to form a candidate player list; the global image features of all players in the candidate list are extracted using the visual encoder and the text encoder respectively. and global name text features Among them, N e Represents the number of players in the candidate player list; add the image features and name text features according to the corresponding positions to obtain the multimodal player feature E m ; E m =E p +E n W1, (2) where W1 is a learnable matrix that maps video features into text space.

4. The method according to claim 1, characterized in that In step (3), the visual encoder first encodes the 18 frames of image into visual features Then convert it into visual features Along the spatial dimension, the video features are averaged to obtain spatial features Along the time dimension, the video features are averaged to obtain the time features Combine the spatial features and temporal features to obtain the spatiotemporal features V of a given video st ; Among them, [,] represents the splicing function.

5. The method according to claim 1, characterized in that In step (4), 18 learnable query vectors are first randomly initialized Then the learnable query vector interacts with the scene knowledge C and spatiotemporal video features in turn to obtain scene-related video features; The interaction process is shown in the following formula: in, and is a randomly initialized learnable matrix, i is the number of attention heads; V i ′ is the interactive feature obtained by the attention mechanism between the query vector and the scene knowledge; δ(·) represents the Softmax normalization function; the concatenated output features of i attention heads [V′1, V′2, …, V′ i ]Through the fully connected layer Integrate to obtain the intermediate feature V′; V i ″ is the feature obtained by the attention mechanism of the intermediate features and spatiotemporal features; the concatenated output features of i attention heads [V″1, V″2, …, V″ i ]Through the fully connected layer Integrate to obtain scene-related features V″ i ; FFN(·) represents the feedforward layer; a total of 4 VKIM layers are stacked.

6. The method according to claim 1, characterized in that In step (5), the video features related to entities are learned based on the attention mechanism Utilize the scaled dot product attention function to associate entities with scenes and utilize the residual connection operation to enhance feature representation; in, and is a randomly initialized learnable matrix, V scene is the scene-related feature V″ obtained by formula (7).

7. The method according to claim 1, characterized in that In step (6), the designed scene-to-entity decoder first decodes the scene-related text features, and then decodes the text features with entity information based on this; Scene cross attention is used to generate scene-related features in, Represents layer normalization, l represents the current decoding layer, l-1 represents the previous decoding layer; t is the current decoding time, 1:t is the time from the first decoding moment to the current decoding moment of the autoregressive decoding operation, V scene is the scene-related feature V″ obtained by formula (7); is the previous self-attention layer M self (·) Sentence features output at decoding time t; M S-Att (·) represents the scene cross attention layer; the first layer embedding representation As shown in formula (11): Among them, token ω is the result obtained by the word segmenter after performing word segmentation on the input sentence, ω 0:t-1 represents the token at time t-1, ε(·) represents the token embedding representation extractor, τ pe (·) represents triangular position embedding, which enables the model to distinguish different positions in the sequence and thus better understand the sequence structure; Leveraging entity cross attention M E-Att (·)Entity information V entity Scene-related features Interact so that entity information can be integrated into text features with scene context; use residual connection operations and feedforward networks to enhance feature representation and output entity-related features in, Represents layer normalization; FFN(·) represents the feedforward layer; the final decoder output is obtained by stacking L layers of decoding layers The value of L is 3; the decoder combines the probability and the word dictionary to decode the corresponding word; the output probability O is obtained through the multi-layer perception layer MLP(·) and the Softmax function δ(·) p ;