Identity perception basketball video subtitle generation method taking player as center
Through deep learning technology, combined with player identification network and large language model, the precise description of player identity and fine-grained actions in video subtitles is achieved, and the problem of lack of accuracy and semantic richness of subtitles in the prior art is solved.
Patent Information
- Application Number
- CN202510051388.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video subtitle generation methods are difficult to effectively capture the identity and fine-grained behavioral characteristics of a specific player, resulting in a lack of precision and semantic richness of generated subtitles.
A deep learning-based method is adopted, and a trained player identification network is used to extract the player's visual characteristics and identity information from the video, and a two-way semantic interaction between video features and player characteristics is achieved through a cross-attention mechanism. A large language model is combined to generate text containing player identity and fine-grained action descriptions.
It realizes direct recognition of player identities from a visual perspective, generates more accurate and semantic video subtitles, improving the accuracy and fine-grained description capabilities of subtitles.
Smart Images

Figure CN119990312A_ABST
Abstract
Description
Technical Field
[0001] Based on deep learning technology, a player-centric identity-aware subtitle generation method for basketball videos is studied. Background Art
[0002] With the rapid development of artificial intelligence technology, video subtitle generation has gradually become an important research direction across the fields of computer vision and natural language processing. Sports video subtitle generation methods aim to automatically generate accurate and detailed text descriptions by analyzing video content to help viewers better understand the game process. As a global sport, basketball video content often contains complex multimodal information, such as scene dynamics, player movements, game events, etc. In basketball games, player identity and action information play a key role in accurately describing game events. However, traditional video subtitle generation methods mainly analyze the overall features of the video and generate broad text descriptions, which makes it difficult to effectively capture the identity and fine-grained behavioral characteristics of specific players, resulting in the lack of accuracy and semantic richness of the generated subtitles.
[0003] Identity-aware sports video caption generation is an important multimodal computer vision research topic. Current research can be divided into anonymous methods and external knowledge-based methods. These methods have their own limitations. Among them, anonymous methods replace entities in descriptions with special characters, such as [PLAYER] and [TEAM]. The text generated by this method can provide fine-grained actions and distinguish different entity types. However, the text descriptions of this type of method lack entity identities, especially the names of players. External knowledge-based methods generate text descriptions containing player identities and fine-grained actions by introducing additional information (such as game news, player statistics, player lists, etc.). However, since this additional information is independent of the video clips, the player identities in the generated text descriptions mainly rely on static external knowledge rather than being directly extracted from the video clips. This means that the models of this type of method do not actually realize the recognition of the player identities in the video, but infer the player identities through external knowledge. However, this inference-based approach is usually less accurate and difficult to fully reflect the details and dynamic changes in the actual video content. Therefore, directly identifying the player identities from a visual perspective is crucial for player identity-aware video captions. Summary of the invention
[0004] In order to effectively address the limitations of existing anonymous methods and external knowledge-based methods, a deep learning-based player-centric identity-aware basketball video subtitle generation method is proposed. The method first uses a trained player recognition network to extract the player's visual features and identity information from the video. Then, combined with the powerful generation and reasoning capabilities of a large language model, text containing player identities and fine-grained action descriptions is generated.
[0005] Data from 40 games were collected from professional basketball websites, including game event descriptions and corresponding video clips. The key player coordinate boxes of the video clips were manually annotated, and the players were cropped out of the video clips to obtain player sequences. Each player was associated with its corresponding sequence clip and organized into a player-centered sequence clip set. Based on this set, a player recognition network was trained to extract the visual features and identity information of the players from the video clips, thereby providing accurate player recognition support for subsequent caption generation.
[0006] Based on the cross-attention mechanism, bidirectional semantic interaction between video features and player features is achieved, so that the two features can enhance each other. Under this mechanism, the video content and player information are aligned with each other, ensuring that the dynamic scenes in the video are fully integrated with the identity and action features of the players. In addition, a learnable method is adopted to adaptively learn visual context information using randomly initialized query vectors, which are then concatenated with video features and player multimodal features as prompts for the large language model to guide it to generate text descriptions with player identities and fine-grained actions. In order to evaluate the proposed method, an identity-aware basketball video captioning dataset NBA-Identity is proposed, which involves 9726 videos, 321 players with bounding boxes, and 9 basketball events.
[0007] Identifying players in video clips from a visual perspective has been well applied in the task of generating subtitles for basketball videos with player identity awareness, and has a good effect on improving the performance of subsequent tasks. The specific steps are as follows:
[0008] 1) Player Identification Network
[0009] According to the player names in the basketball game event description, the player coordinate frame is annotated for each video clip. According to the annotated player coordinate frame, the player in each frame is cropped to obtain the player sequence of the video clip. Each player is associated with all his player sequences and organized into a player-centered player sequence set. The player recognition network consists of a linear tile mapping layer, a TimeSformer visual backbone network, and a classification head based on a fully connected layer. Using the TimeSformer visual backbone network M time Extract player sequence P B ={p1, p2, ..., p n Visual features of where p n represents the nth frame of the player sequence, Represents the feature dimension.
[0010] o P =W f (F P )=Wf (M time (P B )), (1)
[0011] in, is a classification head based on a fully connected layer, D C Indicates the number of player names, which is 321 in this invention. Using cross entropy loss Train a player recognition network.
[0012]
[0013] Among them, N b Indicates the training batch size, the value is 128. m,i and p m,i is the true label and predicted probability of sample m and category i. When the total number of training rounds (epochs) reaches 50, the training is stopped.
[0014] 2) Bidirectional semantic interaction between video and player features
[0015] The video features and player visual features are first enhanced through the self-attention mechanism, and then the cross-attention mechanism is used to achieve information exchange and semantic interaction. This design effectively integrates the video content and player features, and enhances the model's understanding of the video content and its feature expression capabilities. The self-attention mechanism MSA(·) is defined by the following formula:
[0016]
[0017] Among them, W Q , W K , W V Denote the randomly initialized query matrix, key matrix and value matrix respectively. T is the matrix transposition operator. I is the input vector, which is the player visual features or video features in this invention. D is the dimension of the input vector I. δ(·) is the Softmax normalization function. The cross attention mechanism MCA(·) is a variant of the self-attention mechanism. The formula is defined as follows:
[0018]
[0019] Wherein, I1 and I2 are input vectors. In the present invention, when I1 is a player feature, I2 is a video feature; when I1 is a video feature, I2 is a player feature.
[0020] Concatenate the global features of k player sequences into player features Player Characteristics F k and video features V v The features are self-enhanced through the self-attention mechanism.
[0021]
[0022] Among them, MDA1(·) and MDA2(·) represent the self-attention modules, and the parameters of these two modules are initialized with different specific values and distributions. In addition, l1(·) and l2(·) represent layer normalization functions, and their parameter initialization values are also different. Subsequently, the self-enhanced video feature V′ v and player characteristics F′ k Semantic information interaction through cross-attention mechanism.
[0023]
[0024] Among them, W u1 and W u2 Two linearly increasing matrices with different parameter initialization values. d1 and W d2 represents two linear dimension reduction matrices with different parameter initialization values. l3(·) and l4(·) represent two layer normalization functions with different parameter initialization values. MCA ev (·) and MCA ve (·) respectively represent the entity-video multi-head cross attention mechanism and the video-entity multi-head cross attention mechanism. V″ v and F″ k Represents mutually enhanced video features and player features. Then, the final semantically interactive video feature V is obtained through a multi-layer perceptron. bsi and player characteristics F bso .
[0025]
[0026] Among them, l5(·) and l6(·) represent two layer normalization functions with different parameter initialization values. and Denotes two DropOut layers with different parameter initialization values. φ1(·) and φ2(·) denote two GELU activation functions with different parameter initialization values. MLP1(·) and MLP2(·) denote two multilayer perceptrons with different parameter initialization values.
[0027] 3) Learning visual context information
[0028] Randomly initialize 32 learnable query vectors A kind of global semantic information is extracted from the original video features through the learnable query vector, which helps the model capture the overall dynamic content of the video. The learnable vector not only provides contextual information, but also acts as a bridge between the visual space and the text space. These vectors can be automatically adjusted during the training process to better adapt to the semantic mapping between vision and text. The self-attention and cross-attention mechanisms are used to aggregate the video context information, as shown in the following formula:
[0029]
[0030] V pv =τ p +V v , (9)
[0031]
[0032] V c =FFN(MLP(V cross )), (11)
[0033] Among them, is the video feature extracted by the visual encoder TimeSformer. represents the number of video frames. τ p Represents position embedding, helping the model distinguish frame sequences and capture temporal dependencies between frames. Key Matrix Value Matrix Query Matrix Key Matrix Sum Matrix are random parameters to initialize different matrices. MLP(·) represents a multilayer perceptron. FFN(·) represents a feedforward neural network.
[0034] 4) Decoder based on large language model
[0035] First, use the pre-trained large language model (GPT2, Llama3.2 and Qwen2.5) to extract the text features of each player’s name in the player sequence Then concatenate the k player name features into player text features Among them, D llm is the hidden layer dimension of the large language model. When the large language model is GPT2, the hidden layer dimension D llm The value is 768; when the large language model is Llama3.2-1B, the hidden layer dimension D llm The value is 2048; when the large language model is Llama3.2-3B, the hidden layer dimension D llm The value is 3072; when the large language model is Qwen2.5-0.5B, the hidden layer dimension D llmThe value is 896; when the large language model is Qwen2.5-1.5B, the hidden layer dimension D llm The value is 1536; when the large language model is Qwen2.5-3B, the hidden layer dimension D llm The value is 2048. To enhance the generalization and adaptability of the model, a multimodal cue is provided to the large language model (LLM), including visual context information V c 、Video Features V bsi , Player visual characteristics F k and player text feature E k This multimodal cue guides the model to generate text descriptions containing player identities and fine-grained actions.
[0036]
[0037] Among them, the visual context information mapping matrix Video feature map matrix Player visual feature mapping matrix and the player text feature mapping matrix It is a linear mapping layer with different parameters that can map visual features to the vector space of the language model. llm (·) indicates a decoder based on a large language model. [·] indicates a concatenation operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart.
[0039] Figure 2 is a schematic diagram of the player identification network.
[0040] Figure 3 It is a schematic diagram of two-way semantic interaction. DETAILED DESCRIPTION
[0041] First, 40 game data were collected from professional basketball websites, including text descriptions of game events and game videos. Through manual annotation, the coordinate boxes of key players and the text description of the event were annotated for each video clip of the basketball event. According to the annotated player coordinate boxes, the players in the video clips were cropped to obtain player sequences. Each player was associated with all the corresponding sequence clips and organized into a player-centered sequence clip set. Based on this set, a player identity recognition network was trained to extract the player's visual features and identity information. On this basis, a cross-attention mechanism was used to achieve bidirectional semantic interaction between video features and player features, so that the two features enhanced each other. At the same time, a learnable way was adopted to adaptively learn the visual context information of the video using randomly initialized query vectors. Finally, the context information was spliced with the video features and the multimodal features of the players as input prompts for the large language model, thereby guiding the large language model to generate text descriptions containing player identity information. To evaluate the effectiveness of the proposed method, a player identity-aware basketball video subtitle dataset NBA-Identit was constructed based on the 40 games collected. y The dataset contains 9726 videos, 321 players with bounding boxes and 9 basketball events. The present invention belongs to the field of multimodal computer vision, and specifically relates to technologies such as deep learning, multimodal learning, video understanding, and video subtitle generation.
[0042] A deep learning-based player-centric identity-aware basketball video subtitle generation method is proposed. The specific implementation steps of the invention are as follows:
[0043] Step 1: Based on the player names in the basketball game event description, the player coordinate boxes are annotated for each video clip. Based on the annotated player coordinate boxes, the players in each frame are cropped to obtain the player sequence of the video clip. Each player is associated with all his player sequences and organized into a player-centered player sequence set. The player recognition network consists of a linear tile mapping layer, a TimeSformer visual backbone network, and a classification head based on a fully connected layer. Using the TimeSformer visual backbone network M time Extract player sequence P B ={p1, p2, ..., p n Visual features of where p n represents the nth frame of the player sequence, Represents the feature dimension.
[0044] o P =W f (F P )=W f (M time (P B )), (1)
[0045] in, is a classification head based on a fully connected layer, D C Indicates the number of player names, which is 321 in this invention. Using cross entropy loss Train a player recognition network.
[0046]
[0047] Among them, N b Indicates the training batch size, the value is 128. m,i and p m,i is the true label and predicted probability of sample m and category i. When the total number of training rounds (epochs) reaches 50, the training is stopped.
[0048] Step 2: The video features and player visual features are first enhanced through the self-attention mechanism, and then the cross-attention mechanism is used to achieve information exchange and semantic interaction. This design effectively integrates the video content and player features, and enhances the model's understanding of the video content and feature expression capabilities. The self-attention mechanism MSA(·) is defined by the following formula:
[0049]
[0050] Among them, W Q , W K , W V Denote the randomly initialized query matrix, key matrix and value matrix respectively. T is the matrix transposition operator. I is the input vector, which is the player visual features or video features in this invention. D is the dimension of the input vector I. δ(·) is the Softmax normalization function. The cross attention mechanism MCA(·) is a variant of the self-attention mechanism. The formula is defined as follows:
[0051]
[0052] Wherein, I1 and I2 are input vectors. In the present invention, when I1 is a player feature, I2 is a video feature; when I1 is a video feature, I2 is a player feature.
[0053] Concatenate the global features of k player sequences into player features Player Characteristics F k and video features V v The features are self-enhanced through the self-attention mechanism.
[0054]
[0055] Among them, MDA1(·) and MDA2(·) represent the self-attention modules, and the parameters of these two modules are initialized with different specific values and distributions. In addition, l1(·) and l2(·) represent layer normalization functions, and their parameter initialization values are also different. Subsequently, the self-enhanced video feature V′ v and player characteristics F′ k Semantic information interaction through cross-attention mechanism.
[0056]
[0057] Among them, W u1 and W u2 Two linearly increasing matrices with different parameter initialization values. d1 and W d2 represents two linear dimension reduction matrices with different parameter initialization values. l3(·) and l4(·) represent two layer normalization functions with different parameter initialization values. MCA ev (·) and MCA ve (·) respectively represent the entity-video multi-head cross attention mechanism and the video-entity multi-head cross attention mechanism. V″ v and F″ k Represents mutually enhanced video features and player features. Then, the final semantically interactive video feature V is obtained through a multi-layer perceptron. bsi and player characteristics F bsi .
[0058]
[0059] Among them, l5(·) and l6(·) represent two layer normalization functions with different parameter initialization values. and Denotes two DropOut layers with different parameter initialization values. φ1(·) and φ2(·) denote two GELU activation functions with different parameter initialization values. MLP1(·) and MLP2(·) denote two multilayer perceptrons with different parameter initialization values.
[0060] Step 3: Randomly initialize 32 learnable query vectors A kind of global semantic information is extracted from the original video features through the learnable query vector, which helps the model capture the overall dynamic content of the video. The learnable vector not only provides contextual information, but also acts as a bridge between the visual space and the text space. These vectors can be automatically adjusted during the training process to better adapt to the semantic mapping between vision and text. The self-attention and cross-attention mechanisms are used to aggregate the video context information, as shown in the following formula:
[0061]
[0062] V pv =τ p +V v , (9)
[0063]
[0064] V c =FFN(MLP(V cross )), (11)
[0065] Where, is the video feature extracted by the visual encoder TimeSformer. represents the number of video frames. τ p Represents position embedding, helping the model distinguish frame sequences and capture temporal dependencies between frames. Key Matrix Value Matrix Query Matrix Key Matrix Sum Matrix are random parameters to initialize different matrices. MLP(·) represents a multilayer perceptron. FFN(·) represents a feedforward neural network.
[0066] Step 4: First, use the pre-trained large language model (GPT2, Llama3.2 and Qwen2.5) to extract the text features of each player’s name in the player sequence Then concatenate the k player name features into player text features Among them, D llm is the hidden layer dimension of the large language model. When the large language model is GPT2, the hidden layer dimension D llm The value is 768; when the large language model is Llama3.2-1B, the hidden layer dimension D llm The value is 2048; when the large language model is Llama3.2-3B, the hidden layer dimension D llm The value is 3072; when the large language model is Qwen2.5-0.5B, the hidden layer dimension D llm The value is 896; when the large language model is Qwen2.5-1.5B, the hidden layer dimension D llm The value is 1536; when the large language model is Qwen2.5-3B, the hidden layer dimension D llm The value is 2048. To enhance the generalization and adaptability of the model, a multimodal cue is provided to the large language model (LLM), including visual context information V c , Video Features V bsi , Player visual characteristics F k and player text feature E k This multimodal cue guides the model to generate text descriptions containing player identities and fine-grained actions.
[0067]
[0068] Among them, the visual context information mapping matrix Video feature map matrix Player visual feature mapping matrix and the player text feature mapping matrix It is a linear mapping layer with different parameters that can map visual features to the vector space of the language model. llm (·) indicates a decoder based on a large language model. [·] indicates a concatenation operation.
[0069] In order to verify the effectiveness of the proposed method, a performance comparison experiment was conducted on the newly proposed entity-aware sports dataset NBA-Identity. As shown in Table 1, it achieved better performance comparison results than the current best method, namely the event knowledge guidance method "A Simple yet Effective Knowledge Guided Method for Entity-aware Video Captioning" (KEANet) proposed by Wu Lifang's team and the universal video understanding model "OmniViD: AGenerative Framework for Universal Video Understanding" proposed by Wu Zuxuan's team.
[0070] Table 1: Performance comparison on NBA-Identity dataset
[0071]
[0072]
Claims
1. A player-centric identity-aware basketball video subtitle generation method, characterized by: Step (1) pre-training a player recognition network to extract player visual features and identity information; Step (2) uses self-attention mechanism and cross-attention mechanism to enhance the video features and player visual features, and link the players with the video content; Step (3) uses the learnable query vector to adaptively extract global semantic information to help the model capture the overall dynamic content of the video; Step (4) leverages the powerful generation and reasoning capabilities of the large language model to generate text descriptions with player identities and fine-grained actions.
2. The method according to claim 1, characterized in that In step (1), the player coordinate frame is annotated for each video clip according to the player name in the basketball game event description; the player in each frame is cropped out according to the annotated player coordinate frame to obtain the player sequence of the video clip; each player is associated with all his player sequences and organized into a player sequence set centered on the player; the player recognition network consists of a linear tile mapping layer, a TimeSformer visual backbone network and a classification head based on a fully connected layer; the TimeSformer visual backbone network M is used to identify the player. time Extract player sequence P B ={p1,p2,…,p n Visual features of where p n represents the nth frame of the player sequence, Represents feature dimension; o P =W f (F P )=W F (M time (P B )), (1) in, is a classification head based on a fully connected layer, D C Indicates the number of player names, the value is 321; using cross entropy loss Training a player identification network; Among them, N b Indicates the training batch size, the value is 128; y m,i and p m,i are the true labels and predicted probabilities of sample m and category i; training is stopped when the total number of training epochs reaches 50 or more.
3. The method according to claim 1, characterized in that The video features and player visual features are first enhanced through the self-attention mechanism, and then the cross-attention mechanism is used to achieve information exchange and semantic interaction. This design effectively integrates the video content and player features, and enhances the model's understanding of the video content and feature expression capabilities. The self-attention mechanism MSA(·) is defined by the following formula: Among them, W Q ,W K , W V represents the randomly initialized learnable query matrix, key matrix and value matrix respectively; T is the matrix transposition operator; I is the input vector, which is the player visual features or video features in this invention; D is the dimension of the input vector I; δ(·) is the Softmax normalization function; the cross attention mechanism MCA(·) is a variant of the self-attention mechanism; the formula is defined as follows: Among them, I1 and I2 are input vectors; when I1 is a player feature, I2 is a video feature; when I1 is a video feature, I2 is a player feature; Concatenate the global features of l player sequences into player features Player Characteristics F k and video features V v The features are enhanced through the self-attention mechanism respectively; Among them, MSA1(·) and MSA2(·) represent the self-attention mechanism, and the parameters of these two modules are initialized with different specific values and distributions; in addition, and Representation layer normalization function, their parameter initialization values are also different; then, the self-enhanced video feature V′ v and player characteristics F′ k Interact semantic information through cross-attention mechanism; Among them, W u1 and W u2 Represents two linearly increasing matrices with different parameter initialization values; W d1 and W d2 Represents two linear dimensionality reduction matrices with different parameter initialization values; and Represents two layer normalization functions with different parameter initialization values; MCA ev (·) and MCA ve (·) denotes the entity-video multi-head cross attention mechanism and the video-entity multi-head cross attention mechanism respectively; V v ″ and F k ″ represents the mutually enhanced video features and player features; then, the final semantically interactive video features V are obtained through a multi-layer perceptron. bsi and player characteristics F bsi ; in, and Represents two layer normalization functions with different parameter initialization values; and represents two DropOut layers with different parameter initialization values; φ1(·) and φ2(·) represent two GELU activation functions with different parameter initialization values; MLP1(·) and MLP2(·) represent two multilayer perceptrons with different parameter initialization values.
4. The method according to claim 1, characterized in that In step (3), 32 learnable query vectors are randomly initialized The self-attention and cross-attention mechanisms are used to aggregate video context information, as shown in the following formula: V pv =τ p +V v , (9) In c =FFN(MLP(V cross )), (11) Among them, is the video feature extracted by the visual encoder TimeSformer; represents the number of video frames; τ p Represents position embedding, which helps the model distinguish frame sequences and capture temporal dependencies between frames; query matrix Key Matrix Value Matrix Query Matrix Key Matrix Sum Matrix are matrices initialized with random parameters; MLP(·) represents a multilayer perceptron; FFN(·) represents a feedforward neural network.
5. The method according to claim 1, characterized in that In step (4), we first use the pre-trained large language models (GPT2, Llama3.2 and Qwen2.5) to extract the text features of each player’s name in the player sequence. Then concatenate the k player name features into player text features Among them, D llm is the hidden layer dimension of the large language model; when the large language model is GPT2, the hidden layer dimension D llm The value is 768; when the large language model is Llama3.2-1B, the hidden layer dimension D llm The value is 2048; when the large language model is Llama3.2-3B, the hidden layer dimension D llm The value is 3072; when the large language model is Qwen2.5-0.5B, the hidden layer dimension D llm The value is 896; when the large language model is Qwen2.5-1.5B, the hidden layer dimension D llm The value is 1536; when the large language model is Qwen2.5-3B, the hidden layer dimension D llm The value is 2048; to enhance the generalization and adaptability of the model, a multimodal cue is provided to the large language model (LLM), including visual context information V c , Video Features V bsi , Player visual characteristics F k and player text feature E k ; This multimodal cue guides the model to generate text descriptions containing player identities and fine-grained actions Among them, the visual context information mapping matrix Video feature map matrix Player visual feature mapping matrix and the player text feature mapping matrix It is a linear mapping layer with different parameters, which maps the visual features to the vector space of the language model; llm (·) denotes a decoder based on a large language model; [·] denotes a concatenation operation.
Citation Information
Patent Citations
Emotion recognition method based on visual language pre-training and multi-modal collaborative fusion
CN119026071A
Systems and methods for a vision-language pretraining framework
US20240160853A1