An interactive perception multi-person sports competition understanding analysis method and system
Patent Information
- Application Number
- CN202610730456.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,现有人体运动理解方法大多针对单人场景设计,对于真实场景中普遍存在的多人交互行为缺乏有效建模能力
[0036]1、本发明通过设置交互感知筛选模块和社会—时序联合建模模块,解决了现有方法仅对多个人体特征进行简单拼接、难以显式建模人物之间交互关系的问题,达到了提升多人体场景下交互理解能力、角色关系分析能力及因果推理能力的效果。
Smart Images

Figure CN122842006A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an interactive perception method and system for understanding and analyzing multi-player sports competitions, belonging to the field of sports competition movement analysis technology, and applied to the scenario of correcting the tactical movement posture of sports competition groups. Background Technology
[0002] In recent years, with the development of human pose estimation, video understanding, and large language model technologies, motion understanding methods based on human keypoints, skeleton sequences, or video sequences have been extensively studied. Existing technologies typically encode single-person motion sequences and then combine them with text generation, action description, behavior recognition, or question-answering reasoning to achieve semantic understanding of human behavior. These methods have demonstrated application value in fields such as motion analysis, digital human interaction, training assistance, and rehabilitation assessment.
[0003] However, most existing methods for human motion understanding are designed for single-person scenarios and lack effective modeling capabilities for multi-person interactions prevalent in real-world scenarios. Especially in social interactions, group collaborations, and competitive sports, simply encoding individual motion features and then directly concatenating them fails to effectively capture temporal dependencies, role relationships, and causal interactions between individuals. Furthermore, as the number of participants increases, direct concatenation leads to a rapid increase in sequence length, resulting in high computational complexity, increased inference latency, and insufficient scalability for large-scale scenarios. In addition, existing public datasets are mostly focused on single-person or two-person interactions, making it difficult to support unified training and evaluation in complex multi-person scenarios.
[0004] Therefore, in multiplayer sports competition scenarios, how to improve the ability to understand and reason about motion under the interaction of the character and time dimensions has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this invention is to address the technical problem of improving motion understanding and reasoning capabilities in multi-person sports competition scenarios involving interactions between the human and temporal dimensions. It proposes an interactive perception-based method and system for motion understanding and analysis in multi-person sports competitions. This method achieves behavioral understanding and semantic output in multi-person scenarios by constructing human body encoding, interactive perception filtering, social-temporal joint modeling, and cross-modal alignment generation mechanisms. This improves the accuracy, coherence, and real-time processing capabilities of understanding complex group scenarios.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] This invention discloses an interactive, multi-player sports competition understanding and analysis method, comprising the following steps:
[0008] Step 1: From the scene to be analyzed, the human motion sequence or the human motion sequence extracted from the video sequence by the motion estimation module is flipped according to the human dimension and the time dimension to form a multi-human motion sequence tensor;
[0009] Step 2: Use cross attention to align and fuse the video features and motion features extracted by the Video-LLaVA video encoder and LLaMo motion encoder to form a multi-person initial representation;
[0010] Step 2.1: For the multi-person motion sequence tensor, the Video-LLaVA video encoder and the LLaMo motion encoder are used to extract the video features and motion features of the multi-person motion sequence tensor respectively;
[0011] Step 2.2: Utilize the cross-modal alignment module to perform feature alignment and fusion of video features and motion features of multiple people using a cross-attention mechanism to form an initial representation of the multiple people;
[0012] Step 3: The importance weights of language perception obtained by pooling the initial multi-person representations and embedding the text through self-attention calculation are injected into the initial multi-person representations using Hamada product broadcasting to obtain language-guided interactive perception features. The compressed interactive perception representation is then formed by pooling the person dimension and filtering with Top-K selection.
[0013] Step 3.1: Perform average pooling on the initial representations of multiple users according to the time dimension to obtain global temporal features;
[0014] Step 3.2: Use the BERT text encoder to extract features from the text in the scene to be analyzed, forming text embeddings;
[0015] Step 3.3: The importance weights of language perception obtained by calculating text embedding and global temporal features through self-attention mechanism are injected into the initial representation of multiple people using Hamada product broadcasting to form language-guided interactive perception features;
[0016] Step 3.3.1: Pool the global temporal features according to the time dimension, and use the self-attention mechanism to calculate the importance weights of the text embedding and global temporal features to obtain language perception.
[0017] Step 3.3.2: Use Hamada product to broadcast the importance weights of language perception into the initial representations of multiple users, forming language-guided interactive perception features;
[0018] Step 3.4: The importance scores obtained from the pooled language-guided interactive perception features of the character dimension using the self-attention mechanism are filtered using the Top-K selection method to form a compressed interactive perception representation;
[0019] Step 3.4.1: Pool the language-guided interactive perception features according to the character dimension, and use the self-attention mechanism to calculate the importance score of the language-guided interactive perception features;
[0020] Step 3.4.2: Set a filtering threshold K, use the Top-K selection method to filter the top K keyframes in the importance scores of the language-guided interactive perception features, and splice the selected keyframes together to form a compressed interactive perception representation.
[0021] Step 4: Adaptively gating the attention branch features, global branch features, and individual branch features constructed for the current character through multilayer perceptron and normalized probability distribution, and perform bidirectional cross-attention fusion with the social interaction representation obtained through self-attention calculation on the character dimension to obtain a multi-character fused representation;
[0022] Step 4.1: In terms of temporal sequence, construct attention branches, individual branches, and global branches for the motion changes of the current character in consecutive frames; further, attention branches are used to locate the interaction objects that the current character depends on and model their temporal relationship, individual branches are used to model the temporal continuity of the current character itself, and global branches are used to introduce scene-level global context.
[0023] Step 4.2: Perform self-attention calculation on the current person's features in the time dimension to obtain individual branch features; perform cross-attention calculation on the current person's features and the global features of other people to obtain global branch features; use the cross-individual importance weights obtained by cross-attention calculation on the current person's features and the global features pooled in the time dimension to filter the individuals to be followed, and perform cross-attention calculation on the current person's features and the features of the individuals to be followed to obtain the attention branch features;
[0024] Step 4.3: The features of the focus branch, global branch, and individual branch are passed through a multilayer perceptron and then subjected to normalized probability distribution and adaptive gating fusion with the corresponding branch features to obtain a temporal representation;
[0025] Step 4.4: Perform self-attention calculation on the character dimension according to the spatial relationships and social interaction dependencies between multiple characters based on time frames to obtain the social interaction representation;
[0026] Step 4.5: Perform bidirectional cross-attention fusion of temporal representation and social interaction representation to obtain multi-person fused representation;
[0027] Step 5: By aligning the text embedding and multi-character fusion representation through bidirectional cross-attention, we obtain the alignment features of language generation, and output text descriptions applicable to one or more of the following in multi-player sports competition scenarios: behavior description, interaction analysis, question and answer results, tactical analysis, commentary text, or character relationship judgment.
[0028] This invention discloses an interactive perception-based multi-player sports competition understanding and analysis system for implementing the aforementioned method. The system comprises an input acquisition module, a human body encoding module, an interactive perception filtering module, a social-temporal joint modeling module, a cross-modal alignment module, and a semantic generation module.
[0029] The input acquisition module is used to acquire video sequences and / or human motion sequences of the scene to be analyzed, and to construct multi-human input data; this data will be used as input to the human coding module.
[0030] The human body coding module is used to extract and fuse video features and motion features to obtain initial representations of multiple human bodies; it will serve as input to the interactive perception filtering module.
[0031] The interaction-aware filtering module is used to calculate the importance of each time frame in a multi-person sequence and filter key frames to obtain a compressed interaction-aware representation; this will serve as the input to the social temporal joint modeling module.
[0032] The social-temporal joint modeling module is used to model the temporal dynamics of an individual and the social interaction relationships between individuals, and outputs a multi-person fusion representation; it will serve as the input to the cross-modal alignment module.
[0033] A cross-modal alignment module is used to align multi-person fusion representations with language embeddings; it will serve as input to the semantic generation module.
[0034] The semantic generation module is used to generate behavioral descriptions, interaction understanding results, or analysis text for multi-human scenes based on alignment features;
[0035] Compared with existing technologies, it has the following beneficial effects:
[0036] 1. This invention solves the problem that existing methods simply splice together multiple human features and have difficulty explicitly modeling the interaction relationships between characters by setting up an interaction perception filtering module and a social-temporal joint modeling module. This achieves the effect of improving the interaction understanding ability, role relationship analysis ability and causal reasoning ability in multi-human scenes.
[0037] 2. By introducing attention branches, individual branches, and global branches simultaneously in temporal modeling, this invention solves the problem that existing technologies struggle to simultaneously consider the dynamic continuity of individuals, local interaction relationships, and global scene context, thereby improving the accuracy of behavior description and semantic coherence.
[0038] 3. This invention solves the problems of excessively long input sequences, excessive inference computation, and poor real-time performance caused by the increase in the number of users through keyframe filtering and fusion marking mechanisms, and achieves the effect of maintaining high inference efficiency and good scalability in complex large-scale multi-user scenarios.
[0039] 4. This invention supports dual-modal input of video and human motion, and solves the problems of insufficient utilization of multi-source information and limited output semantic analysis capabilities by combining a cross-modal alignment module with a large language model. It achieves the effect of being able to adapt to various application tasks such as group behavior analysis, sports tactical analysis, intelligent question answering, and automatic commentary. Attached Figure Description
[0040] Figure 1 is a schematic diagram of the method flow of the present invention;
[0041] Figure 2 This is a schematic diagram of the system flow of the present invention; Detailed Implementation
[0042] To better illustrate the purpose and advantages of this invention, the invention will be further described below with reference to the accompanying drawings and examples. It should be noted that the implementation of this invention is not limited to the following embodiments, and any modifications or alterations made to this invention will fall within the scope of protection of this invention.
[0043] Example
[0044] like Figure 1 As shown in the figure, the specific implementation steps of the interactive perception multi-player sports competition understanding and analysis method in this embodiment are as follows:
[0045] Step 1: From the scene to be analyzed, the human motion sequence or the human motion sequence extracted from the video sequence by the motion estimation module is flipped according to the human dimension and the time dimension to form a multi-human motion sequence tensor;
[0046] Step 2: Use cross attention to align and fuse the video features and motion features extracted by the Video-LLaVA video encoder and LLaMo motion encoder to form a multi-person initial representation;
[0047] Step 2.1: For the multi-person motion sequence tensor, the Video-LLaVA video encoder and the LLaMo motion encoder are used to extract the video features and motion features of the multi-person motion sequence tensor respectively;
[0048] Step 2.2: Utilize the cross-modal alignment module to perform feature alignment and fusion of video features and motion features of multiple people using a cross-attention mechanism to form an initial representation of the multiple people;
[0049] Step 3: The importance weights of language perception obtained by pooling the initial multi-person representations and embedding the text through self-attention calculation are injected into the initial multi-person representations using Hamada product broadcasting to obtain language-guided interactive perception features. The compressed interactive perception representation is then formed by pooling the person dimension and filtering with Top-K selection.
[0050] Step 3.1: Perform average pooling on the initial representations of multiple users according to the time dimension to obtain global temporal features;
[0051] Step 3.2: Use the BERT text encoder to extract features from the text in the scene to be analyzed, forming text embeddings;
[0052] Step 3.3: The importance weights of language perception obtained by calculating text embedding and global temporal features through self-attention mechanism are injected into the initial representation of multiple people using Hamada product broadcasting to form language-guided interactive perception features;
[0053] Step 3.3.1: Pool the global temporal features according to the time dimension, and use the self-attention mechanism to calculate the importance weights of the text embedding and global temporal features to obtain language perception.
[0054] Step 3.3.2: Use Hamada product to broadcast the importance weights of language perception into the initial representations of multiple users, forming language-guided interactive perception features;
[0055] Step 3.4: The importance scores obtained from the pooled language-guided interactive perception features of the character dimension using the self-attention mechanism are filtered using the Top-K selection method to form a compressed interactive perception representation;
[0056] Step 3.4.1: Pool the language-guided interactive perception features according to the character dimension, and use the self-attention mechanism to calculate the importance score of the language-guided interactive perception features;
[0057] Step 3.4.2: Set a filtering threshold K, use the Top-K selection method to filter the top K keyframes in the importance scores of the language-guided interactive perception features, and splice the selected keyframes together to form a compressed interactive perception representation.
[0058] Step 4: Adaptively gating the attention branch features, global branch features, and individual branch features constructed for the current character through multilayer perceptron and normalized probability distribution, and perform bidirectional cross-attention fusion with the social interaction representation obtained through self-attention calculation on the character dimension to obtain a multi-character fused representation;
[0059] Step 4.1: In terms of temporal sequence, construct attention branches, individual branches, and global branches for the motion changes of the current character in consecutive frames; further, attention branches are used to locate the interaction objects that the current character depends on and model their temporal relationship, individual branches are used to model the temporal continuity of the current character itself, and global branches are used to introduce scene-level global context.
[0060] Step 4.2: Perform self-attention calculation on the current person's features in the time dimension to obtain individual branch features; perform cross-attention calculation on the current person's features and the global features of other people to obtain global branch features; use the cross-individual importance weights obtained by cross-attention calculation on the current person's features and the global features pooled in the time dimension to filter the individuals to be followed, and perform cross-attention calculation on the current person's features and the features of the individuals to be followed to obtain the attention branch features;
[0061] Step 4.3: The features of the focus branch, global branch, and individual branch are passed through a multilayer perceptron and then subjected to normalized probability distribution and adaptive gating fusion with the corresponding branch features to obtain a temporal representation;
[0062] Step 4.4: Perform self-attention calculation on the character dimension according to the spatial relationships and social interaction dependencies between multiple characters based on time frames to obtain the social interaction representation;
[0063] Step 4.5: Perform bidirectional cross-attention fusion of temporal representation and social interaction representation to obtain multi-person fused representation;
[0064] Step 5: By aligning the text embedding and multi-character fusion representation through bidirectional cross-attention, we obtain the alignment features of language generation, and output text descriptions applicable to one or more of the following in multi-player sports competition scenarios: behavior description, interaction analysis, question and answer results, tactical analysis, commentary text, or character relationship judgment.
[0065] This invention discloses an interactive perception-based multi-player sports competition understanding and analysis system for implementing the aforementioned method. The system comprises an input acquisition module, a human body encoding module, an interactive perception filtering module, a social-temporal joint modeling module, a cross-modal alignment module, and a semantic generation module.
[0066] The input acquisition module is used to acquire video sequences and / or human motion sequences of the scene to be analyzed, and to construct multi-human input data; this data will be used as input to the human coding module.
[0067] The human body coding module is used to extract and fuse video features and motion features to obtain initial representations of multiple human bodies; it will serve as input to the interactive perception filtering module.
[0068] The interaction-aware filtering module is used to calculate the importance of each time frame in a multi-person sequence and filter key frames to obtain a compressed interaction-aware representation; this will serve as the input to the social temporal joint modeling module.
[0069] The social-temporal joint modeling module is used to model the temporal dynamics of an individual and the social interaction relationships between individuals, and outputs a multi-person fusion representation; it will serve as the input to the cross-modal alignment module.
[0070] A cross-modal alignment module is used to align multi-person fusion representations with language embeddings; it will serve as input to the semantic generation module.
[0071] The semantic generation module is used to generate behavioral descriptions, interaction understanding results, or analysis text for multi-human scenes based on alignment features;
[0072] In this embodiment, a video sequence V={v1,…,vT} of a multi-human scene is obtained, and a human motion sequence M corresponding to the video sequence is obtained; the human motion sequence contains key point information of P individuals at T time points. When the input scene already has a skeleton sequence, the skeleton sequence can be directly called; when the input scene only has video, the human motion sequence can be recovered from the video using an existing pose estimation model.
[0073] The video sequence is input into the video encoder, and the human motion sequence is input into the motion encoder to obtain video features and motion features respectively. Then, the video features and motion features are interactively fused using a cross-modal attention fusion module to form a multi-human initial representation that combines visual appearance information and human dynamic information.
[0074] The initial representation of multiple human figures is input into the interaction perception filtering module. First, pooling is performed along the time dimension to obtain global temporal features. Then, these global temporal features are interacted with the language embeddings corresponding to the task text or prompt text to calculate the importance weight of language guidance. This weight is then applied to the original multi-human figure features to highlight action and interaction segments relevant to the current task. Subsequently, frame-level importance scores are calculated for the time series aggregated along the human figure dimension, and the top K frames with the highest scores are selected as keyframes, thereby obtaining the compressed interaction perception representation.
[0075] The compressed interaction-aware representation is input into the social-temporal joint modeling module. For each person, a temporal modeling submodule is first constructed on their time series. For any query feature, the most relevant interaction object is first determined from the person-level global representation. Then, attention branches are executed for the interaction object, individual branch attention branches for their own sequence, and global branch attention branches for the scene-level context. Finally, the three attention outputs are adaptively fused through a gating mechanism to obtain the temporal representation.
[0076] At each time frame, social relationship modeling is performed along the person dimension to obtain a social interaction representation. Subsequently, a bidirectional cross-attention mechanism is used to fuse the temporal representation and the social interaction representation, and the transposed temporal information is injected back into the social interaction stream to further refine the semantics of group relationships. Finally, a multi-person fusion representation is output using a person perception aggregation function.
[0077] The multi-person fusion representation is input into the cross-modal alignment module and aligned with the language embedding through a bidirectional cross-attention structure to generate aligned features suitable for decoding large language models.
[0078] The alignment features are input into a large language model, which outputs natural language results corresponding to multi-person scenes. In social scenarios, it can output descriptions of group interactions, relationships between characters, and causal explanations; in sports scenarios, it can output action analysis, tactical judgments, match commentary, or training suggestions.
[0079] During the training phase, a multi-human sample set including video, human motion, and text annotations was constructed, and some basic encoder parameters were frozen, with only fine-tuning performed on the large language model or related adaptation parameters. The negative log-likelihood loss of the target text was used as the training objective to complete model training.
[0080] The multi-human sample set can consist of social interaction data, sports competition data, and publicly available data with human body key points and text annotations, which can be used to support the unified training of the multi-human understanding model in both everyday and professional scenarios.
[0081] The semantic output can be configured as description generation, question-and-answer analysis, tactical reasoning, automatic commenting, or training-assisted feedback, depending on the application task, to meet the needs of different business scenarios.
[0082] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for interactive perception-based understanding and analysis of multi-player sports competitions, characterized in that: Includes the following steps, Step 1: From the scene to be analyzed, the human motion sequence or the human motion sequence extracted from the video sequence through the motion estimation module is flipped according to the human dimension and the time dimension to form a multi-human motion sequence tensor; Step 2: Use cross attention to align and fuse the video features and motion features extracted by the Video-LLaVA video encoder and LLaMo motion encoder to form a multi-person initial representation; Step 3: The importance weights of language perception obtained by pooling the initial multi-person representations and embedding the text through self-attention calculation are injected into the initial multi-person representations using Hamada product broadcasting to obtain language-guided interactive perception features. The compressed interactive perception representation is then formed by pooling the person dimension and filtering with Top-K selection. Step 4: Adaptively gating the attention branch features, global branch features, and individual branch features constructed for the current character through multilayer perceptron and normalized probability distribution, and perform bidirectional cross-attention fusion with the social interaction representation obtained through self-attention calculation on the character dimension to obtain a multi-character fused representation; Step 4.1: In terms of temporal sequence, construct attention branches, individual branches, and global branches for the current character's motion changes in consecutive frames; Furthermore, the focus branch is used to locate the interaction objects that the current character depends on and model their temporal relationship, the individual branch is used to model the temporal continuity of the current character itself, and the global branch is used to introduce the scene-level global context. Step 4.2: Perform self-attention calculation on the current person's features in the time dimension to obtain individual branch features; perform cross-attention calculation on the current person's features and the global features of other people to obtain global branch features; use the cross-individual importance weights obtained by cross-attention calculation on the current person's features and the global features pooled in the time dimension to filter the individuals to be followed, and perform cross-attention calculation on the current person's features and the features of the individuals to be followed to obtain the attention branch features; Step 4.3: The features of the focus branch, global branch, and individual branch are passed through a multilayer perceptron and then subjected to normalized probability distribution and adaptive gating fusion with the corresponding branch features to obtain a temporal representation; Step 4.4: Perform self-attention calculation on the person dimension according to the spatial relationships and social interaction dependencies between multiple people based on time frames to obtain the social interaction representation; Step 4.5: Perform bidirectional cross-attention fusion of temporal representation and social interaction representation to obtain multi-person fused representation; Step 5: By aligning the text embedding and multi-character fusion representation through bidirectional cross-attention, we obtain the alignment features of language generation, and output text descriptions applicable to one or more of the following in multi-player sports competition scenarios: behavior description, interaction analysis, question and answer results, tactical analysis, commentary text, or character relationship judgment.
2. The interactive perception-based multi-player sports competition understanding and analysis method as described in claim 1, characterized in that: Step 2 is implemented as follows: Step 2.1: For the multi-person motion sequence tensor, the Video-LLaVA video encoder and the LLaMo motion encoder are used to extract the video features and motion features of the multi-person motion sequence tensor respectively; Step 2.2: Use the cross-modal alignment module to perform feature alignment and fusion of video features and motion features of multiple people using a cross-attention mechanism to form an initial representation of multiple people.
3. The interactive perception-based multi-player sports competition understanding and analysis method as described in claim 1, characterized in that: Step 3 is implemented as follows: Step 3.1: Perform average pooling on the initial representations of multiple users according to the time dimension to obtain global temporal features; Step 3.2: Use the BERT text encoder to extract features from the text in the scene to be analyzed, forming text embeddings; Step 3.3: The importance weights of language perception obtained by calculating text embedding and global temporal features through self-attention mechanism are injected into the initial representation of multiple people using Hamada product broadcasting to form language-guided interactive perception features; Step 3.4: The importance scores obtained from the pooled language-guided interactive perception features of the character dimension using the self-attention mechanism are filtered using the Top-K selection method to form a compressed interactive perception representation.
4. The interactive perception-based multi-player sports competition understanding and analysis method as described in claim 3, characterized in that: Step 3.3 is implemented as follows: Step 3.3.1: Pool the global temporal features according to the time dimension, and use the self-attention mechanism to calculate the importance weights of the text embedding and global temporal features to obtain language perception. Step 3.3.2: Use Hamada product to broadcast the importance weights of language perception into the initial representations of multiple people, forming language-guided interactive perception features.
5. The interactive perception-based multi-player sports competition understanding and analysis method as described in claim 3, characterized in that: Step 3.4 is implemented as follows: Step 3.4.1: Pool the language-guided interactive perception features according to the character dimension, and use the self-attention mechanism to calculate the importance score of the language-guided interactive perception features; Step 3.4.2: Set a filtering threshold K, use the Top-K selection method to filter the top K keyframes in the importance scores of the language-guided interactive perception features, and splice the selected keyframes together to form a compressed interactive perception representation.
6. A multi-player sports competition understanding and analysis system that implements the method described in claim 1, characterized in that: It includes an input acquisition module, a human body coding module, an interaction perception and filtering module, a social-temporal joint modeling module, a cross-modal alignment module, and a semantic generation module; The input acquisition module is used to acquire video sequences and / or human motion sequences of the scene to be analyzed, and to construct multi-human input data; It will be used as input to the human body coding module; The human body coding module is used to extract and fuse video features and motion features to obtain initial representations of multiple human bodies; It will be used as input for the interactive sensing filtering module; The interactive perception filtering module is used to calculate the importance of each time frame in the multi-human sequence and filter key frames to obtain a compressed interactive perception representation. It will be used as input to the social temporal joint modeling module; The social-temporal joint modeling module is used to model the temporal dynamics of an individual and the social interaction relationships between individuals, and output a multi-person fusion representation. It will be used as input to the cross-modal alignment module; The cross-modal alignment module is used to align the multi-human fusion representation with the language embedding; It will be used as input to the semantic generation module; The semantic generation module is used to generate behavioral descriptions, interactive understanding results, or analysis text for multi-human scenes based on alignment features.