Video data processing method, system and equipment for sports events and medium
By extracting behavioral semantic features at different scales and cascading and fusing features from sports event video data, the problem of insufficient accuracy and stability in athlete trajectory recognition in traditional methods is solved, and high-precision motion trajectory analysis is achieved.
Patent Information
- Application Number
- CN202511545899.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-24
AI Technical Summary
Traditional video analysis methods struggle to achieve high-precision and stable motion trajectory recognition when tracking athletes' trajectories, especially in rapidly changing or multi-person intertwined scenarios. They neglect the deep semantic relationships between image frames, resulting in low accuracy and stability in trajectory analysis.
By receiving video data from sports events, we extract behavioral semantic features at different scales. Combining convolutional neural networks and graph neural networks, we determine the perceptual attention and behavioral dependencies of motion associations and trajectory contour features. We then use perceptual attention and salient factors to perform feature cascade fusion to achieve labeling of motion trajectories.
It improves the labeling accuracy of athletes' movement trajectories, enhances the stability and precision of trajectory tracking, maintains high accuracy in complex scenarios, and provides more accurate motion analysis support.
Smart Images

Figure CN121564601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and more specifically, to methods, systems, devices, and media for processing video data for sports events. Background Technology
[0002] With the digitalization and intelligentization of sports events, video processing applications are no longer limited to simple image playback. Instead, they combine advanced technologies such as image recognition, motion capture, and real-time playback, enabling viewers to experience richer information and visual effects during matches. Through video data analysis, details such as athletes' technical movements and trajectories can be revealed in depth, providing coaches with more precise tactical guidance. The application of video processing technology makes sports events more intelligent and efficient, becoming an important bridge for interaction between athletes, coaches, and viewers, and promoting the further development of the sports industry.
[0003] In traditional video analytics methods, athletes' trajectories are typically tracked using simple motion detection algorithms. These methods often rely solely on positional changes within the image, neglecting the deeper semantic relationships between adjacent frames. Because athletes' movements are usually complex and dynamic, traditional methods frequently suffer from trajectory recognition errors or instability in rapidly changing scenes and situations involving multiple athletes. This is especially true when athletes are moving at high speeds or performing complex movements; simple motion detection algorithms struggle to accurately capture minute changes in the trajectory, resulting in low accuracy and stability in motion trajectory analysis. This is particularly pronounced in high-intensity events or densely packed scenes, where traditional trajectory tracking methods often fail to meet the demands for high accuracy. Therefore, achieving feature fusion of contours and behaviors during motion in video data labeling has become a significant challenge for the industry. Summary of the Invention
[0004] This application provides a method, system, device, and medium for processing video data of sports events, which can achieve feature fusion of contours and behaviors during motion in the tagging process of video data.
[0005] In a first aspect, this application provides a method for processing video data for sports events, comprising the following steps: Receive video data of sports events, and extract behavioral semantic features at different scales from each frame of the video data to obtain shallow semantic features and deep semantic features of each frame. The motion association relationship of the trajectory pixels during the movement of the target athlete is determined based on the semantic similarity of the shallow semantic features between adjacent image frames. Then, the perceptual attention of the trajectory contour features of the target athlete during movement is determined by the difference between the motion association relationship and the semantic segmentation mask information between adjacent image frames. The behavioral dependencies of the target athlete are determined by deep semantic features between adjacent image frames, and then significant factors of the trajectory behavior features of the target athlete during movement are determined based on the behavioral dependencies and the global trend information of the target athlete's movement trajectory in the video data. Based on the perceived attention and the saliency factor, the trajectory contour features and trajectory behavior features are cascaded and fused to obtain the cascaded fusion features of the target athlete's movement trajectory. Then, the movement trajectory of the target athlete in the video data is labeled based on the cascaded fusion features.
[0006] Preferably, the extraction of behavioral semantic features at different scales for each frame of the video data to obtain the shallow semantic features and deep semantic features of each frame specifically includes: Each frame of the video data is preprocessed to obtain the preprocessed frame of each image. The shallow convolutional layers in the convolutional neural network are used to perform semantic convolution on each preprocessed image frame to extract the shallow semantic features of each image frame. The shallow semantic features of each frame of the image are used as input parameters for the deep convolutional layers in the convolutional neural network, and then the deep semantic features of each frame of the image are extracted through the deep convolutional layers.
[0007] Preferably, determining the motion correlation of trajectory pixels during the movement of a target athlete based on the semantic similarity of shallow semantic features between adjacent image frames specifically includes: Determine the semantic similarity of shallow semantic features between any two adjacent image frames; The similarity matrix of pixel locations in the video data is determined by using all semantic similarities; Based on the similarity matrix, track pixel pairs with matching similarity are selected from adjacent image frames of the video data; The motion relationships between trajectory pixels during the target athlete's movement are determined by all trajectory pixel pairs.
[0008] Preferably, the perceptual attention that determines the trajectory contour features of the target athlete during movement based on the motion correlation and the difference in semantic segmentation mask information between adjacent image frames specifically includes: For every two adjacent image frames, the semantic segmentation mask information of the adjacent image frames is obtained, and then the degree of difference of the semantic segmentation mask information between adjacent image frames is determined. Trajectory pixel matching is performed on the semantic segmentation mask information of adjacent image frames to obtain the matching degree of trajectory pixels in the two semantic segmentation mask information. The state transition amount of the target athlete's trajectory contour between adjacent image frames is determined by the difference in semantic segmentation mask information between adjacent image frames and the matching degree of the trajectory pixels, thereby obtaining the state transition amount of the target athlete's trajectory contour between every two adjacent image frames. Perceptual attention is determined based on all state transitions and the motion correlations to identify the trajectory contour features of the target athlete during movement.
[0009] Preferably, determining the behavioral dependencies of the target athlete through deep semantic features between adjacent image frames specifically includes: The deep semantic features of each frame of the image are used as nodes in a directed graph. The edge weights between nodes in a directed graph are determined by the similarity of deep semantic features between adjacent image frames. The feature dependency graph of the target athlete's movement process is determined by all directed graph nodes and all edge weights; Extract the behavioral dependencies of the target athlete from the feature dependency graph.
[0010] Preferably, determining the significant factors of the target athlete's trajectory behavior characteristics based on the behavioral dependency relationship and the global trend information of the target athlete's movement trajectory in the video data specifically includes: Perform trend analysis on the trajectory of the target athlete throughout the entire video data to extract global trend information of the target athlete's movement trajectory; Determine the trajectory behavior characteristics between image frames within a sliding window during the target athlete's movement; Based on the behavioral dependency, the global trend information of the motion trajectory is temporally correlated with the trajectory behavior features between adjacent image frames, thereby determining the significant factors of the trajectory behavior features of the target athlete during movement.
[0011] Preferably, the cascaded fusion of trajectory contour features and trajectory behavior features based on the perceived attention and the saliency factor to obtain the cascaded fusion features of the target athlete's trajectory specifically includes: Based on the dual attention mechanism, the fusion coefficients of trajectory contour features and trajectory behavior features are determined according to the perceptual attention and the salient factors. By using the fusion coefficients of trajectory contour features and trajectory behavior features, the time-aligned trajectory contour features and trajectory behavior features are cascaded and fused to obtain the cascaded fused features of the target athlete's movement trajectory.
[0012] Secondly, this application provides a video data processing system for sports events, comprising: The feature extraction module is used to receive video data of sports events and extract behavioral semantic features at different scales from each frame of the video data to obtain shallow semantic features and deep semantic features of each frame. The processing module is used to determine the motion association relationship of the trajectory pixels during the movement of the target athlete based on the semantic similarity of the shallow semantic features between adjacent image frames, and then determine the perceptual attention of the trajectory contour features of the target athlete during movement based on the difference between the motion association relationship and the semantic segmentation mask information between adjacent image frames. The processing module is further configured to determine the behavioral dependency relationship of the target athlete through deep semantic features between adjacent image frames, and then determine the significant factors of the trajectory behavior features of the target athlete during movement based on the behavioral dependency relationship and the global trend information of the target athlete's movement trajectory in the video data. The execution module is used to perform feature cascade fusion of trajectory contour features and trajectory behavior features based on the perceived attention and the saliency factor to obtain the cascade fusion features of the target athlete's movement trajectory, and then perform labeling processing on the movement trajectory of the target athlete in the video data based on the cascade fusion features.
[0013] Thirdly, this application provides a computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described video data processing method for sports events.
[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video data processing method for sports events.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: In this embodiment, video data of a sports event is received, and behavioral semantic features of each frame in the video data are extracted at different scales to obtain shallow semantic features and deep semantic features of each frame. The motion association relationship of trajectory pixels during the movement of the target athlete is determined based on the semantic similarity of the shallow semantic features between adjacent image frames. Then, the perceptual attention of the trajectory contour features of the target athlete during movement is determined based on the difference between the motion association relationship and the semantic segmentation mask information between adjacent image frames. The behavioral dependency relationship of the target athlete is determined through the deep semantic features between adjacent image frames. Then, the saliency factor of the trajectory behavioral features of the target athlete during movement is determined based on the behavioral dependency relationship and the global trend information of the target athlete's movement trajectory in the video data. The trajectory contour features and trajectory behavioral features are cascaded and fused based on the perceptual attention and the saliency factor to obtain the cascaded fusion feature of the target athlete's movement trajectory. Finally, the movement trajectory of the target athlete in the video data is labeled based on the cascaded fusion feature.
[0016] Therefore, this application achieves cascaded fusion of trajectory contour features and trajectory behavior features through perceptual attention and saliency factors to obtain cascaded fusion features of the target athlete's movement trajectory. Then, based on these cascaded fusion features, the movement trajectory of the target athlete in the video data is labeled. First, by extracting shallow and deep semantic features from each frame, detailed movements of the athlete can be captured at different scales, effectively improving the ability to capture complex dynamic behaviors and avoiding the problem of neglecting rapid movements or losing details in complex scenes in traditional methods. Second, the perceptual attention of the trajectory contour features during the target athlete's movement is determined based on the motion correlation of trajectory pixels and the difference in semantic segmentation mask information between adjacent image frames. This allows the tracking of the athlete's trajectory to not only rely on positional changes but also consider the semantic correlation between different frames, thereby improving the stability and accuracy of trajectory tracking, especially in complex scenes with drastic changes in athlete movements or multiple athletes intertwined, maintaining high accuracy. Then, based on the behavior of the target athlete... By using the global trend information of the target athlete's movement trajectory in the video data and the dependency relationship, salient factors of the target athlete's trajectory behavior characteristics are determined. These salient factors allow for a deeper understanding of athlete behavior beyond static position, accurately reflecting the temporal dependency between behavioral patterns and movement trajectories, thus enhancing the temporality and forward-looking nature of behavioral analysis. Finally, by using perceptual attention and salient factors, trajectory contour features and trajectory behavior features are cascaded and fused to obtain cascaded fusion features of the target athlete's movement trajectory. This considers not only the athlete's movement contour features but also their movement behavior features, providing more accurate movement trajectory recognition for labeling. This fusion enhances the integration capability of multi-dimensional information, providing more comprehensive and accurate support for the analysis of highly complex sports events. Furthermore, the cascaded fusion features are used to classify the target athlete's movement trajectory patterns and assign corresponding labels. In summary, this application's solution can achieve feature fusion of movement contour and behavior in video data labeling, thereby improving the labeling accuracy of athlete movement trajectories in video data. Attached Figure Description
[0017] Figure 1 This is an exemplary flowchart of a video data processing method for sports events, according to some embodiments of this application; Figure 2 This is a flowchart illustrating the determination of motion correlation relationships according to some embodiments of this application; Figure 3 This is a flowchart illustrating the determination of behavioral dependencies according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of a video data processing system for sports events, according to some embodiments of this application; Figure 5 This is a schematic diagram of the structure of a computer device that implements a video data processing method for sports events, according to some embodiments of this application. Detailed Implementation
[0018] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] refer to Figure 1 The figure is an exemplary flowchart of a video data processing method for sports events according to some embodiments of this application. The video data processing method 100 for sports events mainly includes the following steps: In step 101, video data of a sports event is received, and behavioral semantic features of each frame in the video data are extracted at different scales to obtain shallow semantic features and deep semantic features of each frame.
[0020] It should be noted that receiving video data from sports events in this application refers to acquiring video image data generated during the sports event through a data interface or acquisition device.
[0021] In some embodiments, extracting behavioral semantic features at different scales from each frame of the video data to obtain the shallow semantic features and deep semantic features of each frame can be achieved through the following steps: Each frame of the video data is preprocessed to obtain the preprocessed frame of each image. The shallow convolutional layers in the convolutional neural network are used to perform semantic convolution on each preprocessed image frame to extract the shallow semantic features of each image frame. The shallow semantic features of each frame of the image are used as input parameters for the deep convolutional layers in the convolutional neural network, and then the deep semantic features of each frame of the image are extracted through the deep convolutional layers.
[0022] It should be noted that, in this application, shallow convolutional layers refer to convolutional layers located at the front end of a convolutional neural network, with a small receptive field, used to extract local texture and edge features of an image. Specifically, shallow convolutional layers in this application refer to the first three convolutional layers in a convolutional neural network. Deep convolutional layers in this application refer to convolutional layers located at the back end of a convolutional neural network, with a large receptive field, used to extract the overall structure and higher-order semantic information of an image. Specifically, deep convolutional layers in this application refer to convolutional layers after the third layer in a convolutional neural network. It should also be noted that the number of convolutional layers in the convolutional neural network in this application is not less than 6 layers.
[0023] Additionally, it should be noted that the shallow semantic features in this application refer to low-level visual features such as local texture, edge information, and color distribution of the image extracted through shallow convolutional layers, which are used to capture the basic contour information of the target athlete; the deep semantic features in this application refer to high-level semantic features such as the overall shape of the target, action patterns, and behavioral information extracted through deep convolutional layers, which are used to characterize the movement trends and behavioral relationships of the target athlete.
[0024] In specific implementation, preprocessing each frame of the video data to obtain the preprocessed frame can be achieved in the following way: For each frame of the video data, Gaussian filtering is used for noise reduction, normalization is used to improve the stability of the image frame, and scale transformation (such as bilinear interpolation) is used to improve the multi-scale adaptability of image frame feature extraction. The processed image is then used as the preprocessed frame image, thus obtaining the preprocessed frame image for each frame. Semantic convolution processing is performed on the preprocessed frame image through shallow convolutional layers in a convolutional neural network to extract the shallow semantic features of each frame image. This can be achieved in the following way: The preprocessed image frame is input into a shallow convolutional layer (e.g., the first three layers of ResNet, VGG, or MobileNet), and then a small-sized convolutional kernel (e.g., 1×1) preset in the shallow convolutional layer is used to extract the edge information, local texture information, and color information of the image frame. Color distribution information is used to extract features, and the extracted features are used as shallow semantic features of image frames. These shallow semantic features can be used to capture the basic contour information of the target athlete during movement. The shallow semantic features of each frame are used as input parameters for deep convolutional layers in a convolutional neural network. The extraction of deep semantic features of each frame can be achieved by passing the shallow semantic features of each frame as input to deep convolutional layers, where deep convolutional layers are convolutional layers after the third layer in a convolutional neural network, such as the Bottleneck structure of ResNet. The overall shape, movement pattern, and behavior information of the target athlete are extracted through deeper network stacking (convolutional layers after the third layer), larger receptive fields (such as 3×3 stacking structure), and global pooling. The extracted overall shape, movement pattern, and behavior information are used as the deep semantic features of each frame.
[0025] In step 102, the motion association relationship of the trajectory pixels during the movement of the target athlete is determined based on the semantic similarity of the shallow semantic features between adjacent image frames. Then, the perceptual attention of the trajectory contour features of the target athlete during movement is determined by the difference between the motion association relationship and the semantic segmentation mask information between adjacent image frames.
[0026] In some embodiments, reference Figure 2As shown in the figure, this is a flowchart illustrating the process of determining motion association in some embodiments of this application. In this embodiment, determining the motion association of trajectory pixels during the movement of a target athlete based on the semantic similarity of shallow semantic features between adjacent image frames can be achieved using the following steps: Determine the semantic similarity of shallow semantic features between any two adjacent image frames; The similarity matrix of pixel locations in the video data is determined by using all semantic similarities; Based on the similarity matrix, track pixel pairs with matching similarity are selected from adjacent image frames of the video data; The motion relationships between trajectory pixels during the target athlete's movement are determined by all trajectory pixel pairs.
[0027] It should be noted that the semantic similarity in this application is an indicator that measures the degree of similarity of shallow semantic features between adjacent image frames; the similarity matrix in this application is an indicator that quantifies the degree of similarity of pixels between adjacent image frames in the shallow semantic feature space; the trajectory pixel pair in this application refers to a pair of pixels that match the same motion trajectory between adjacent image frames; and the motion association relationship of trajectory pixels in this application is an indicator that measures the motion consistency of trajectory pixel pairs between adjacent image frames in spatial semantics.
[0028] In specific implementation, determining the semantic similarity of shallow semantic features between two adjacent image frames can be achieved in the following way: the cosine similarity of shallow semantic features between two adjacent image frames can be used as the semantic similarity of shallow semantic features between two adjacent image frames. In other embodiments, other similarity metrics can also be used to quantify the semantic similarity of shallow semantic features between two adjacent image frames, which is not specifically limited here. Determining the similarity matrix of pixel positions in video data through all semantic similarities can be achieved in the following way: the semantic similarity of shallow semantic features between two adjacent image frames can be used as the similarity label between the two adjacent image frames. Then, using pixel position as an index, the matching score of corresponding pixels between adjacent image frames is used as the similarity score of pixel positions between every two adjacent image frames. The similarity labels and similarity scores between every two adjacent image frames are then arranged in chronological order of the video data to obtain a matrix, which serves as the similarity matrix of pixel positions in the video data. This similarity matrix can be used to quickly filter out close-up data segments of athletes in the video data. Filtering out trajectory pixel pairs with similarity matching from adjacent image frames of the video data based on the similarity matrix can be achieved in the following way: a threshold can be set in the similarity matrix, which can be set based on historical similarity. Pixels with similarity higher than the threshold can be filtered out. For adjacent image frames with a threshold, an optical flow-based matching algorithm (such as Lucas-Kanade) is used to perform pixel-level motion trajectory matching on the adjacent image frames. The pixel pairs with matched trajectories are then used as the trajectory pixel pairs with similarity matching in the adjacent image frames. The motion correlation of trajectory pixels during the target athlete's movement can be determined by considering the selected trajectory pixel pairs as nodes in a graph structure, using the similarity between pixel pairs as edge weights, and connecting all nodes through these edge weights to obtain a weighted graph. This weighted graph can then be used to describe the motion correlation of trajectory pixels during the target athlete's movement. It should be noted that... The construction of the weighted graph in this application specifically includes: first, treating the selected trajectory pixel pairs as nodes in the graph structure, and using the matching degree or similarity between pixels as the weights of the edges to construct a weighted graph model; second, adopting trajectory connection strategies, such as the Euclidean distance method, to connect temporally adjacent trajectory points at the pixel level to form a coherent motion trajectory; then, using temporal smoothing methods (such as Kalman filtering or Hidden Markov Model HMM) to eliminate matching errors and improve the continuity and stability of the trajectory; finally, eliminating abnormal matching points through motion constraints (such as velocity continuity and angle change constraints) to ensure that the trajectory conforms to the real motion law; and finally, using the output of the weighted graph model as the weighted graph.
[0029] In some embodiments, the perceptual attention that determines the trajectory contour features of a target athlete during movement based on the motion association and the difference in semantic segmentation mask information between adjacent image frames can be implemented using the following steps: For every two adjacent image frames, the semantic segmentation mask information of the adjacent image frames is obtained, and then the degree of difference of the semantic segmentation mask information between adjacent image frames is determined. Trajectory pixel matching is performed on the semantic segmentation mask information of adjacent image frames to obtain the matching degree of trajectory pixels in the two semantic segmentation mask information. The state transition amount of the target athlete's trajectory contour between adjacent image frames is determined by the difference in semantic segmentation mask information between adjacent image frames and the matching degree of the trajectory pixels, thereby obtaining the state transition amount of the target athlete's trajectory contour between every two adjacent image frames. Perceptual attention is determined based on all state transitions and the motion correlations to identify the trajectory contour features of the target athlete during movement.
[0030] It should be noted that the semantic segmentation mask information in this application refers to a binary annotation matrix (i.e., athlete and background) used to represent the semantic category of each pixel in an image frame, where the value of each pixel corresponds to a specific semantic category, used to segment the region contour of the target athlete, and to achieve accurate localization and region differentiation of a specific target (such as an athlete) in the image; the state transition quantity in this application refers to the change in the trajectory contour of the target athlete between adjacent image frames; and the perceptual attention in this application is an indicator that measures the importance of the trajectory contour features in perceiving the athlete's motion contour features at different time stages.
[0031] In specific implementation, obtaining the semantic segmentation mask information of adjacent image frames and then determining the difference between the semantic segmentation mask information of adjacent image frames can be achieved in the following way: A semantic segmentation model in deep learning (such as DeepLabV3+ or Mask R-CNN) can be used to obtain the semantic segmentation mask information of adjacent image frames, and the Dice coefficient of the semantic segmentation mask information between adjacent frames can be used as the difference between the semantic segmentation mask information between adjacent image frames. Trajectory pixel matching is performed on the semantic segmentation mask information of adjacent image frames to obtain the matching degree of trajectory pixels in the two semantic segmentation mask information. This can be achieved in the following way: Trajectory pixel pairs between adjacent image frames are obtained, and the corresponding pixels in each semantic segmentation mask information are marked as trajectory pixels based on the pixel positions of these trajectory pixel pairs. Then, based on a pixel-level matching strategy using cosine similarity, the trajectory pixels in the semantic segmentation mask information of adjacent image frames are matched. Finally, the matching result is used as the matching degree of trajectory pixels in the two semantic segmentation mask information. Matching can quantify the trajectory consistency of semantic segmentation mask information between adjacent image frames. The determination of the state transition amount of the target athlete's trajectory contour between adjacent image frames by using the difference degree of semantic segmentation mask information and the matching degree of the trajectory pixels can be achieved in the following way: the difference degree of semantic mask information and the matching degree of trajectory pixels can be weighted and fused, and the result of the weighted fusion can be used as the state transition amount of the target athlete's trajectory contour between adjacent image frames. By weighted fusion of these two quantities, the degree of trajectory change of the target athlete during the movement can be comprehensively evaluated. The weights of the difference degree and the matching degree can be preset based on historical fusion data. In this application, the weight of the difference degree is set to 0.4, and the weight of the matching degree is set to 0.6. In other embodiments, other weight values can be set according to the specific experimental purpose, which are not specifically limited here; the perceptual attention for determining the trajectory contour features of the target athlete during movement based on all state transition quantities and the motion correlation can be implemented in the following way: the state transition quantities and motion correlation corresponding to every two frames of images can be used as input to a Transformer model based on an attention mechanism. The Transformer model calculates the attention score of the athlete's trajectory contour features in the video data, and the obtained attention score is used as the perceptual attention for the trajectory contour features of the target athlete during movement. Through perceptual attention, it can be determined which frames are more important for the recognition of trajectory contour features. This method can adaptively strengthen the focus on the athlete's behavioral features, thereby extracting the key actions and behavioral patterns of the athlete during movement; it should be noted that in this application, the Transformer model uses the state transition quantities and motion correlation between the two frames as input to a Transformer model based on an attention mechanism. The Transformer model calculates the attention score of the athlete's trajectory contour features in the video data, and uses the obtained attention score ... target athlete's trajectory contour features. Through perceptual attention, it can be determined which frames are more important for the recognition of trajectory contour features. This method can adaptively strengthen the focus on the athlete's behavioral features, thereby extracting the key actions and behavioral patterns of the athlete during movement; it should be noted that in this application, the Transformer model calculates the attention score of the athlete's trajectory contour features as input to a Transformer model based The Transformer model calculates attention scores for athlete trajectory contour features in video data by first taking the state transitions and motion associations between every two frames as input. These inputs are mapped to query, key, and value vectors. The self-attention mechanism calculates the similarity between the query and key to obtain an attention score for each input pair. These scores represent the importance of each frame for identifying the target athlete's trajectory contour features; a higher score indicates a greater contribution. Then, these scores are weighted and applied to the corresponding value vectors to highlight important image frames. Finally, a perceptual attention to the athlete's trajectory contour features during movement is generated. In this way, the Transformer model can automatically learn and adjust which frames are most critical for athlete trajectory recognition, effectively focusing on key regions and movements of the athlete's behavior.
[0032] In step 103, the behavioral dependencies of the target athlete are determined by the deep semantic features between adjacent image frames, and then the significant factors of the target athlete's trajectory behavior features during movement are determined based on the behavioral dependencies and the global trend information of the target athlete's movement trajectory in the video data.
[0033] In some embodiments, reference Figure 3 As shown in the figure, this is a flowchart illustrating the process of determining behavioral dependencies in some embodiments of this application. In this embodiment, determining the behavioral dependencies of a target athlete through deep semantic features between adjacent image frames can be achieved using the following steps: In step 1031, the deep semantic features of each frame of image are used as nodes in a directed graph. In step 1032, the edge weights between directed graph nodes are determined by the similarity of deep semantic features between adjacent image frames; In step 1033, the feature dependency graph of the target athlete's movement process is determined by all directed graph nodes and all edge weights; In step 1034, the behavioral dependencies of the target athlete are extracted from the feature dependency graph.
[0034] It should be noted that the directed graph nodes in this application refer to nodes in a directed graph within a graph neural network. The directed graph acts as the core of the data structure in the graph neural network, primarily used to represent directional relational information. A directed graph consists of nodes and directed edges, where nodes represent objects (such as deep semantic features of image frames), and directed edges represent the dependencies or information flow between these objects. Compared to undirected graphs, directed graphs can preserve the directionality of relationships, making information transmission sequential and hierarchical. For example, in temporal modeling of video data, directed graphs can represent causal relationships between time steps, enabling graph neural networks to more accurately model temporal correlations. Furthermore, in the propagation of athletes' trajectory behavior features, the adjacency matrix of a directed graph can be used to guide information aggregation, allowing graph neural networks to transmit information from source nodes to target nodes based on edge directions, thus better conforming to dependency structures in the real world. The feature dependency graph in this application refers to a directed graph structure constructed based on the correlation between features; the behavioral dependency in this application refers to the correlation between the behavioral trajectories of athletes in different movement modes during exercise.
[0035] In specific implementation, treating the deep semantic features of each frame as directed graph nodes can be achieved in the following way: In the graph neural network, the deep semantic features of each frame can be mapped to a high-dimensional vector, and the mapped high-dimensional vector is used as a node of the directed graph. Determining the edge weights between directed graph nodes based on the similarity of deep semantic features between adjacent image frames can be achieved in the following way: Similarity measurement methods (such as cosine similarity or Euclidean distance) can be used to calculate the similarity of deep semantic features between adjacent image frames, and the normalized similarity value is used as the edge weight between the corresponding directed graph nodes of adjacent image frames. Determining the feature dependency graph of the target athlete's movement process using all directed graph nodes and all edge weights can be achieved in the following way: Connecting all directed graph nodes using edge weights, and using the resulting directed graph as the feature dependency graph of the target athlete's movement process. Extracting the behavioral dependencies of the target athlete from the feature dependency graph can be achieved in the following way: Graph clustering algorithms can be used. (Such as spectral clustering) identifies key motion patterns in the feature dependency graph and uses the dependencies between motion patterns as the behavioral dependencies of the target athlete. It should be noted that the motion pattern refers to a stable and identifiable combination of motion features exhibited by the target athlete within a specific time window, usually determined by its key posture, movement trajectory, or mechanical properties. In the feature dependency graph, motion patterns correspond to clusters of nodes with similar deep semantic features. To extract these motion patterns, a spectral clustering algorithm can be used. This method first performs feature dimensionality reduction on the feature dependency graph using the Laplacian matrix, mapping high-dimensional motion features to a low-dimensional space, thereby highlighting the clustering structure between nodes. Then, K-means or other clustering methods are used to classify the dimensionality-reduced features, grouping similar motion states into the same pattern. Finally, the dependencies between different motion patterns are determined by the edge structure of the feature dependency graph, that is, the connectivity and weights between motion patterns reflect the behavioral evolution trajectory of the athlete in different states, thereby constructing their behavioral dependencies.
[0036] In some embodiments, determining the salient factors of the target athlete's trajectory behavior characteristics based on the behavioral dependencies and the global trend information of the target athlete's movement trajectory in the video data can be achieved using the following steps: Perform trend analysis on the trajectory of the target athlete throughout the entire video data to extract global trend information of the target athlete's movement trajectory; Determine the trajectory behavior characteristics between image frames within a sliding window during the target athlete's movement; Based on the behavioral dependency, the global trend information of the motion trajectory is temporally correlated with the trajectory behavior features between adjacent image frames, thereby determining the significant factors of the trajectory behavior features of the target athlete during movement.
[0037] It should be noted that the global trend information in this application refers to the overall change pattern of the target athlete's movement trajectory in the entire video data; the sliding window in this application is a short window (i.e., 3-5 frames); the trajectory behavior features in this application refer to the spatiotemporal change features of the trajectory points of the target athlete during the movement; and the significance factor in this application is an indicator that measures the importance of trajectory behavior features in extracting the athlete's movement behavior features at different time stages.
[0038] In specific implementation, trend analysis of the target athlete's trajectory in the entire video data can be performed to extract global trend information of the target athlete's motion trajectory. This can be achieved in the following way: Global analysis of the target athlete's motion trajectory in the video can be performed using optical flow methods (such as the RAFT algorithm), extracting motion direction features, velocity change features, and acceleration features from the motion trajectory. Then, trend analysis is performed on these motion direction features, velocity change features, and acceleration features, and the trend analysis results are used as the global trend information of the target athlete's motion trajectory. Determining the trajectory behavior features between image frames within a sliding window during the target athlete's movement can be achieved in the following way: Trajectory points of the target athlete are detected within the sliding window, and geometric features such as velocity, acceleration, and motion direction of the athlete within the sliding window are calculated based on the trajectory point positions. Then, a convolutional neural network is used to extract the athlete's local motion model from the geometric features such as velocity, acceleration, and motion direction. The resulting local motion patterns are used as trajectory behavior features between image frames. It should be further explained that local motion patterns refer to the typical movements of an athlete within a specific time period. Specifically, key point positions of the athlete in each frame can be obtained through trajectory point detection methods (such as OpenPose), and motion vectors (velocity, acceleration, direction of motion, etc.) between adjacent frames can be calculated. These geometric features can then be used as input for feature extraction using a convolutional neural network (CNN). CNN automatically learns and extracts spatial features of the athlete's local behavior, such as gait, running style, or turning movements, through hierarchical convolutional operations. Through multi-layer convolutional operations, CNN can extract the athlete's motion patterns within a specific time period, such as running, pausing, or changing speed, thus providing specific trajectory behavior features for each sliding window. These local motion patterns can be further used for athlete behavior analysis, pattern recognition, and behavior prediction.
[0039] Furthermore, in specific implementation, the determination of significant factors of the trajectory behavior characteristics of the target athlete during movement can be achieved by temporally associating the global trend information of the motion trajectory with the trajectory behavior characteristics between adjacent image frames based on the aforementioned behavioral dependencies. This can be accomplished in the following way: global trend information and trajectory behavior characteristics between adjacent image frames can be used as input through temporal modeling methods (such as LSTM and Transformer). By learning temporal dependencies, the evolution of athlete behavior can be captured. Behavioral dependencies provide temporal connection information between nodes, which helps the model accurately predict behavioral changes between different time points. Through temporal attention mechanisms (such as the self-attention module in Transformer), the model can automatically weight the trajectory behavior characteristics of different time periods and use the weights corresponding to the trajectory behavior characteristics of different time periods as significant factors of the trajectory behavior characteristics of the target athlete during movement. Through significant factors, the model can focus on key frames or key action sequences, thereby revealing the key driving factors of athlete behavior and providing more accurate understanding and prediction of motion behavior.
[0040] In step 104, the trajectory contour features and trajectory behavior features are cascaded and fused according to the perceptual attention and the saliency factor to obtain the cascaded fusion features of the target athlete's movement trajectory, and then the movement trajectory of the target athlete in the video data is labeled based on the cascaded fusion features.
[0041] In some embodiments, the cascaded fusion of trajectory contour features and trajectory behavior features based on the perceived attention and the saliency factor to obtain the cascaded fused features of the target athlete's trajectory can be achieved through the following steps: Based on the dual attention mechanism, the fusion coefficients of trajectory contour features and trajectory behavior features are determined according to the perceptual attention and the salient factors. By using the fusion coefficients of trajectory contour features and trajectory behavior features, the time-aligned trajectory contour features and trajectory behavior features are cascaded and fused to obtain the cascaded fused features of the target athlete's movement trajectory.
[0042] It should be noted that the cascaded fusion feature in this application is a fusion feature used to characterize the morphological structure and temporal behavior pattern of the motion trajectory.
[0043] In specific implementation, the determination of the fusion coefficients of trajectory contour features and trajectory behavior features based on the dual attention mechanism, according to the perceptual attention and the salient factors, can be achieved in the following way: A dual attention mechanism can be used, where perceptual attention and salient factors work together. Perceptual attention uses channel attention (such as Squeeze-and-Excitation, SE) to weight the trajectory contour features, highlighting important trajectory morphological information (such as curvature changes and inflection point features). The salient factors use temporal attention (such as Transformer's multi-head self-attention) to analyze the temporal relevance of the athlete's movement behavior, highlighting key time steps (such as rapid acceleration and sudden changes in direction). The weights calculated by both are then used as the fusion coefficients of the trajectory contour features and trajectory behavior features, respectively. The cascaded fusion of the time-aligned trajectory contour features and trajectory behavior features using the fusion coefficients of the trajectory contour features and trajectory behavior features, to obtain the cascaded fusion features of the target athlete's trajectory, can be achieved in the following way: First, a Dynamic Time Warping algorithm can be used. DTW (Trajectory-Driven Warping) aligns the extracted trajectory contour features and trajectory behavior features along the time dimension to obtain aligned trajectory contour features and trajectory behavior features. The fusion coefficients of the trajectory contour features are then used as the fusion weights of the aligned trajectory contour features, and the fusion coefficients of the trajectory behavior features are used as the fusion weights of the aligned trajectory behavior features. The trajectory contour features and trajectory behavior features are then cascaded and fused according to their corresponding fusion weights. The result of this fusion process is used as the cascaded fused feature of the target athlete's trajectory. It should be further noted that the cascaded fusion process in this application refers to the weighted combination of features from different sources (i.e., trajectory contour features and trajectory behavior features) during feature cascading to form a unified and more information-rich feature representation. In this example, the cascaded fusion process first aligns the trajectory contour features and trajectory behavior features in time to ensure their matching at time steps. Then, it uses the fusion coefficients of the trajectory contour features and trajectory behavior features as weighting factors to adjust the contribution of the two types of features. Specifically, cascaded fusion can be done by directly splicing the two types of features, that is, connecting them in the channel dimension or feature dimension so that they share the representation space, or by using a weighted summation method, that is, linearly weighting the features according to the fusion coefficient and then adding them together. This enhances the data's expressive power while retaining important information about the trajectory morphology and behavior features. By using cascaded fusion, the geometric features and temporal behavior features of the athlete's trajectory can be effectively combined, thereby improving the ability to describe the movement trajectory and laying the foundation for subsequent behavior analysis and prediction.
[0044] It should be noted that, in this application, labeling the motion trajectory of a target athlete in video data based on the cascaded fusion features means classifying the motion trajectory pattern of the target athlete using the cascaded fusion features and assigning corresponding labels.
[0045] On the other hand, in some embodiments, this application provides a video data processing system for sports events, with reference to... Figure 4 The figure is a schematic diagram of the structure of a video data processing system for sports events according to some embodiments of this application. The video data processing system 400 for sports events includes: a feature extraction module 401, a processing module 402, and an execution module 403, which are described below: Feature extraction module 401, in this application, is mainly used to receive video data of sports events and extract behavioral semantic features at different scales from each frame of the video data to obtain shallow semantic features and deep semantic features of each frame. Processing module 402, in this application, is used to determine the motion association relationship of trajectory pixels during the movement of the target athlete based on the semantic similarity of shallow semantic features between adjacent image frames, and then determine the perceptual attention of the trajectory contour features of the target athlete during movement based on the motion association relationship and the difference between the semantic segmentation mask information between adjacent image frames. In this application, the processing module 402 is further configured to determine the behavioral dependency relationship of the target athlete through the deep semantic features between adjacent image frames, and then determine the significant factors of the trajectory behavioral features of the target athlete during movement based on the behavioral dependency relationship and the global trend information of the target athlete's movement trajectory in the video data. The execution module 403 in this application is mainly used to perform feature cascade fusion of trajectory contour features and trajectory behavior features according to the perceptual attention and the saliency factor to obtain the cascade fusion features of the target athlete's movement trajectory, and then perform label processing on the movement trajectory of the target athlete in the video data based on the cascade fusion features.
[0046] In addition, this application also provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described video data processing method for sports events.
[0047] In some embodiments, reference Figure 5 The figure is a schematic diagram of the structure of a computer device implementing a video data processing method for sports events, according to some embodiments of this application. The video data processing method for sports events in the above embodiments can be implemented through... Figure 5The computer device shown is used to implement this, and the computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.
[0048] Processor 501 can be a general-purpose central processing unit (CPU) or an application-specific integrated circuit (ASIC).
[0049] The communication bus 502 can be used to transmit information between the aforementioned components.
[0050] Memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CDROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 503 may exist independently and be connected to processor 501 via communication bus 502. Memory 503 may also be integrated with processor 501.
[0051] The memory 503 stores program code for executing the solution of this application, and its execution is controlled by the processor 501. The processor 501 executes the program code stored in the memory 503. The program code may include one or more software modules. The video data processing method for sports events in the above embodiments can be implemented by the processor 501 and one or more software modules in the program code in the memory 503.
[0052] Communication interface 504 uses any transceiver-like device to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0053] In a specific implementation, as one example, a computer device may include multiple processors, each of which may be a single-core (single CPU) processor or a multi-core (multi CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0054] The aforementioned computer device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device can be a desktop computer, a portable computer, a network server, a handheld digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. This application does not limit the type of computer device.
[0055] In addition, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video data processing method for sports events.
[0056] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0057] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for processing video data for sports events, characterized in that, Includes the following steps: Receive video data of sports events, and extract behavioral semantic features at different scales from each frame of the video data to obtain shallow semantic features and deep semantic features of each frame. The motion association relationship of the trajectory pixels during the movement of the target athlete is determined based on the semantic similarity of the shallow semantic features between adjacent image frames. Then, the perceptual attention of the trajectory contour features of the target athlete during movement is determined by the difference between the motion association relationship and the semantic segmentation mask information between adjacent image frames. The behavioral dependencies of the target athlete are determined by deep semantic features between adjacent image frames, and then significant factors of the trajectory behavior features of the target athlete during movement are determined based on the behavioral dependencies and the global trend information of the target athlete's movement trajectory in the video data. Based on the perceived attention and the saliency factor, the trajectory contour features and trajectory behavior features are cascaded and fused to obtain the cascaded fusion features of the target athlete's movement trajectory. Then, the movement trajectory of the target athlete in the video data is labeled based on the cascaded fusion features.
2. The method as described in claim 1, characterized in that, The video data is processed by extracting behavioral semantic features at different scales for each frame, resulting in shallow and deep semantic features for each frame. Specifically, these features include: Each frame of the video data is preprocessed to obtain the preprocessed frame of each image. The shallow convolutional layers in the convolutional neural network are used to perform semantic convolution on each preprocessed image frame to extract the shallow semantic features of each image frame. The shallow semantic features of each frame of the image are used as input parameters for the deep convolutional layers in the convolutional neural network, and then the deep semantic features of each frame of the image are extracted through the deep convolutional layers.
3. The method as described in claim 1, characterized in that, Determining the motion correlation of trajectory pixels during the movement of a target athlete based on the semantic similarity of shallow semantic features between adjacent image frames specifically includes: Determine the semantic similarity of shallow semantic features between any two adjacent image frames; The similarity matrix of pixel locations in the video data is determined by using all semantic similarities; Based on the similarity matrix, track pixel pairs with matching similarity are selected from adjacent image frames of the video data; The motion relationships between trajectory pixels during the target athlete's movement are determined by all trajectory pixel pairs.
4. The method as described in claim 1, characterized in that, The perceptual attention that determines the trajectory contour features of the target athlete during movement based on the motion correlation and the difference in semantic segmentation mask information between adjacent image frames specifically includes: For every two adjacent image frames, the semantic segmentation mask information of the adjacent image frames is obtained, and then the degree of difference of the semantic segmentation mask information between adjacent image frames is determined. Trajectory pixel matching is performed on the semantic segmentation mask information of adjacent image frames to obtain the matching degree of trajectory pixels in the two semantic segmentation mask information. The state transition amount of the target athlete's trajectory contour between adjacent image frames is determined by the difference degree of semantic segmentation mask information between adjacent image frames and the matching degree of the trajectory pixels, thereby obtaining the state transition amount of the target athlete's trajectory contour between every two adjacent image frames. Perceptual attention is determined based on all state transitions and the motion correlations to identify the trajectory contour features of the target athlete during movement.
5. The method as described in claim 1, characterized in that, Determining the behavioral dependencies of a target athlete through deep semantic features between adjacent image frames specifically includes: The deep semantic features of each frame of the image are used as nodes in a directed graph. The edge weights between nodes in a directed graph are determined by the similarity of deep semantic features between adjacent image frames. The feature dependency graph of the target athlete's movement process is determined by all directed graph nodes and all edge weights; Extract the behavioral dependencies of the target athlete from the feature dependency graph.
6. The method as described in claim 1, characterized in that, The significant factors for determining the trajectory behavior characteristics of the target athlete during movement based on the behavioral dependencies and the global trend information of the target athlete's movement trajectory in the video data specifically include: Perform trend analysis on the trajectory of the target athlete throughout the entire video data to extract global trend information of the target athlete's movement trajectory; Determine the trajectory behavior characteristics between image frames within a sliding window during the target athlete's movement; Based on the behavioral dependency, the global trend information of the motion trajectory is temporally correlated with the trajectory behavior features between adjacent image frames, thereby determining the significant factors of the trajectory behavior features of the target athlete during movement.
7. The method as described in claim 1, characterized in that, Based on the perceived attention and the saliency factor, the trajectory contour features and trajectory behavior features are cascaded and fused to obtain the cascaded fused features of the target athlete's movement trajectory, specifically including: Based on the dual attention mechanism, the fusion coefficients of trajectory contour features and trajectory behavior features are determined according to the perceptual attention and the salient factors. By cascading and fusing the trajectory contour features and trajectory behavior features after alignment in the time dimension using the fusion coefficients of the trajectory contour features and the trajectory behavior features, the cascaded fused features of the target athlete's movement trajectory are obtained.
8. A video data processing system for sports events, characterized in that, include: The feature extraction module is used to receive video data of sports events and extract behavioral semantic features at different scales from each frame of the video data to obtain shallow semantic features and deep semantic features of each frame. The processing module is used to determine the motion association relationship of the trajectory pixels during the movement of the target athlete based on the semantic similarity of the shallow semantic features between adjacent image frames, and then determine the perceptual attention of the trajectory contour features of the target athlete during movement based on the difference between the motion association relationship and the semantic segmentation mask information between adjacent image frames. The processing module is further configured to determine the behavioral dependency relationship of the target athlete through deep semantic features between adjacent image frames, and then determine the significant factors of the trajectory behavior features of the target athlete during movement based on the behavioral dependency relationship and the global trend information of the target athlete's movement trajectory in the video data. The execution module is used to perform feature cascade fusion of trajectory contour features and trajectory behavior features based on the perceived attention and the saliency factor to obtain the cascade fusion features of the target athlete's movement trajectory, and then perform labeling processing on the movement trajectory of the target athlete in the video data based on the cascade fusion features.
9. A computer device comprising a memory and a processor, the memory storing code, characterized in that, The processor is configured to acquire the code and execute the video data processing method for sports events as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the video data processing method for sports events as described in any one of claims 1 to 7.