Video processing method and device based on attitude estimation, equipment, medium and program
Through the video processing method based on pose estimation, the problem of unstable detection effect or low accuracy in complex scenarios of traditional methods is solved, and high-precision capture and video analysis of athletes' movements are achieved.
Patent Information
- Application Number
- CN202510177921.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional detection and skeleton recognition methods have unstable detection effects or low accuracy in complex scenarios, making it difficult to meet actual needs.
A video processing method based on pose estimation is adopted to improve the accuracy of athletes' motion capture by acquiring video frame sequences, object detection, pose estimation, motion trajectory generation and behavioral feature extraction.
It improves the capture accuracy and robustness of athletes' movements, and can more accurately capture the context information and category semantics of the target object, achieving efficient and low-cost video analysis and behavior strategy classification.
Smart Images

Figure CN119992664A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine vision technology, and in particular to a video processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on posture estimation. Background Art
[0002] The rapid development of digitization and science and technology has led to the rise of computer vision technology. Among them, motion capture technology digitizes human posture by analyzing and inferring human motion data in images or videos, and realizes the reconstruction of human movements and postures. It provides an important source of information for human motion analysis, human-computer interaction, virtual reality, and human behavior understanding, and has important research and application value.
[0003] At present, in the field of sports, detection and skeleton recognition methods are mainly used to predict or estimate human posture. However, traditional detection methods usually rely on algorithms based on rules or manually designed features, such as background subtraction, optical flow, or template matching based on edges and colors. These methods are sensitive to the complexity and changes of the scene. Especially in sports fields, such as courts, where there are occlusions, changes in lighting, and dynamic interference (such as non-target objects such as spectators and referees), the detection effect is unstable and it is difficult to achieve robustness and high accuracy. Traditional skeleton recognition methods usually rely on algorithms based on manually designed features or statistical models (such as PCA-based posture estimation), which have low recognition accuracy for key points of players, especially in complex scenes (such as multi-person interaction, partial occlusion of the target). Although deep learning methods have made some progress in the field of skeleton recognition, early methods based on convolutional neural networks (CNNs) still have limitations in capturing global features, especially when dealing with cross-frame associations and complex backgrounds, the low accuracy makes it difficult to meet actual needs.
[0004] In summary, traditional detection and skeleton recognition methods have the problem of unstable detection effects in complex scenes or low accuracy to meet actual needs. Therefore, it is necessary to propose improved technical means to solve this problem. Summary of the invention
[0005] Based on this, it is necessary to provide a video processing method, device, computer equipment, computer-readable storage medium and computer program product based on posture estimation that can improve the accuracy of athlete motion capture in response to the above technical problems.
[0006] In a first aspect, a video processing method based on posture estimation comprises:
[0007] Acquire a video frame sequence; the video frame sequence includes at least one video frame;
[0008] Perform target detection on each video frame to obtain the target object corresponding to each video frame and the category label corresponding to the target object;
[0009] Perform posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object;
[0010] Based on the same body key points of the same target object in each video frame, a motion trajectory sequence of the target object is generated;
[0011] The key behavior features of each target object are obtained based on the motion trajectory sequence, and the key behavior features are used for motion evaluation.
[0012] In one embodiment, target detection is performed on each video frame to obtain a target object corresponding to each video frame and a category label corresponding to the target object, including:
[0013] Extract multi-scale spatial features of each video frame;
[0014] Perform feature fusion on multi-scale spatial feature information and optimize feature extraction efficiency to generate feature pyramid;
[0015] Perform initial detection on the feature pyramid based on a single direct regression method to generate several candidate boxes in the video frame;
[0016] The candidate boxes are processed using the non-maximum suppression algorithm to obtain the target object and the category label corresponding to the target object.
[0017] In one embodiment, posture estimation is performed based on the target object and the category label corresponding to the target object to obtain the body key points of each target object, including:
[0018] Based on the target object, extract the spatial feature information, temporal feature information, and key point feature information of each video frame;
[0019] The spatial feature information, the temporal feature information and the key point feature information are fused and encoded to generate the global posture feature corresponding to the video frame; the global posture feature includes the key point query vector, the key point key vector and the key point value vector;
[0020] The key point query vector is used to perform self-attention decoding on the key point key vector and the key point value vector to generate each body key point and the corresponding coordinates of each body key point.
[0021] In one embodiment, the body key points include one or more of shoulder key points, hip key points, hand key points, and foot key points, and the coordinates corresponding to the body key points are used to characterize the spatial position and confidence of the body key points in the video frame.
[0022] In one embodiment, based on the same body key points of the same target object in each video frame, generating a motion trajectory sequence of the target object includes:
[0023] Calculating the body orientation of the target object based on the body key points of the target object in each video frame and the coordinate set corresponding to the body key points;
[0024] Using a motion trajectory generation method, a time series of the target object's body changes is generated based on the target object's body orientation;
[0025] The time series of the target object's physical changes are input into the observation system to obtain the target object's behavioral variables.
[0026] In one embodiment, the category labels include a goalie label and a penalty taker label;
[0027] After obtaining the key behavior features of each target object based on the motion trajectory sequence, it includes:
[0028] The strategy classification corresponding to each goalkeeper label and penalty taker label is obtained by fusing the key point query vector, key point key vector, key point value vector and time series using the logistic regression classification model.
[0029] Based on the key body points, key behavioral characteristics and corresponding strategy classifications of each target object, combined with time series modeling, the free throw success rate corresponding to each strategy classification is generated.
[0030] In a second aspect, the present application also provides a posture estimation device, comprising:
[0031] An acquisition module, used to acquire a video frame sequence; the video frame sequence includes at least one video frame;
[0032] The recognition module is used to perform target detection on each video frame to obtain the target object corresponding to each video frame and the category label corresponding to the target object;
[0033] An estimation module is used to perform posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object;
[0034] A processing module, for generating a motion trajectory sequence of the target object based on the same body key points of the same target object in each video frame;
[0035] The evaluation module is used to obtain key behavior features of each target object based on the motion trajectory sequence, and the key behavior features are used for motion evaluation.
[0036] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0037] In a fourth aspect, the present application further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0038] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0039] The above-mentioned video processing method, device, computer equipment, computer-readable storage medium and computer program product based on posture estimation input the acquired video frame sequence into the recognition module, extract and fuse multi-scale spatial features, and optimize the detection results using non-maximum suppression, and output the bounding box and category label of the target object corresponding to each video frame. Compared with the traditional method that only uses low-level features such as color, texture or edge, the high-level features extracted by the video processing method based on posture estimation provided by the present application can more accurately capture the contextual information and category semantics of the target object, and improve the detection accuracy and robustness.
[0040] Then, the target object corresponding to each video frame and the category label corresponding to the target object are input into the estimation module to extract the spatial position and confidence of the body key points of each target object in each video frame, remove abnormal key points and complete missing key points to achieve tracking of the target object identification.
[0041] Finally, the motion trajectory of the target object is generated based on the time series, and the classification model is combined to predict the shooting strategy of the penalty taker and the defensive strategy of the goalkeeper, thus achieving efficient and low-cost video analysis and behavioral strategy classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments of the present application or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 A diagram showing an application environment of a video processing method based on posture estimation in one embodiment;
[0044] Figure 2 is a flow chart of a video processing method based on posture estimation in one embodiment;
[0045] Figure 3 is a schematic flow chart of step 204 in one embodiment;
[0046] Figure 4 is a schematic flow chart of step 206 in one embodiment;
[0047] Figure 5 is a schematic flow chart of step 208 in one embodiment;
[0048] Figure 6 is a schematic diagram of an image of a target scene at a certain moment in an embodiment;
[0049] Figure 7 is a framework diagram of a posture estimation model in one embodiment;
[0050] Figure 8 is a schematic diagram of a posture estimation image of a target object in one embodiment;
[0051] Fig. 9 A structural block diagram of a video processing method and device based on posture estimation in one embodiment;
[0052] Fig.10 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] The video processing method based on posture estimation provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones and tablet computers, and the server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0055] In one embodiment, Figure 2As shown, a video processing method based on posture estimation is provided. This embodiment takes the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0056] Step 202: Acquire a video frame sequence; the video frame sequence includes at least one video frame.
[0057] The video frame sequence refers to a group of temporally continuous frame data including a multi-player sports scene.
[0058] Optionally, the terminal receives a video selection instruction, and selects an initial video including a multiplayer competitive sports scene based on the video selection instruction. After that, the terminal sends the initial video to the server. The server receives the initial video sent by the terminal and pre-processes the initial video, such as cropping the initial video according to time, so as to obtain a video frame sequence. The initial video is obtained by recording the multiplayer competitive game by a video acquisition device, and the video acquisition device may include a camera, a still camera, or a webcam. The server obtains a video frame sequence; the video frame sequence includes at least one video frame, and each video frame in the video frame sequence may be continuous in time, or arranged at a preset time interval.
[0059] Optionally, after the terminal selects the video frame sequence, the terminal sends the video frame sequence to the server, and the server receives the video frame sequence sent by the terminal. The disclosed embodiment does not limit the way in which the server obtains the video frame sequence, and it can be set according to actual conditions.
[0060] Step 204: Perform target detection on each video frame to obtain a target object corresponding to each video frame and a category label corresponding to the target object.
[0061] The target object refers to each target appearing in each video frame. In a football video, the target queue may be a person, such as a player.
[0062] The category label refers to the category of the target, such as an obstacle or a person, etc. In the multiplayer competitive sports scene, the category label may refer to the sports identity corresponding to the person, such as a goalkeeper, etc.
[0063] Optionally, the server pre-trains the target detection model, and the training process includes data enhancement, feature extraction and multi-target classification. In actual operation, the server will identify the sample objects and the actual positions corresponding to the sample objects in each video frame of the sample video, that is, mark the actual positions of the people in the sample video, where the sample video includes a multi-player competitive sports scene. The target detection model is obtained by training the sample objects and the actual positions corresponding to the target objects in each video frame of the sample video frame using a convolutional neural network.
[0064] Optionally, after the server obtains the video frame sequence, it inputs the video frame sequence into a pre-trained target detection model. After the target detection model performs a convolution operation on each video frame, it outputs the detection results corresponding to each video frame sequence. The detection results include the target object corresponding to each video frame and the category label corresponding to the target object.
[0065] Step 206: performing posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object.
[0066] The body key points refer to key points that can measure the posture of the target object, for example, they can be representative and important parts of the human skeletal structure.
[0067] Optionally, the posture estimation model is obtained by training sample video frames using a Transformer attention mechanism, where the attention mechanism may include a self-attention mechanism and a cross-attention mechanism.
[0068] Optionally, after obtaining the detection results corresponding to each video frame sequence, the server inputs the video frame sequence into a pre-trained posture estimation model, extracts the global features of each video frame according to actual needs, and represents it as a high-dimensional feature matrix, and then uses the decoder to restore the high-dimensional features generated by the encoder to specific body key point positions.
[0069] Step 208: Generate a motion trajectory sequence of the target object based on the same body key points of the same target object in each video frame.
[0070] Among them, the motion trajectory sequence refers to a set of ordered trajectory points whose positions change over time in a continuous period of time. In a football video, it can be the motion path of a player or goalkeeper during the time period covered by the video.
[0071] Optionally, after acquiring the body key points of the target object, the server processes the body key points of each identical target object in each video frame, calculates the body orientation of the target object to obtain key motion features, and converts the changes of the body key points over time into the motion trajectory of the target object.
[0072] Step 210: Obtain key behavior features of each target object based on the motion trajectory sequence; the key behavior features are used for motion evaluation.
[0073] Among them, key behavioral features refer to the characteristics of behavioral performance that play a key role in achieving goals, judging status or understanding the nature of behavior in multiplayer competitions, such as running speed, line of sight direction, and deceptive actions.
[0074] Among them, motion assessment refers to the systematic observation, measurement, analysis and evaluation of the sports performance, sports ability, sports status and other aspects of each target object. In football videos, this motion assessment can be the smoothness of the penalty taker's shooting action and the goalkeeper's ability to adjust his posture during the save.
[0075] Optionally, the server may describe the dynamic trajectory and action continuity of the key points over time through a statistical analysis module.
[0076] For example, the server may calculate the overall motion pattern and behavior variables of the target object through a feature extraction module.
[0077] For example, the server may perform dynamic change analysis on the time series data of key points of the target object's body based on a time series analysis model.
[0078] For example, the server performs a correlation analysis between the target object's movement pattern and the scoring success rate through the key behavior feature association analysis module to reveal the impact of the behavior features on the results of the competition.
[0079] In the above-mentioned video processing method based on posture estimation, first of all, compared with a single image or static data, a video frame sequence can present the complete process of the target object's movement. The target detection model can identify a specific target object from each frame of the video and determine its category label, which ensures the pertinence of the analysis and solves the problem that the analysis object is unclear or the analysis scope is too large in the background technology. By accurately locating the target object and its category, it is possible to eliminate the interference of other irrelevant information in the video, focus on the analysis of specific targets, and improve the efficiency and accuracy of the analysis.
[0080] Secondly, the motion trajectory sequence converts discrete body key points into continuous motion trajectories, intuitively showing the motion path of the target object in spatial and temporal dimensions, solving the problem that the motion path and rules of the target object cannot be clearly presented in the background technology.
[0081] Ultimately, the key behavioral features obtained by further abstracting and refining the motion trajectory can more accurately reflect the motion state of the target object and lay the foundation for comprehensive and objective motion evaluation of the target object.
[0082] In an exemplary embodiment, Figure 3 As shown, step 204 includes steps 302 to 308. Among them:
[0083] Step 302: extract multi-scale spatial features of each video frame.
[0084] Among them, multi-scale spatial features include edge features, texture features and background information of the target object; edge features indicate the boundary information of the target object in the image, including the clarity of the object outline, the shape and direction of the boundary, etc.; texture features indicate the information of the surface details of the target object, including the thickness, repeatability and directionality of the texture, etc.; background information indicates the environmental information around the target object, including the color distribution, lighting conditions and complexity of the background, etc.
[0085] Optionally, the server performs target detection on each video frame in the target sequence through a deep learning network, such as a convolutional neural network, and extracts multi-scale spatial feature information, including edge features, texture features, and background information of the target object, as input to the detection layer. Assuming that the size of the feature map output by the detection module of the convolutional neural network is (Wmodel, Hmodel, Dmodel), the detection module converts the feature into a feature representation of the size of (Wmodel×Hmodel, Dmodel) for direct prediction of the position coordinates, category label, and confidence of the target object.
[0086] Step 304: perform feature fusion on the multi-scale spatial feature information, and optimize the feature extraction efficiency to generate a feature pyramid.
[0087] For example, the edge features, texture features and background information of the target object are integrated into the multi-scale spatial feature information through the coding layer, and the cross-stage partial network module is used to optimize the feature extraction efficiency to generate a multi-scale feature pyramid.
[0088] Step 306: Perform initial detection on the feature pyramid based on a single direct regression method to generate several candidate boxes in the video frame.
[0089] Optionally, the decoding layer maps the global features in the multi-scale feature pyramid to the position and category of a specific target object through a single direct regression method, locates the bounding box of the target object layer by layer according to the pyramid structure of the multi-scale feature fusion, generates several candidate boxes, accurately predicts their positions in the video frame, and uses the classification module in the decoding layer to determine the category label of the target object according to its high-level semantic features.
[0090] Step 308: Process the candidate box using a non-maximum suppression algorithm to obtain a target object and a category label corresponding to the target object.
[0091] For example, the target detection prediction results are optimized by non-maximum suppression, overlapping bounding boxes are removed, and finally a set of bounding boxes of all target objects in each target video frame and their corresponding categories and confidence levels are output.
[0092] In this process, non-maximum suppression (NMS) calculates the confidence score of each target object, removes those bounding boxes that overlap more with higher confidence target boxes, and retains the target with the highest confidence. Among them, all candidate boxes are sorted and arranged from high to low according to the confidence score. Each candidate box is traversed, and the overlap between the current box and other candidate boxes is calculated. The commonly used calculation method is the intersection over union (IoU). If the overlap between a box and other boxes exceeds the set threshold, the overlapping boxes are removed and the most representative boxes are retained. Finally, a set of bounding boxes containing all target objects is output, and each bounding box is accompanied by a corresponding category label and confidence.
[0093] It will be appreciated that in some embodiments, the category tags include a goalkeeper tag and a penalty taker tag.
[0094] In this embodiment, the high-level features extracted by the deep learning network for target detection can more accurately capture the contextual information and category semantics of the target object, thereby improving detection accuracy and robustness, compared with the traditional method that only uses low-level features such as color, texture or edge.
[0095] In one embodiment, if Figure 4 As shown, step 206 includes steps 402 to 406. Among them:
[0096] Step 402: Based on the target object, extract the spatial feature information, temporal feature information, and key point feature information of each video frame.
[0097] Among them, spatial feature information refers to the spatial characteristics of the target object related to the bounding box in each video frame, including the position, size, shape, etc. of the target object; temporal feature information refers to the temporal dependency and dynamic relationship between video frames, including the order, interval and duration of movement between frames. By analyzing the bounding box changes of the target object in consecutive frames, the dynamic behavior pattern of the target object can be captured; body key point feature information: refers to the position and posture of the joints and body parts of the target object, the coordinates and confidence of the key points. This information can describe the movement, direction and posture of the target object, for example, the tilt angle or movement amplitude of the target object's body while running.
[0098] Step 404: Fusion-encode the spatial feature information, the temporal feature information and the key point feature information to generate a global posture feature corresponding to the video frame; the global posture feature includes a key point query vector Query, a key point key vector Key and a key point value vector Value.
[0099] Among them, the key point query vector Query represents the identification information of the target object, including the unique description of the body key points, which is used to track the action and posture of the target object between each video frame; the key point query vector Query is mapped to the key point key vector Key, which is used to construct the feature relationship of the target object's body key points and associate the key point features of the target object in different time and space; based on the mapping of the key point query vector Query to the key point key vector Key, the feature that best describes the posture and action characteristics of the target object is called the key point value vector Value.
[0100] For example, the server encodes the spatial feature information, temporal feature information, and body key point feature information extracted by the convolutional neural network into global posture features through the encoding layer, and adds the position encoding in the original video frame to the key point query vector Query and the key point key vector Key to capture the spatial and temporal dependencies of the key points.
[0101] Step 406: Use the key point query vector Query to perform self-attention decoding on the key point key vector Key and the key point value vector Value to generate each body key point and the coordinates corresponding to each body key point.
[0102] For example, the key point query vector Query is used to perform self-attention decoding on the key point key vector Key and the key point value vector Value through a decoding layer to generate each body key point and the coordinates corresponding to each body key point.
[0103] The encoded global pose features are input into the Transformer decoder, and the output size of the decoder is the same as the input, which is maintained as (Wmodel×Hmodel, Dmodel).
[0104] During the decoding process, the key point query vector Query is initialized into several detection queries for the pose estimation task. Each detection query corresponds to a key point of a target object and is responsible for detecting and tracking the newly appearing target key points in the current video frame.
[0105] It can be understood that in some embodiments, the body key points include one or more of shoulder key points, hip key points, hand key points, and foot key points, and the coordinates corresponding to the body key points are used to characterize the spatial position and confidence of the body key points in the video frame.
[0106] In this embodiment, the more the number of detection queries is, the more key points can be identified and located. The decoder fuses the encoded features to finally generate high-precision target object posture features for subsequent action analysis and behavior classification tasks.
[0107] After step 206, in some embodiments, as Figure 5 As shown, step 208 includes steps 502 to 506. Among them:
[0108] Step 502: Calculate the body orientation of the target object based on the body key points of the target object in each video frame and the coordinate sets corresponding to the body key points.
[0109] Among them, body orientation includes key movement characteristics such as the direction angles of the shoulders and hips, and the position of the non-kicking foot.
[0110] For example, for the body posture estimation result set of the target object, statistical features of body key points of each video frame are extracted, wherein the statistical features of body key points include spatial distribution of key points, skeleton angle relationship and symmetry features.
[0111] Step 504: Generate a time series of the target object's body changes based on the target object's body orientation using a motion trajectory generation method.
[0112] For example, for the body posture estimation result set of the target object, a time series of body key points of each video frame is extracted, wherein the time series is used to describe the dynamic trajectory and action continuity of the key points over time.
[0113] Step 506: Input the time series of the target object's body changes into the observation system to obtain the target object's behavior variables.
[0114] The behavioral variables include the kicker's running speed, shoulder direction, non-support foot position, and changes in fake moves before shooting, the goalkeeper's moving direction, advance movement amplitude, and timing of save reaction, and other characteristics.
[0115] For example, the server calculates the overall motion pattern and behavior variables of the target object through a feature extraction module, wherein the motion pattern includes the target object's movement amplitude, body posture balance, and motion path distribution.
[0116] In some embodiments, step 210 includes steps 602 to 604. Among them:
[0117] Step 602: Utilize a logistic regression classification model to fuse the key point query vector Query, the key point key vector Key, the key point value vector Value, and the time series to obtain a strategy classification corresponding to each of the goalkeeper labels and the penalty taker labels.
[0118] Among them, the penalty taker label strategy is classified into: the penalty taker is not affected by the goalkeeper's action when shooting, and decides the shooting direction in advance; or the penalty taker adjusts the shooting direction according to the goalkeeper's movement. The goalkeeper label strategy is classified into: the goalkeeper moves in advance and tries to predict the penalty taker's shooting direction; or the goalkeeper waits for the penalty taker to move before reacting to avoid exposing the defense direction too early.
[0119] For example, based on the time series analysis model, the dynamic change analysis of the time series data of the key points of the target object's body is performed. First, the motion trajectory of the target object in continuous video frames is calculated through motion trajectory analysis, and the key turning points, movement speed changes and overall path curves are extracted to comprehensively characterize the motion pattern of the target object. Secondly, combined with the dynamic change analysis of posture, the relative position changes of the key points of the body are used to evaluate the smoothness of the penalty kicker's shooting action and the goalkeeper's posture adjustment ability during the save. In addition, through the action continuity evaluation, the time series smoothing algorithm is used to calculate the target object's action stability indicators, such as the average speed change rate, turning frequency and pause duration, so as to quantitatively measure the coordination and stability of the action.
[0120] Step 604: Based on the body key points, key behavioral features and corresponding strategy classifications of each target object, combined with time series modeling, a free throw success rate corresponding to each strategy classification is generated.
[0121] For example, through the behavioral feature association analysis module, the movement pattern of the target object and the penalty kick success rate are correlated to reveal the impact of behavioral features on the game results. In the penalty kicker behavioral feature analysis, combined with variables such as running speed, feint frequency, and shooting direction adjustment, the success rate distribution of different behavioral patterns in penalty kick scenarios is statistically analyzed to explore the relationship between action features and shooting effects. In the goalkeeper behavioral feature analysis, the focus is on studying the rationality of the amplitude and direction selection of the early save action, and combined with the save reaction timing, the actual impact of different defensive strategies on the penalty kick defense effect is evaluated to provide a scientific basis for strategy optimization.
[0122] In this embodiment, the fused key point query vector Query, key point key vector Key, key point value vector Value and time series features are input into the logistic regression classification model, and the model is trained to classify the shooting strategy of the penalty kicker and the defensive strategy of the goalkeeper. Based on the extracted body key point data and the key behavioral characteristics of the penalty kick observation system, combined with time series modeling, the model can also quantify the impact of different strategies on the success rate of penalty kicks. The final output includes the classification results of the shooting strategy of the penalty kicker in the current video frame, the classification results of the goalkeeper's defensive strategy, and the correlation analysis between strategy classification and success rate, which is used to deeply understand the effectiveness of penalty kicks and defensive strategies.
[0123] In one feasible embodiment, step 202 is performed to obtain a video frame sequence.
[0124] See also Figure 6 , obtain an image at the Tth moment in the target scene (football field), wherein the image includes multiple target objects, where T=1, 2, ..., t-1, t, t+1, ...
[0125] Specifically, a video frame having a motion track of a target object at time T is obtained, and at least four first marked positions and first positions to be projected corresponding to the first marked positions are obtained from the video frame, wherein the first marked positions include the corner positions of the football field in the video frame, and the first projection positions include the corner positions of the two-dimensional court. Subsequently, a first homography matrix is determined based on the first marked positions and the first positions to be projected, and the positions to be projected of the motion tracks of each target object in the two-dimensional court are determined based on the first homography matrix and the motion tracks of each target object.
[0126] The above process is executed to obtain video frames corresponding to the first half of the football video captured from the first perspective, and video frames corresponding to the second half of the football video captured from the second perspective.
[0127] Step 204 is executed to perform target detection on each video frame to obtain a target object corresponding to each video frame and a category label corresponding to the target object.
[0128] Specifically, the video frames corresponding to the target football video are input, and the video frames at T moments are input into the improved feature extraction network to obtain the corresponding detection results at T moments.
[0129] Among them, after performing a convolution operation on the image at each moment, a downsampled feature map is obtained, and a multi-layer convolution operation is performed on the downsampled feature map to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map. A convolution operation with a convolution kernel of 3x3 is used for feature extraction, and the third-scale feature map is processed by a spatial pyramid module to obtain a processed third-scale feature map. The processed third-scale feature map is upsampled, and the feature map of the second scale is jump-connected to obtain a processed second-scale feature map. The processed second-scale feature map is upsampled, and the feature map of the first scale is jump-connected to obtain a processed first-scale feature map. The processed first-scale feature map, the processed second-scale feature map, and the processed third-scale feature map are spliced and fused to extract the detection result of each target object in each of the images.
[0130] Step 206 is executed to perform posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object.
[0131] See also Figure 7 , the video frame after target detection is input into the Patch Embedding module, which divides the input video frame into fixed-size image blocks (patches) and embeds each image block into a low-dimensional feature representation to form an initial feature vector sequence. The main task of this module is to map each image block into a low-dimensional feature representation space to form an initial feature vector sequence.
[0132] The output feature vector sequence (Embedding Sequence) will be used as the input of the subsequent Transformer encoder. The feature vector sequence is input into the Transformer encoder. The encoder consists of multiple Transformer Blocks, each of which is responsible for capturing global features and performing context modeling. Among them:
[0133] The multi-head self-attention mechanism captures the global dependencies of the input data through parallel multi-head attention operations. Each attention head calculates the correlations of different subspaces respectively, focusing on the complex interactions between key points (such as body joints), and calculates the attention weights through the dot product operation of the key point query vector Query, the key point key vector Key, and the key point value vector Value, thereby selectively focusing on important feature areas.
[0134] The feedforward network performs nonlinear transformation on the features extracted by the self-attention module to further enhance the feature representation capability, and includes two fully connected layers and an activation function.
[0135] The residual connection allows the output of each Transformer Block to be directly superimposed on its input, thereby alleviating the gradient vanishing problem, and layer normalization improves the stability and convergence speed of training.
[0136] The encoder extracts the global features of each video frame through these modules and represents them as a high-dimensional feature matrix, which contains the fusion information of spatial, temporal and motion features.
[0137] The output of the Transformer encoder is fed into the decoder module to generate the final pose estimation result. The task of the decoder is to restore the high-dimensional features generated by the encoder to specific key point locations and generate corresponding visualization results. The decoding process includes:
[0138] Initialize several query vectors, each query corresponds to the features of a human joint (such as head, shoulder, knee, etc.), perform dot product operation on the query vector and the features output by the encoder, extract feature representations related to key points, convert the extracted key point features into specific coordinate values (such as (x, y) position) through a fully connected layer, perform temporal modeling on multi-frame features (optional), further improve the stability of prediction, and generate a bounding box surrounding the target object based on the key points of posture estimation.
[0139] Finally, the decoder output includes: the pose estimation result of the target object in each video frame, including the specific location of all key points, the bounding box of the target object for locating the overall target, and the visualization output, which superimposes the key points and bounding box onto the original video frame. Figure 8 shown.
[0140] For the penalty taker, the direction angles of the shoulder, hip and non-supporting foot are extracted in each video frame to generate the key point query vector Query, key point key vector Key and key point value vector Value of the current frame; for the goalkeeper, the angles of the neck and hip are extracted, the advance motion amplitude is calculated, and the key point query vector Query, key point key vector Key and key point value vector Value are generated at the same time.
[0141] When the current video frame is not the first video frame, the key point prediction results and tracking vector of the previous frame are fused with the key point features of the current frame, and the self-attention mechanism is used to generate the key point prediction results and tracking vector of the current frame to capture the action continuity and dynamic changes in the time series, thereby improving the dynamic analysis capabilities of the penalty taker and goalkeeper strategies.
[0142] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0143] Based on the same inventive concept, the embodiment of the present application also provides a posture estimation device for implementing the above-mentioned video processing method based on posture estimation. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in one or more posture estimation device embodiments provided below can refer to the above-mentioned limitations on the video processing method based on posture estimation, which will not be repeated here.
[0144] In an exemplary embodiment, see Fig. 9 , a posture estimation device is provided, comprising: an acquisition module 901, a recognition module 902, an estimation module 903, a processing module 904 and an evaluation module 905, wherein:
[0145] The acquisition module 901 is used to acquire a video frame sequence; the video frame sequence includes at least one video frame;
[0146] The recognition module 902 is used to perform target detection on each video frame to obtain the target object corresponding to each video frame and the category label corresponding to the target object;
[0147] An estimation module 903 is used to perform posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object;
[0148] A processing module 904 is used to generate a motion trajectory sequence of the target object based on the same body key points of the same target object in each video frame;
[0149] The evaluation module 905 is used to obtain key behavior features of each target object based on the motion trajectory sequence; the key behavior features are used for motion evaluation.
[0150] In an exemplary embodiment, the identification module 902 includes:
[0151] The first acquisition module is used to extract multi-scale spatial features of each video frame;
[0152] The first optimization module is used to fuse multi-scale spatial feature information and optimize feature extraction efficiency to generate a feature pyramid;
[0153] The first detection module is used to perform initial detection on the feature pyramid based on a single direct regression method to generate several candidate frames in the video frame;
[0154] The first processing module is used to process the candidate frame using a non-maximum suppression algorithm to obtain a target object and a category label corresponding to the target object.
[0155] The estimation module 903 includes:
[0156] The second acquisition module is used to extract spatial feature information, temporal feature information, and key point feature information of each video frame based on the target object;
[0157] The second processing module is used to fuse and encode the spatial feature information, the temporal feature information and the key point feature information to generate a global posture feature corresponding to the video frame; the global posture feature includes a key point query vector, a key point key vector and a key point value vector;
[0158] The third processing module is used to use the key point query vector to perform self-attention decoding on the key point key vector and the key point value vector to generate each body key point and the coordinates corresponding to each body key point.
[0159] The processing module 904 includes:
[0160] A first calculation module, configured to calculate the body orientation of the target object based on the body key points of the target object in each video frame and the coordinate sets corresponding to the body key points;
[0161] A second calculation module is used to generate a time series of body changes of the target object based on the body orientation of the target object by using a motion trajectory generation method;
[0162] The observation module inputs the time series of the target object’s physical changes into the observation system to obtain the target object’s behavioral variables.
[0163] The evaluation module 905 includes:
[0164] A classification module is used to fuse key point query vectors, key point key vectors, key point value vectors and time series using a logistic regression classification model to obtain strategy classifications corresponding to each goalkeeper label and penalty taker label;
[0165] The analysis module generates the free throw success rate corresponding to each strategy classification based on the body key points, key behavioral characteristics and corresponding strategy classification of each target object, combined with time series modeling.
[0166] Each module in the above-mentioned posture estimation device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0167] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.10 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a video processing method based on posture estimation is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0168] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0169] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0170] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0171] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0173] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0174] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0175] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A video processing method based on posture estimation, characterized in that: include: Acquire a video frame sequence, wherein the video frame sequence includes at least one video frame; Performing target detection on each of the video frames to obtain a target object corresponding to each of the video frames and a category label corresponding to the target object; Performing posture estimation based on the target object and the category label corresponding to the target object to obtain body key points of each target object; Based on the same body key points of the same target object in each of the video frames, generating a motion trajectory sequence of the target object; Based on the motion trajectory sequence, key behavior features of each target object are obtained, and the key behavior features are used for motion evaluation.
2. The method according to claim 1, characterized in that The performing target detection on each of the video frames to obtain a target object corresponding to each of the video frames and a category label corresponding to the target object includes: Extract multi-scale spatial features of each video frame; Performing feature fusion on the multi-scale spatial feature information and optimizing feature extraction efficiency to generate a feature pyramid; Performing initial detection on the feature pyramid based on a single direct regression method to generate a number of candidate frames in the video frame; The candidate box is processed using a non-maximum suppression algorithm to obtain a target object and a category label corresponding to the target object.
3. The method according to claim 1, characterized in that The step of performing posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object includes: Based on the target object, extracting the spatial feature information, the temporal feature information, and the key point feature information of each of the video frames; The spatial feature information, the temporal feature information and the key point feature information are fused and encoded to generate a global posture feature corresponding to the video frame; the global posture feature includes a key point query vector, a key point key vector and a key point value vector; The key point query vector is used to perform self-attention decoding on the key point key vector and the key point value vector to generate each of the body key points and the coordinates corresponding to each of the body key points.
4. The method according to claim 3, characterized in that The body key points include one or more of shoulder key points, hip key points, hand key points, and foot key points, and the coordinates corresponding to the body key points are used to characterize the spatial position and confidence of the body key points in the video frame.
5. The method according to claim 3, characterized in that: The step of generating a motion trajectory sequence of the target object based on the same body key points of the same target object in each of the video frames comprises: Calculating the body orientation of the target object based on the body key points of the target object in each of the video frames and the coordinate sets corresponding to the body key points; Using a motion trajectory generation method, a time series of body changes of the target object is generated based on the body orientation of the target object; The time series of the target object's physical changes are input into an observation system to obtain the target object's behavioral variables.
6. The method according to claim 5, characterized in that The category labels include a goalkeeper label and a penalty taker label; After obtaining the key behavior features of each target object based on the motion trajectory sequence, the method includes: The key point query vector, the key point key vector, the key point value vector and the time series are fused by using a logistic regression classification model to obtain a strategy classification corresponding to each of the goalkeeper labels and the penalty taker labels; Based on the key body points, key behavioral characteristics and corresponding strategy classifications of each target object, combined with time series modeling, the free throw success rate corresponding to each strategy classification is generated.
7. A posture estimation device, characterized in that: The device comprises: An acquisition module, used for acquiring a video frame sequence, wherein the video frame sequence includes at least one video frame; An identification module is used to perform target detection on each of the video frames to obtain a target object corresponding to each of the video frames and a category label corresponding to the target object; An estimation module, used for performing posture estimation based on the target object and the category label corresponding to the target object to obtain the body key points of each target object; A processing module, configured to generate a motion trajectory sequence of the target object based on the same body key points of the same target object in each of the video frames; An evaluation module is used to obtain key behavior features of each target object based on the motion trajectory sequence, and the key behavior features are used for motion evaluation.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Football juggling counting method and system based on attitude constraint and time sequence analysis
CN120747830A