Golf swing key node interception and video generation method and system
By extracting swing features using high frame rate cameras and deep learning models, key nodes of the golf swing are identified and compressed, solving the problems of data redundancy and insufficient recognition accuracy in existing technologies, and achieving efficient swing analysis and data compression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-27
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to accurately identify key moments and effectively compress data in golf swing analysis, resulting in high data storage and transmission costs, and a lack of comprehensive utilization of club trajectory and body center of gravity changes.
The swing process is captured by a high frame rate camera, and the human posture, club trajectory and center of gravity change features are extracted by a deep learning model. Key nodes are identified by weighted fusion, and representative images are selected based on image quality evaluation to generate a simplified swing video with encoding and compression.
It enables efficient identification of key swing points, reduces data volume, and maintains image clarity and analytical value, facilitating teaching and remote transmission.
Smart Images

Figure CN121838010A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of golf swing critical point selection technology, and in particular to a method, system, computer equipment, and storage medium for golf swing critical point selection and video generation. Background Technology
[0002] In golf teaching and training, the common method for swing analysis is to capture swing videos with a camera, and then have coaches or analysis software observe and compare different swing phases. Existing solutions often use ordinary frame rate cameras or mobile phone cameras for recording, which can only roughly divide the backswing and downswing into general phases, making it difficult to capture subtle changes in movement during high-speed swings and frame-level key moments such as the moment of impact. On the other hand, there are solutions that use high-frame-rate, high-definition cameras to continuously capture complete swing image sequences and combine them with simple time slicing, threshold judgment, or single human posture estimation algorithms for analysis. However, these solutions usually only roughly divide the swing phases based on single modal features, lacking comprehensive utilization of golf-specific movement characteristics such as club trajectory and changes in body center of gravity. Key moments still require manual screening and confirmation. The large amount of image data generated by high-frame-rate acquisition is often stored and played back in its raw form, resulting in high data storage, transmission, and processing costs. Furthermore, there is no specific evaluation and screening of image quality for key moments, making it difficult to achieve effective data compression and simplified video generation while ensuring clear presentation of key information. Summary of the Invention
[0003] The purpose of this application is to propose a method, system, computer equipment, and storage medium for selecting key nodes of a golf swing and generating video, so as to solve the technical problem of difficulty in achieving a balance between key node identification accuracy, analysis efficiency, and data utilization.
[0004] To address the aforementioned technical problems, this application provides a method for selecting key moments in a golf swing and generating video, employing the following technical solution: A camera with a frame rate no lower than a preset high frame threshold is used to capture the golfer's swing process, and a high frame image sequence covering the entire swing action is obtained; The high-frame image sequence is input into a pre-trained recognition model to extract human posture features, club trajectory features, and center of gravity change features arranged in chronological order. Each feature is normalized and fused according to a preset weight to obtain a fused feature sequence for swing event determination. Based on the fused feature sequence and the preset swing key node determination rules, multiple swing key nodes, including the ready posture, take-off, top of the backswing, downswing, moment of impact, follow-through and finish, are automatically identified in the high frame image sequence, and candidate key node images are extracted from several frames adjacent to each swing key node. Calculate image quality evaluation index for each candidate key node image, and select representative images of key nodes that meet preset quality conditions based on the image quality evaluation index. The images representing each key node are combined in chronological order to form a simplified swing video. The simplified swing video is then encoded and compressed to obtain a golf swing key node video that maintains a predetermined image quality at each key swing node and has a lower data volume than the high frame image sequence.
[0005] To address the aforementioned technical problems, this application also provides a golf swing key node capture and video generation system, which employs the following technical solution: The acquisition module is configured to use a camera with a frame rate not lower than a preset high frame threshold to capture the golfer's swing process and acquire a high frame image sequence covering the entire swing action process; The input module is configured to input the high-frame image sequence into a pre-trained recognition model, extract human posture features, club trajectory features and center of gravity change features arranged in chronological order, normalize each feature and fuse them according to preset weights to obtain a fused feature sequence for swing event determination. The recognition module is configured to automatically identify multiple key swing nodes, including the ready posture, take-off, top of backswing, downswing, moment of impact, follow-through, and finish, in the high frame image sequence based on the fused feature sequence and preset key swing node determination rules, and to extract candidate key node images from several frames adjacent to each key swing node. The selection module is configured to calculate image quality evaluation indicators for each candidate key node image, and select representative images of key nodes that meet preset quality conditions based on the image quality evaluation indicators. The compression module is configured to combine representative images of each key node in chronological order into a simplified swing video, and to encode and compress the simplified swing video to obtain a golf swing key node video that maintains a predetermined image quality at each key swing node and has a data volume lower than that of the high frame image sequence.
[0006] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the golf swing key node selection and video generation method as described above.
[0007] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the golf swing key node selection and video generation method described above.
[0008] Compared with the prior art, the embodiments of this application have the following main advantages: The golf swing key node selection and video generation method disclosed in this application, based on a high frame image sequence, jointly extracts human posture features, club trajectory features, and center of gravity change features, and automatically determines multiple key swing nodes such as the ready posture, take-off, top of the backswing, downswing, moment of impact, follow-through, and finish after weighted fusion. Then, it combines image quality evaluation to select representative images of key nodes and generate a simplified swing video that is encoded and compressed. This allows for automatic and stable key node identification and image selection even when a single swing generates a large amount of high frame raw data. It significantly reduces the amount of data while maintaining the clarity and analytical value of key node images, thereby improving the accuracy and efficiency of swing analysis and making it suitable for use in teaching, training, and remote transmission scenarios. Attached Figure Description
[0009] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart of an embodiment of the golf swing key point selection and video generation method according to this application; Figure 2 This is a schematic diagram of an embodiment of the golf swing key node selection and video generation system according to this application; Figure 3 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0012] refer to Figure 1 The diagram illustrates a flowchart of an embodiment of the golf swing key point selection and video generation method according to this application. The golf swing key point selection and video generation method includes the following steps: Step S101: Use a camera with a frame rate not lower than a preset high frame threshold to capture the golfer's swing process and obtain a high frame image sequence that covers the entire swing action.
[0013] In this embodiment, the electronic device running on the golf swing key node capture and video generation method can send or receive data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future known wireless connection methods.
[0014] In this embodiment, during the golf swing acquisition phase, a camera with a frame rate not lower than a preset high frame threshold captures the golfer's swing process, obtaining a high-frame image sequence covering the entire swing motion. In a preferred embodiment, the camera is positioned to the side of the golfer to capture the swing process from the side, ensuring complete capture of the backswing, downswing, and the change in the angle between the club and the ground at the moment of impact. Alternatively, the camera can be positioned directly in front of the golfer to capture the swing from the front, allowing for clearer observation of club deviation in the target direction and body sway. Two cameras can also be positioned at the front and side respectively, simultaneously acquiring high-frame image sequences from both perspectives, which are then processed separately or jointly by the recognition model. This invention does not limit the specific camera perspective; as long as a high-frame image sequence covering the entire swing motion can be acquired, subsequent feature extraction and key node recognition can be achieved. High-frame-rate image sequences refer to continuous image frame sequences with a frame rate of at least several hundred frames per second. For example, in one embodiment, the camera frame rate is configured to be 480-960fps and the resolution is 1920×1080 to 3840×2160. It is positioned to the side of the player (e.g., at approximately 45° to the right, at a distance of approximately 6.5m, and at a height of approximately 1.8m) to ensure that the entire swing motion from the ready position to the follow-through position is continuously and clearly recorded within the field of view. This allows for the acquisition of thousands of high-resolution images within the 2.8-3.2 seconds of a single swing, providing a sufficient data foundation for subsequent frame-level detail-based analysis.
[0015] Step S102: Input the high frame image sequence into the pre-trained recognition model, extract human posture features, club trajectory features and center of gravity change features arranged in chronological order, normalize each feature and fuse them according to preset weights to obtain a fused feature sequence for swing event determination.
[0016] In this embodiment, after obtaining the high-frame image sequence, the image sequence is input into a pre-trained recognition model. The recognition model can be a deep learning model composed of sub-networks such as target detection, human posture estimation, deep feature extraction, and temporal analysis, used to automatically extract multi-dimensional motion features of the golfer and club from the images. Among them, human posture features can be understood as feature vectors representing the positions of key joints of the golfer's body. For example, a posture estimation algorithm is used to extract the two-dimensional or three-dimensional coordinates of 18 key points such as the head, shoulder, elbow, wrist, hip, knee, and ankle joints from each frame of the image. Club trajectory features can be understood as a feature sequence representing the spatial trajectory and angle changes of the clubhead, shaft, and clubface in consecutive frames. For example, the position of the clubhead is recorded at a sampling frequency of approximately 480Hz, forming a club motion curve on the time axis. Center of gravity change features represent the distribution changes of the golfer's body mass on the support surface. For example, by the pressure distribution of the two feet or the center of gravity inference based on posture estimation, a quantitative curve of the forward and backward and left and right movement of the center of gravity during the swing is obtained. After outputting these features arranged in chronological order, the recognition model normalizes the various features to make different physical quantities comparable in numerical scale. Then, it performs weighted fusion according to preset weights (e.g., 0.4 for human posture features, 0.35 for club trajectory features, and 0.25 for center of gravity change features) to form a fused feature sequence for swing event determination, thereby simultaneously reflecting the comprehensive state of body movement, club movement, and center of gravity transfer on a single time axis.
[0017] Step S103: Based on the fused feature sequence and the preset swing key node determination rules, automatically identify multiple swing key nodes in the high frame image sequence, including the ready posture, take-off, top of the backswing, downswing, moment of impact, follow-through and finish-out, and extract candidate key node images from several frames adjacent to each swing key node.
[0018] In this embodiment, after obtaining the fused feature sequence, the system automatically identifies multiple key swing nodes in the high-frame image sequence based on preset swing key node determination rules, including the ready posture, take-off, top of the backswing, downswing, impact, follow-through, and finish. The swing key node determination rules quantify the determination conditions for each stage in the fused feature space. For example, the ready posture can be identified by the club-ground angle being within a predetermined range and the body posture remaining stable over several frames; the take-off can be determined by the club movement angle change rate exceeding a certain threshold accompanied by a change in wrist angle; the top of the backswing can be located by the club-ground parallelism approaching its extreme value and the shoulder rotation angle reaching its maximum; the downswing can be identified by the club angular velocity exceeding a preset threshold combined with a significant shift of weight from the back foot to the front foot; the impact can be accurately located by the clubhead linear velocity reaching its peak and identifying the ball-clutch face contact characteristics; the follow-through can be determined by the continuity of club movement and the elbow angle exceeding a certain threshold; and the finish posture can be identified by the body's center of gravity stabilizing in the front and the club position remaining basically stable. In practice, the temporal analysis subnetwork can output the probability of each frame belonging to the aforementioned key nodes. The system finds the probability peak on the time axis as the representative time of each key node, and then extracts candidate key node images from several frames adjacent to each key node (e.g., a small time window composed of several frames before and after) to ensure that there is a certain redundancy in subsequent selection, so that the frame with the best quality can be selected from multiple images at similar times.
[0019] Step S104: Calculate image quality evaluation index for each candidate key node image, and select representative key node images that meet preset quality conditions based on the image quality evaluation index.
[0020] In this embodiment, for each candidate key node image obtained at each key node, the system further calculates image quality evaluation indicators for each candidate image. These indicators can include quantifiable metrics such as sharpness, motion blur, and exposure. For example, image sharpness can be measured using the variance of the Laplacian operator, motion blur can be measured using the edge strength of Sobel edge detection, and exposure can be measured by converting the image to the HSV color space and statistically analyzing the brightness channel and saturation distribution. In this embodiment, sharpness thresholds, blur thresholds, and effective brightness ratio thresholds can be set. The above indicators are calculated for each candidate key node image, and only images that simultaneously meet the preset quality conditions are retained as the representative images of that key node. For example, for the moment of impact, the system may extract more than ten candidate images from several frames before and after that moment. After the above quality evaluation, only one or two image frames with optimal sharpness, blur, and exposure are retained. This ensures that the outlines of the club and ball are clear at the moment of impact and avoids images that are blurry, too dark, or overexposed, affecting technical analysis, from entering subsequent videos.
[0021] Step S105: Combine the representative images of each key node in chronological order into a simplified swing video, and encode and compress the simplified swing video to obtain a golf swing key node video that maintains a predetermined image quality at each key swing node and has a data volume lower than that of the high frame image sequence.
[0022] In this embodiment, after obtaining the representative images of key nodes corresponding to each key swing node, the system combines these representative images into a simplified swing video in chronological order. To ensure viewing continuity and information integrity, the combination can be understood as sorting according to the timestamps of the key nodes and inserting appropriate transition frames between adjacent representative images or using interpolation algorithms to generate smooth transition images. For example, cubic Bézier curve interpolation is used to smooth the visual changes between adjacent key nodes. At the same time, different display durations are allocated according to the importance of each node, so that key nodes such as the moment of impact have a slightly longer time in the final video, while nodes such as the preparation posture and follow-through have a relatively shorter time. This ensures that key information is fully presented within a limited total duration (e.g., 3-4 seconds). Subsequently, the generated simplified swing video is encoded and compressed. Encoding and compression can adopt existing high-efficiency video coding standards (such as H.265 / HEVC). By setting appropriate bitrate control strategies and entropy coding methods (such as CABAC), the video data volume is significantly reduced while maintaining the image clarity of key nodes as much as possible. In one embodiment, the original high-frame image data of a single swing can reach several gigabytes, while the compressed key node video file can be controlled to the order of several megabytes. This ensures that the golf swing key node video maintains the predetermined image quality at each key swing node while significantly reducing the data volume compared to the original high-frame image sequence, making it convenient for storage, transmission, and playback on teaching terminals, mobile devices, or the cloud.
[0023] In one specific implementation, when the system selects representative images of key nodes for each swing critical point and constructs a simplified swing video, it can determine the total playback duration of the simplified swing video according to the action duration corresponding to the original high-frame-rate image sequence. This ensures that the relative positions of key nodes such as the take-off, top of the backswing, and impact moment in the simplified swing video on the timeline are basically consistent with the actual swing process. While keeping the total duration unchanged, non-critical frames in the original high-frame-rate image sequence are downsampled, retaining only necessary frames for action transitions near each key swing node. This reduces the effective frame count of the entire video from the order of magnitude of high frames to the number of frames at a lower frame rate, for example, from hundreds of frames per second to tens of frames per second. Furthermore, the remaining frames are compressed using efficient video encoding methods, resulting in a video file size significantly smaller than the original high-frame-rate image data. Because representative images of each key swing node and their adjacent transition frames are preferentially retained during downsampling, the resulting simplified swing video achieves a significant reduction in frame rate and data volume while maintaining a relatively constant playback duration and complete key node footage. This reduces storage and transmission costs while preserving the original swing rhythm and visual appeal. This application, based on high-frame-rate image sequences, jointly extracts human posture features, club trajectory features, and center-of-gravity change features. After weighted fusion, it automatically determines multiple key swing nodes, including the setup posture, take-off, backswing peak, downswing, impact moment, follow-through, and finish. Combined with image quality evaluation, representative images of key nodes are selected, and a coded and compressed simplified swing video is generated. This allows for automatic and stable key node identification and image selection even with a large amount of high-frame-rate raw data generated per swing. It significantly reduces data volume while maintaining the clarity and analytical value of key node images, thereby improving the accuracy and efficiency of swing analysis and facilitating its use in teaching, training, and remote transmission scenarios.
[0024] In some optional implementations of this embodiment, the above-mentioned identification model includes: A target detection subnetwork for detecting the golfer region and club region in the high-frame image sequence; A human pose estimation subnetwork for outputting a sequence of human key point coordinates based on the player's region; A temporal analysis sub-network is used to model the human posture features, the club trajectory features, and the center of gravity change features in the time dimension and output a temporal feature sequence. The temporal feature sequence is used to form the fused feature sequence for swing event determination.
[0025] In this embodiment, the object detection subnetwork is used to automatically locate the golfer region and club region in each frame of the high-frame image sequence. It can output candidate boxes or segmented regions surrounding the golfer's body and club, providing spatial constraints for subsequent detailed analysis. The human pose estimation subnetwork takes the golfer region output by the object detection subnetwork as input and extracts the coordinate sequences of multiple human key points in the region, such as the two-dimensional or three-dimensional coordinates of key joints such as the head, shoulders, elbows, hips, and knees, thereby forming human pose features to describe the swing action. The temporal analysis subnetwork models the human pose features, club trajectory features, and center of gravity change features in the time dimension. It can be understood as a network module that models the feature sequence of continuous frames, such as a recurrent neural network or a temporal network based on an attention mechanism, to output a temporal feature sequence representing the evolution of the action over time. This temporal feature sequence is not used as the final output alone, but participates in the formation of a fusion feature sequence for swing event determination. That is, during the fusion process, static attitude / trajectory / center of gravity information and dynamic temporal change trends are considered simultaneously. This allows key node identification to focus on the instantaneous attitude of a certain frame, while also comprehensively considering the changes in several frames before and after, thereby improving the robustness and accuracy of key node determination.
[0026] This application defines the recognition model as consisting of three parts: a target detection subnetwork, a human pose estimation subnetwork, and a temporal analysis subnetwork. This allows for the accurate identification of the golfer and club areas first, followed by the extraction of key points and modeling in the time dimension. This improves the accuracy of multimodal feature extraction and temporal variation characterization, and enhances the reliability of using fused features for key node determination.
[0027] In some optional implementations of this embodiment, the steps of extracting human posture features, club trajectory features, and center of gravity change features arranged in chronological order include: Based on the target detection subnetwork in each frame, the spatial position sequence of the club head in adjacent frames is extracted, and the club trajectory features and club angular velocity features are calculated based on the spatial position sequence. Based on the pressure distribution or acceleration data output by the center of gravity acquisition device deployed under the player's feet or on the body, calculate the characteristics of the player's center of gravity change in the left-right and front-back directions. The cue angular velocity characteristics and the center of gravity change characteristics are then input into the time-series analysis subnetwork.
[0028] In this embodiment, on the one hand, based on the club region output by the target detection sub-network in each frame, the spatial position of the clubhead can be extracted in each frame. For example, by finding feature points at the end of the clubhead within the club region, a time series of clubhead coordinates between adjacent frames can be obtained. Then, the club trajectory features and club angular velocity features are calculated based on this spatial position sequence. For example, parameters such as linear velocity and rate of change of angle are obtained by differentiating the position sequence. This can more accurately reflect the speed and direction changes of the club during the backswing, downswing, and impact. On the other hand, the center of gravity change features are not directly inferred from the image, but can be calculated based on the pressure distribution or acceleration data obtained by the center of gravity acquisition device placed under the golfer's feet or on their body. For example, the left-right movement of the center of gravity can be estimated by obtaining the force ratio of the left and right feet through the pressure plate, and the forward-backward movement of the center of gravity can be estimated by obtaining the acceleration and posture changes through the inertial measurement unit worn on the body. Thus, the center of gravity change curves of the golfer in the left-right and forward-backward directions are obtained. By inputting the aforementioned club angular velocity characteristics and center of gravity change characteristics along with human posture characteristics into the temporal analysis subnetwork, the temporal analysis is not only based on joint position changes, but also jointly considers club movement velocity and center of gravity transfer, thereby better characterizing the dynamic characteristics of the swing action. For example, during the downswing phase, the club angular velocity increases rapidly and the center of gravity clearly shifts from the back foot to the front foot. This multimodal input is beneficial for the accurate identification of subsequent key nodes.
[0029] This application calculates the club trajectory and club angular velocity from the clubhead position sequence in the club area, and calculates the center of gravity change from pressure or acceleration data, and feeds them into the time series analysis sub-network. This allows the swing characteristics to simultaneously reflect the motion form and dynamic information, more sensitively reflect the power rhythm and center of gravity transfer, and helps to distinguish key stages such as backswing, downswing and impact.
[0030] In some optional implementations of this embodiment, in the step of automatically identifying multiple swing key nodes, including the ready stance, take-off, backswing peak, downswing, impact moment, follow-through, and finish, in the high-frame image sequence based on the fused feature sequence and preset swing key node determination rules, and extracting candidate key node images from several frames adjacent to each swing key node, the swing key node determination rules include: When the shoulder rotation angle obtained from the human posture features reaches the first extreme value, the corresponding time point is determined as the top node of the pole. When the clubhead height obtained from the club trajectory characteristics changes from increasing to decreasing and the angle between the club and the ground changes from increasing to decreasing, the corresponding time point is determined as the start point or the downswing point. When the spatial distance between the clubhead and the golf ball obtained from the fused feature sequence is less than a preset distance threshold and the contact feature between the ball and the clubface meets the judgment condition, the corresponding time point is determined as the moment of impact.
[0031] In this embodiment, the top of the backswing can be determined by the change in the shoulder rotation angle in the human posture features. When this angle increases over time and reaches the first extreme value, it indicates that the golfer has completed the rotational motion from the start to the top of the backswing, and the corresponding time point can be used as the top of the backswing node. The start and downswing nodes can be determined by the combination of changes in clubhead height and the angle between the club and the ground in the club trajectory features. For example, during the backswing phase, the clubhead height continuously increases and the angle between the club and the ground gradually increases. When the trend of clubhead height change from increasing to decreasing and the angle between the club and the ground changes, the top of the backswing node is determined. When the angle changes from increasing to decreasing, the corresponding time point can be identified as the transition node from the backswing to the downswing. By combining the changes before and after the movement, the take-off node and the downswing node can be further distinguished. The determination of the moment of impact node depends on the spatial distance and contact characteristics between the clubhead and the golf ball in the fused feature sequence. When the spatial distance between the clubhead and the ball is less than a preset distance threshold and the features representing the contact between the ball and the clubface in the image or sensor data (such as the deformation of the ball, the brightness change of the contact area between the ball and the clubface, etc.) meet the determination conditions, the corresponding time point is determined as the moment of impact node. Through this quantitative rule based on angle extremes, trend changes, and distance thresholds, these key nodes that are crucial to technical analysis can be accurately found in high-frame sequences, and candidate key node images can be extracted from several adjacent frames, providing accurate temporal positioning for subsequent quality screening and video generation.
[0032] This application refines the key node determination rules into quantitative conditions such as the extreme value of the shoulder rotation angle, the changing trend of the clubhead height and the angle between the club and the ground, and the spatial distance and contact characteristics between the clubhead and the golf ball. This transforms the timing of key nodes such as the top of the backswing, the take-off or downswing, and the moment of impact from experience-based judgment into calculable rules, making key node identification more accurate and stable.
[0033] In some optional implementations of this embodiment, the step of extracting candidate key node images from several frames adjacent to each of the said key swing nodes includes: For each swing key node, taking the time point corresponding to the swing key node as the center, select no less than two and no more than a preset upper limit number of consecutive frames as candidate key node images within a preset time window, and limit the time deviation of the timestamp of each candidate key node image relative to the time point of the corresponding swing key node to a preset range.
[0034] In this embodiment, for each key swing node, a symmetrical or asymmetrical preset time window is set on the timeline, centered on the time point corresponding to that node. This window typically contains a range of several frames before and after the node. Within this time window, at least two consecutive frames, but no more than a preset upper limit, are selected as candidate key node images. This ensures that the candidate images are sufficiently close to the node in time to reflect the instantaneous changes in the key movement, while limiting the upper limit avoids excessive redundant data and increased processing burden. Simultaneously, the time deviation of the timestamps of each candidate key node image relative to the corresponding key swing node time point is limited to a preset range. This can be understood as a further constraint on the time window, such as specifying that the deviation does not exceed a certain number of milliseconds or frames, to prevent the candidate images from deviating too far from the key node, causing the final representative image to inaccurately reflect the true movement pattern of that node.
[0035] This application sets a time window around each key swing node time point, limits the number of consecutive frames captured and constrains the time deviation range, ensuring that each node can obtain a moderate number of candidate key node images with time positions close to the actual movement, thus avoiding redundant data and improving the reliability of subsequent selection of the best representative image.
[0036] In some optional implementations of this embodiment, in the steps of calculating image quality evaluation indicators for each of the candidate key node images and selecting representative images of key nodes that meet preset quality conditions based on the image quality evaluation indicators, the image quality evaluation indicators include: A sharpness index calculated based on the Laplacian operator or a high-pass filter is used to characterize the sharpness of image edges. Motion blur indexes calculated based on edge gradient distribution or edge expansion are used to characterize the degree of blurring of the club and player silhouettes. Exposure metrics, calculated based on luminance histograms or luminance mean and variance, are used to characterize image brightness and contrast. The step of selecting key nodes representing images that meet preset quality conditions based on the image quality evaluation index includes: The sharpness index, motion blur index, and exposure index are combined into a quality score according to a preset weight, and the quality score is compared with a quality threshold to determine whether to select the corresponding candidate key node image.
[0037] In this embodiment, the sharpness index can be quantified by applying a Laplacian operator or other high-pass filters to the candidate image and statistically analyzing the variance or energy of the filtering results. A higher value generally means sharper edges. The motion blur index can be calculated based on the edge gradient distribution or edge extension in the image. For example, by analyzing the gradient magnitude and gradient direction changes near the outline of the golf club and player, if there is obvious trailing or edge extension, the motion blur index is larger, which is used to characterize the degree of ghosting. The exposure index can be calculated based on the brightness histogram or the brightness mean and variance. For example, by statistically analyzing the distribution of pixel grayscale in the image, if the brightness is too low or too high, the histogram will be biased towards the dark or bright side, and the brightness variance may also deviate from a reasonable range, thus reflecting underexposure or overexposure problems. Instead of simply setting thresholds for each of the three indicators separately, this invention synthesizes the sharpness, motion blur, and exposure indicators into a unified quality score based on preset weights. For example, it increases the weight of the sharpness indicator and decreases the weight of the motion blur indicator, making the score more focused on images with clear edges and no obvious motion blur. This quality score is then compared with a preset quality threshold. When the quality score is not lower than the threshold, the corresponding candidate key node image is selected as the representative image. When multiple candidate images simultaneously meet the quality conditions, the image with the highest score can be selected as the representative image for that key node. This comprehensive scoring and threshold determination method allows different quality factors to work synergistically, avoiding misjudgments caused by relying on a single indicator, thereby ensuring that the generated video has sufficient sharpness and visual appeal at each key node.
[0038] This application constructs an image quality evaluation system by combining sharpness, motion blur, and exposure metrics, and compares the synthesized quality score with a quality threshold according to preset weights. This enables comprehensive screening of candidate key node images, effectively eliminating blurry, ghosting, or abnormally exposed image frames, and significantly improving the sharpness and technical analysis value of representative images of key nodes.
[0039] In some optional implementations of this embodiment, in the steps of combining the representative images of each key node in chronological order into a simplified swing video and encoding and compressing the simplified swing video, corresponding key node display strategies and encoding and compression parameters are set according to different application modes, wherein: In the teaching mode, the display duration of the moment of impact and the top of the backswing is longer than that of other key swing nodes, and the joint angle annotations calculated from the human posture features are superimposed on the simplified swing video. In the match replay mode, the target bit rate of the simplified swing video is increased to ensure the image quality of the moment of impact, while retaining all the key swing moments; In mobile sharing mode, the total duration and output file size of the simplified swing video are limited, and the number of key swing nodes selected is reduced accordingly.
[0040] In this embodiment, the concept of application mode is introduced, and corresponding key node display strategies and encoding compression parameters are set according to different application scenarios, so that the same set of key node recognition and video generation process can adapt to different usage needs such as teaching, competition review and mobile sharing.
[0041] In teaching mode, the system can allocate a longer display duration for the moment of impact and the top of the backswing than for other key swing points. For example, in a three-second condensed video, hundreds of milliseconds or even several seconds of slow motion can be allocated for the moment of impact, and a relatively long dwell time can be allocated for the top of the backswing. Joint angle annotations calculated from human posture characteristics can be superimposed on the screen corresponding to these points, such as displaying values and auxiliary lines for shoulder rotation angle, hip rotation angle, and wrist flexion angle, so that coaches and students can intuitively observe whether the posture has reached the expected level.
[0042] In the game review mode, the system aims to help golfers review key swing actions. Therefore, it can improve the target bit rate or encoding quality level of the simplified swing video, especially focusing on ensuring the image quality corresponding to the moment of impact, and retaining all key swing nodes, presenting the complete swing rhythm from the preparation position to the follow-through, so as to conduct a comprehensive analysis of rhythm, power transfer and follow-through balance.
[0043] In mobile sharing mode, considering the limitations of mobile network bandwidth, terminal storage space and viewing time, the system can set upper limits on the total duration and output file size of the simplified swing video. For example, the total duration can be limited to a few seconds and the file size to a few MB. The number of key swing nodes selected can be reduced accordingly, and only the most critical nodes such as the moment of impact and the top of the backswing can be retained to obtain a short video that is easy to share quickly and can reflect the characteristics of the action.
[0044] This application differentiates between teaching mode, competition review mode, and mobile sharing mode during the streamlined swing video generation and encoding compression process. It adjusts the display duration, overlay annotation, encoding bit rate, and number of retained nodes for key nodes respectively, so that the same technical solution can flexibly balance image quality, information content, and file size according to specific application needs, thereby improving the applicability and resource utilization efficiency of the method in different use scenarios.
[0045] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0046] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0047] Further reference Figure 2 As a response to the above Figure 1 The implementation of the method shown in this application provides an embodiment of a golf swing key node capture and video generation system. This system embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.
[0048] like Figure 2 As shown, the golf swing key node capture and video generation system 200 described in this embodiment includes: an acquisition module 201, an input module 202, a recognition module 203, a selection module 204, and a compression module 205. Wherein: The acquisition module 201 is configured to use a camera with a frame rate not lower than a preset high frame threshold to capture the golfer's swing process and acquire a high frame image sequence covering the entire swing action process; Input module 202 is configured to input the high frame image sequence into a pre-trained recognition model, extract human posture features, club trajectory features and center of gravity change features arranged in chronological order, normalize each feature and fuse them according to preset weights to obtain a fused feature sequence for swing event determination; The recognition module 203 is configured to automatically identify multiple swing key nodes, including the ready posture, take-off, top of the backswing, downswing, moment of impact, follow-through and finish, in the high frame image sequence based on the fused feature sequence and the preset swing key node determination rules, and to extract candidate key node images from several frames adjacent to each swing key node. The selection module 204 is configured to calculate image quality evaluation index for each candidate key node image, and select representative images of key nodes that meet preset quality conditions based on the image quality evaluation index. Compression module 205 is configured to combine representative images of each key node in chronological order into a simplified swing video, and to encode and compress the simplified swing video to obtain a golf swing key node video that maintains a predetermined image quality at each swing key node and has a data volume lower than that of the high frame image sequence.
[0049] The golf swing key node selection and video generation system provided in this embodiment of the invention can realize all the processes of the golf swing key node selection and video generation method in the above embodiments. The functions and technical effects of each module in the device are the same as those of the golf swing key node selection and video generation method in the above embodiments, and will not be repeated here.
[0050] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0051] The computer device 3 includes a memory 31, a processor 32, and a network interface 33 that are interconnected via a system bus. It should be noted that only the computer device 3 with components 31-33 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0052] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0053] The memory 31 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Of course, the memory 31 may also include both the internal storage unit and its external storage device of the computer device 3. In this embodiment, the memory 31 is typically used to store the operating system and various application software installed on the computer device 3, such as computer-readable instructions for a method of capturing key points of a golf swing and generating video. In addition, the memory 31 can also be used to temporarily store various types of data that have been output or will be output.
[0054] In some embodiments, the processor 32 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 32 is typically used to control the overall operation of the computer device 3. In this embodiment, the processor 32 is used to execute computer-readable instructions stored in the memory 31 or to process data, for example, to execute computer-readable instructions for the golf swing key node capture and video generation method.
[0055] The network interface 33 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 3 and other electronic devices.
[0056] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the golf swing key node selection and video generation method described above.
[0057] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0058] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for selecting key moments of a golf swing and generating video, characterized in that, Includes the following steps: A camera with a frame rate no lower than a preset high frame threshold is used to capture the golfer's swing process, and a high frame image sequence covering the entire swing action is obtained; The high-frame image sequence is input into a pre-trained recognition model to extract human posture features, club trajectory features, and center of gravity change features arranged in chronological order. Each feature is normalized and fused according to a preset weight to obtain a fused feature sequence for swing event determination. Based on the fused feature sequence and the preset swing key node determination rules, multiple swing key nodes, including the ready posture, take-off, top of the backswing, downswing, moment of impact, follow-through and finish, are automatically identified in the high frame image sequence, and candidate key node images are extracted from several frames adjacent to each swing key node. Calculate image quality evaluation index for each candidate key node image, and select representative images of key nodes that meet preset quality conditions based on the image quality evaluation index. The images representing each key node are combined in chronological order to form a simplified swing video. The simplified swing video is then encoded and compressed to obtain a golf swing key node video that maintains a predetermined image quality at each key swing node and has a lower data volume than the high frame image sequence.
2. The method according to claim 1, characterized in that, The recognition model includes: A target detection subnetwork for detecting the golfer region and club region in the high-frame image sequence; A human pose estimation subnetwork for outputting a sequence of human key point coordinates based on the player's region; A temporal analysis sub-network is used to model the human posture features, the club trajectory features, and the center of gravity change features in the time dimension and output a temporal feature sequence. The temporal feature sequence is used to form the fused feature sequence for swing event determination.
3. The method according to claim 2, characterized in that, The steps for extracting human posture features, club trajectory features, and center of gravity change features arranged in chronological order include: Based on the target detection subnetwork in each frame, the spatial position sequence of the club head in adjacent frames is extracted, and the club trajectory features and club angular velocity features are calculated based on the spatial position sequence. Based on the pressure distribution or acceleration data output by the center of gravity acquisition device deployed under the player's feet or on the body, calculate the characteristics of the player's center of gravity change in the left-right and front-back directions. The cue angular velocity characteristics and the center of gravity change characteristics are then input into the time-series analysis subnetwork.
4. The method according to claim 1, characterized in that, In the step of automatically identifying multiple key swing nodes, including the ready position, take-off, top of the backswing, downswing, moment of impact, follow-through, and finish, in the high-frame image sequence based on the fused feature sequence and preset key swing node determination rules, and extracting candidate key node images from several frames adjacent to each key swing node, the key swing node determination rules include: When the shoulder rotation angle obtained from the human posture features reaches the first extreme value, the corresponding time point is determined as the top node of the pole. When the clubhead height obtained from the club trajectory characteristics changes from increasing to decreasing and the angle between the club and the ground changes from increasing to decreasing, the corresponding time point is determined as the start point or the downswing point. When the spatial distance between the clubhead and the golf ball obtained from the fused feature sequence is less than a preset distance threshold and the contact feature between the ball and the clubface meets the judgment condition, the corresponding time point is determined as the moment of impact.
5. The method according to claim 4, characterized in that, The step of extracting candidate key node images from several frames adjacent to each of the key swing nodes includes: For each swing key node, taking the time point corresponding to the swing key node as the center, select no less than two and no more than a preset upper limit number of consecutive frames as candidate key node images within a preset time window, and limit the time deviation of the timestamp of each candidate key node image relative to the time point of the corresponding swing key node to a preset range.
6. The method according to claim 1, characterized in that, In the step of calculating image quality evaluation indicators for each of the candidate key node images and selecting representative images of key nodes that meet preset quality conditions based on the image quality evaluation indicators, the image quality evaluation indicators include: A sharpness index calculated based on the Laplacian operator or a high-pass filter is used to characterize the sharpness of image edges. Motion blur indexes calculated based on edge gradient distribution or edge expansion are used to characterize the degree of blurring of the club and player silhouettes. Exposure metrics, calculated based on luminance histograms or luminance mean and variance, are used to characterize image brightness and contrast. The step of selecting key nodes representing images that meet preset quality conditions based on the image quality evaluation index includes: The sharpness index, motion blur index, and exposure index are combined into a quality score according to a preset weight, and the quality score is compared with a quality threshold to determine whether to select the corresponding candidate key node image.
7. The method according to claim 1, characterized in that, In the step of combining the representative images of each key node in chronological order into a simplified swing video and encoding and compressing the simplified swing video, corresponding key node display strategies and encoding and compression parameters are set according to different application modes, wherein: In the teaching mode, the display duration of the moment of impact and the top of the backswing is longer than that of other key swing nodes, and the joint angle annotations calculated from the human posture features are superimposed on the simplified swing video. In the match replay mode, the target bit rate of the simplified swing video is increased to ensure the image quality of the moment of impact, while retaining all the key swing moments; In mobile sharing mode, the total duration and output file size of the simplified swing video are limited, and the number of key swing nodes selected is reduced accordingly.
8. A system for capturing and generating video of key moments in a golf swing, characterized in that, include: The acquisition module is configured to use a camera with a frame rate not lower than a preset high frame threshold to capture the golfer's swing process and acquire a high frame image sequence covering the entire swing action process; The input module is configured to input the high-frame image sequence into a pre-trained recognition model, extract human posture features, club trajectory features and center of gravity change features arranged in chronological order, normalize each feature and fuse them according to preset weights to obtain a fused feature sequence for swing event determination. The recognition module is configured to automatically identify multiple key swing nodes, including the ready posture, take-off, top of backswing, downswing, moment of impact, follow-through, and finish, in the high frame image sequence based on the fused feature sequence and preset key swing node determination rules, and to extract candidate key node images from several frames adjacent to each key swing node. The selection module is configured to calculate image quality evaluation indicators for each candidate key node image, and select representative images of key nodes that meet preset quality conditions based on the image quality evaluation indicators. The compression module is configured to combine representative images of each key node in chronological order into a simplified swing video, and to encode and compress the simplified swing video to obtain a golf swing key node video that maintains a predetermined image quality at each key swing node and has a data volume lower than that of the high frame image sequence.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the golf swing key node selection and video generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the golf swing key node selection and video generation method as described in any one of claims 1 to 7.