A swing motion segmentation method, system, device and medium

CN122244958BActive Publication Date: 2026-09-25SHANGHAI FUTURE MIND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610702193.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-25
Estimated Expiration
2046-05-21

AI Technical Summary

Technical Problem

[0003]本发明提供了一种挥拍动作分割方法、系统、设备及介质,以解决现有挥拍动作分割方法因采用固定规则而无法识别球员战术意图、导致不同场景下泛化能力不足的技术问题

Benefits of technology

[0009]根据本发明的另一方面,提供了一种计算机程序产品,所述计算机程序产品包括计算机程序,所述计算机程序在被处理器执行时实现本发明任一实施例所述的挥拍动作分割方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244958B_ABST
    Figure CN122244958B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and particularly discloses a swing action segmentation method, system, device and medium. The method comprises the following steps: performing downsampling on a target video to obtain a sampling video frame; extracting macro semantic features from the sampling video frame, and generating constraint information based on the macro semantic features; extracting multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information at least comprises human body skeleton key point coordinates, sphere coordinates and racket head coordinates; performing swing action segmentation on the target video based on the constraint information and the multi-target spatial coordinate information, and outputting time information corresponding to multiple action stages of the swing action. According to the scheme, the constraint information is generated by extracting the macro semantic features to dynamically guide the action segmentation process, so that the segmentation logic can be adaptively adjusted according to different tactical scenes, and therefore the accuracy and robustness of the action stage segmentation in a complex and changeable swing movement scene are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for segmenting racket swing motions. Background Technology

[0002] Existing methods for segmenting racket-based sports (such as tennis, badminton, and table tennis) typically use standard 3D human skeleton point extraction models to obtain the coordinates of key points on the human skeleton. Then, they rely on hard-coded heuristic rules written by the developers to determine state transitions. For example, a backward movement of the right wrist beyond a certain distance is considered a charging phase; a wrist speed reaching its extreme value is considered a contact phase. Because these methods use hard-coded rules with fixed thresholds for action segmentation, they cannot recognize the player's tactical intentions (such as topspin, slice, volley, etc.). Using the same set of criteria for all scenarios results in severely insufficient generalization ability across different players and tactics, leading to poor results in racket swing action segmentation. Summary of the Invention

[0003] This invention provides a method, system, device, and medium for racket swing action segmentation, in order to solve the technical problem that existing racket swing action segmentation methods, due to their use of fixed rules, cannot identify the player's tactical intentions and thus have insufficient generalization ability in different scenarios.

[0004] According to one aspect of the present invention, a method for segmenting a racket swing motion is provided, the method comprising: The target video is downsampled to obtain sampled video frames; Macro-level semantic features are extracted from sampled video frames, and constraint information is generated based on these macro-level semantic features; Extract multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head; Based on constraint information and multi-target spatial coordinate information, the target video is segmented into swing motions, and the time information corresponding to multiple action stages of the swing motion is output.

[0005] According to another aspect of the present invention, a swing motion segmentation system is provided for executing the swing motion segmentation method described in any embodiment of the present invention. The system includes a first subsystem and a second subsystem that operate asynchronously, wherein... The first subsystem is used to downsample the target video to obtain sampled video frames, extract macroscopic semantic features from the sampled video frames, generate constraint information based on the macroscopic semantic features, and send the constraint information to the second subsystem. The second subsystem, connected to the first subsystem, is used to receive constraint information, extract multi-target spatial coordinate information from each video frame of the target video, and perform swing motion segmentation on the target video based on the constraint information and multi-target spatial coordinate information, and output the time information corresponding to multiple action stages of the swing motion; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head.

[0006] According to another aspect of the present invention, a swing motion segmentation device is provided, the device comprising: The downsampling module is used to downsample the target video to obtain sampled video frames; The semantic extraction module is used to extract macroscopic semantic features from sampled video frames and generate constraint information based on the macroscopic semantic features; The coordinate extraction module is used to extract multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head; The action segmentation module is used to segment the swing action of the target video based on constraint information and multi-target spatial coordinate information, and output the time information corresponding to multiple action stages of the swing action.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the swing motion segmentation method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the swing motion segmentation method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the swing motion segmentation method according to any embodiment of the present invention.

[0010] The technical solution of this invention involves downsampling the target video to obtain sampled video frames; extracting macroscopic semantic features from the sampled video frames and generating constraint information based on the macroscopic semantic features; extracting multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head; and segmenting the target video into swing motions based on the constraint information and the multi-target spatial coordinate information, outputting the time information corresponding to multiple action stages of the swing motion. This technical solution extracts macroscopic semantic features from sampled video frames obtained by downsampling the target video and generates constraint information. Simultaneously, it extracts multi-target spatial coordinate information from each frame of the target video, including at least the coordinates of key points of the human skeleton, the coordinates of the ball, and the coordinates of the racket head. This enables the action segmentation process to have both a global understanding of the player's tactical intentions and the ability to accurately capture the spatial motion details of the ball and the racket. The constraint information, derived from macroscopic semantic features, is used to dynamically guide the action segmentation process based on multi-target spatial coordinate information. This allows the segmentation logic to adaptively adjust according to different scenarios, overcoming the shortcomings of traditional fixed-rule methods that use the same standard to deal with all scenarios, resulting in insufficient generalization ability. It significantly improves the accuracy and robustness of action segmentation in complex and varied swing motion scenarios.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a swing motion segmentation method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a swing motion segmentation method provided in Embodiment 2 of the present invention; Figure 3 This is a flowchart of a swing motion segmentation method provided in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of a swing motion segmentation system provided in Embodiment 4 of the present invention; Figure 5 This is a flowchart of a swing motion segmentation method provided in Embodiment 4 of the present invention; Figure 6 This is a flowchart of another swing motion segmentation method provided in Embodiment 4 of the present invention; Figure 7 This is a schematic diagram of a racket swing action segmentation device provided in Embodiment 5 of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device that implements the swing motion segmentation method of this invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart of a swing motion segmentation method provided in Embodiment 1 of the present invention. This embodiment is applicable to the precise segmentation of high-speed swing motion videos. The method can be executed by a swing motion segmentation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown in the figure, the swing motion segmentation method provided in this embodiment includes the following steps: S110. Downsample the target video to obtain sampled video frames.

[0017] The target video can refer to the original video frame sequence to be segmented for the swing action phase. It contains the complete process of a player completing one or more swing actions and has a fixed original frame rate. Downsampling refers to the process of extracting a portion of video frames from the original target video at a sampling density lower than the original frame rate according to preset downsampling rules. Its purpose is to reduce the amount of data for subsequent processing while preserving key visual information about macroscopic movement changes. Sampled video frames refer to the set of video frames extracted from the target video through downsampling, which is fewer than the total number of frames in the target video, and are used for extracting macroscopic semantic features.

[0018] In this embodiment of the invention, a subset of video frames can be extracted from the original target video to be analyzed according to a preset downsampling rule to obtain a set of sparse sampled video frames. The target video can originate from a continuous video stream from a camera, video file, streaming media server, etc., and this embodiment does not impose any restrictions on this.

[0019] S120. Extract macroscopic semantic features from the sampled video frames and generate constraint information based on the macroscopic semantic features.

[0020] Macro-level semantic features refer to semantic information extracted from video footage that describes the high-level tactical intent and overall motion attributes of a swing motion. Examples include, but are not limited to, action type, racket hand, hitting direction, stance, spin intention, on-court position, and player fitness level. Constraint information refers to control information generated based on macro-level semantic features to guide and regulate subsequent motion segmentation processes. Examples include, but are not limited to: semantic condition vectors used to guide feature fusion and attention allocation within the temporal deep prediction network; temporal transition probability correction matrices used to dynamically adjust the probability of transitions between action stages; and decision control parameters used to configure rule-based state machine decision conditions and stage detection switches.

[0021] In this embodiment of the invention, semantic extraction can be performed on sampled video frames to obtain macroscopic semantic features that describe the high-level tactical intent and overall motion attributes of the swing action, such as action type, racket hand, hitting direction, stance, and swing spin intention. The semantic extraction methods may include, but are not limited to: ① inputting the sampled video frames into a preset multimodal large language model, which outputs macroscopic semantic features based on its cross-modal understanding capabilities; ② inputting the sampled video frames into a cascaded classification structure composed of multiple dedicated convolutional neural network classifiers, extracting various sub-features from the macroscopic semantic features sequentially or in parallel; ③ extracting the coordinates of key points of the player's skeleton from the sampled video frames, and inferring macroscopic semantic features based on the skeletal kinematic parameters of consecutive frames using preset geometric and physical rules.

[0022] Next, based on the extracted macroscopic semantic features, corresponding constraint information is generated. This constraint information can be used to provide prior guidance and dynamic control basis in the subsequent fine segmentation stage. The methods for generating constraint information may include, but are not limited to: ① converting each discrete macroscopic semantic feature into a numerical vector according to a preset encoding rule, and concatenating the vector fragments in a fixed order into a one-dimensional tensor to form a semantic condition vector, which is then used as constraint information; ② constructing a state transition probability correction matrix corresponding to the number of action stages based on the macroscopic semantic features, and using this correction matrix as constraint information. The elements in the matrix represent correction values ​​for the initial transition probability between two action stages, and the sign and magnitude of each correction value depend on the specific value of the macroscopic semantic features; ③ querying a pre-configured semantic-parameter mapping knowledge base to obtain the judgment control parameters corresponding to the macroscopic semantic features, and using these judgment control parameters as constraint information. These judgment control parameters may include rotation angle judgment condition parameters, time window judgment condition parameters, speed judgment condition parameters, stage detection switch commands, etc.

[0023] S130. Extract multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head.

[0024] The multi-target spatial coordinate information can refer to the set of position coordinates of multiple independent physical entities extracted from the video frame in three-dimensional space. For example, it can include at least the coordinates of key points on the human skeleton, the coordinates of a sphere, and the coordinates of the racket head. The coordinates of key points on the human skeleton can refer to the coordinates of the major joints of the human body extracted from the video frame in three-dimensional space. These key points can include, but are not limited to, joints such as the shoulder, elbow, wrist, hip, knee, and ankle. The coordinates of a sphere can refer to the coordinates of a moving sphere (such as a tennis ball or badminton shuttlecock) extracted from the video frame in three-dimensional space, used to describe the instantaneous spatial position and flight trajectory of the sphere. In the following embodiments, a tennis ball will be used as an example for illustration. The coordinates of the racket head can refer to the coordinates of the key points at the top of the racket frame extracted from the video frame in three-dimensional space, used to accurately describe the motion state and spatial trajectory of the racket head.

[0025] In this embodiment of the invention, multi-target detection and tracking can be performed on each video frame in the target video. At least three independent entities, namely the player's body, the ball, and the racket, can be located in the image. Then, the coordinates of the key points corresponding to each entity are determined to obtain the corresponding multi-target spatial coordinate information. This information can provide an accurate kinematic basis for subsequent temporal segmentation of the action stage guided by constraint information.

[0026] S140. Based on constraint information and multi-target spatial coordinate information, segment the target video into swing motions and output the time information corresponding to multiple action stages of the swing motion.

[0027] In this context, swing motion segmentation refers to the process of dividing a continuous sequence of video frames into several discrete motion stages with clear physical meanings, based on the phased characteristics of the swing motion. Each motion stage can be a sub-process with specific kinematic characteristics and tactical functions throughout the entire swing motion, such as the preparation, backswing, racket head drop, ball impact, and follow-through. Temporal information refers to the start and end times or instantaneous moments corresponding to each motion stage on the target video timeline.

[0028] In this embodiment of the invention, the aforementioned generated constraint information can be used as a control or guidance signal to apply to the temporal segmentation process based on multi-target spatial coordinate information. During this process, the constraint information can adjust the segmentation logic's focus on spatial features or the judgment conditions for stage transitions in real time, thereby accurately dividing multiple action stages from a continuous frame sequence, such as initiation (S1), power accumulation (S2), drop shot (S3), ball contact (S4), and follow-through (S5), and finally outputting the start and end times or instantaneous moments of each action stage on the target video timeline. The core function of the constraint information is to transform macroscopic semantic understanding into adaptive control of the microscopic segmentation process, enabling the segmentation logic to dynamically adjust according to different action types, tactical intentions, and player characteristics, fundamentally solving the problem of poor generalization ability in traditional methods.

[0029] The technical solution of this invention extracts macroscopic semantic features from sampled video frames obtained by downsampling the target video and generates constraint information. At the same time, it extracts multi-target spatial coordinate information from each frame of the target video, including at least the coordinates of key points of the human skeleton, the coordinates of the ball, and the coordinates of the racket head. This enables the action segmentation process to have both a global understanding of the player's tactical intentions (such as action type, stance, and spin intention) and the ability to accurately capture the spatial motion details of the ball and racket. The constraint information comes from the macroscopic semantic features and is used to dynamically guide the action segmentation process based on the multi-target spatial coordinate information. This allows the segmentation logic to adaptively adjust according to different scenarios, overcoming the shortcomings of traditional fixed-rule methods that use the same standard to deal with all scenarios, resulting in insufficient generalization ability. This significantly improves the accuracy and robustness of action segmentation in complex and varied swing motion scenarios.

[0030] Example 2 Figure 2 This is a flowchart of a swing motion segmentation method provided in Embodiment 2 of the present invention. It is further optimized and extended based on the above embodiments and can be combined with various optional technical solutions in the above embodiments. For example... Figure 2As shown in the figure, the swing motion segmentation method provided in this embodiment is a further refinement of the swing motion segmentation process using a multimodal large language model and a temporal deep prediction network. The method specifically includes the following steps: S210. Downsample the target video to obtain sampled video frames.

[0031] In this embodiment of the invention, downsampling the target video to obtain sampled video frames may include at least one of the following: Video frames are extracted from the target video at preset time intervals and used as sampled video frames; When a preset trigger event is detected, a preset number of video frames are extracted from the target video before and after the time of the preset trigger event, and these are used as sampled video frames.

[0032] The preset time interval can refer to a pre-set time interval parameter used to perform uniform sampling operations, such as sampling one frame every several seconds or frames. Its value can be determined based on the original frame rate of the target video and the processing capability of the first subsystem. The preset trigger event can refer to a specific signal or state that appears in the target video or the output of the second subsystem, indicating that a swing action or a change in the action phase may occur. For example, it may include, but is not limited to, sudden changes in ball speed, human displacement acceleration exceeding a preset threshold, and detection of the sound of hitting the ball.

[0033] Specifically, based on the system's computing power configuration and business accuracy requirements, at least one of the following sampling strategies can be selected for execution: (1) Uniform sampling: video frames are periodically extracted from the target video at preset time intervals to form sampled video frames that are uniformly distributed in time. (2) Event-driven sampling: When a preset trigger event is detected, such as a sudden change in ball speed or human displacement acceleration exceeding a preset threshold, a preset number of video frames are extracted forward and backward based on the time of the event, and used as the sampled video frames.

[0034] This embodiment effectively reduces the computational overhead and inference latency of the first subsystem through downsampling, enabling stable extraction of macroscopic semantic features even with limited computing power. Simultaneously, the uniform sampling method ensures continuous tracking of slowly changing semantic features, while the event-driven sampling method ensures sufficient temporal context is captured when key swing events occur, thereby generating high-quality constraint information to guide the fine-grained action segmentation of the second subsystem.

[0035] S220. Input the sampled video frames into a preset multimodal large language model to obtain macroscopic semantic features.

[0036] Among them, the pre-trained multimodal large language model can refer to a large language model that has been pre-trained and can process image and text information simultaneously, and has the ability to understand across modalities, make global inferences and associate contexts; for example, the pre-trained multimodal large language model includes at least a vision language model (VLM).

[0037] In this embodiment of the invention, the extracted sampled video frames can be input into a pre-trained preset multimodal large language model. Utilizing the built-in image understanding and semantic reasoning capabilities of the large model, deep analysis and semantic mapping are performed on the visual elements in the sampled video frames. Specifically, the large language model can identify basic visual information such as the player's body posture, racket position, and court markings from a single frame image. Through cross-frame temporal analysis, it can understand the player's movement trends, swing trajectory evolution, and stance changes in consecutive frames. Based on the above visual understanding results, the model further combines the professional domain knowledge learned during its pre-training process to transform the underlying visual features into a high-level macro-tactical semantic description through semantic reasoning, thereby outputting structured macro-semantic features. These macro-semantic features may include at least one of the following: action type (forehand / single-handed backhand / double-handed backhand), racket hand (left hand / right hand), hitting direction (straight / diagonal), stance (open / closed), hitting spin intention (topspin / backspin / slice / flat), court position (baseline area / net volley area), and player's physical condition.

[0038] In one embodiment, a player's physical condition can refer to a quantitative indicator of physical fatigue or energy consumption, obtained by assessing the deviation of the player's current posture from a preset baseline posture. The acquisition process is as follows: The second subsystem, based on extracting the coordinates of key points of the human skeleton frame by frame, calculates spatial geometric feature parameters of specific force exertion phases from each swing cycle. These parameters include, for example, the knee flexion angle during the initiation phase, the elbow height on the striking side during the power-gathering phase, and the torso rotation angle during the contact phase. The calculated geometric parameters are then reported to the first subsystem. Upon receiving these geometric parameters, the first subsystem compares them with a pre-established baseline of the player's individual normal threshold range to determine the deviation of each parameter from the baseline. The degree of deviation from the baseline is measured, such as whether the knee flexion angle decreases by more than a preset proportional threshold, whether the height of the elbow on the hitting side remains below the baseline, or whether the trunk rotation angle is significantly reduced. When a parameter is detected to deviate from the baseline to a certain extent, the first subsystem inputs the various deviation values ​​as independent variables into a preset linear weighted evaluation model and outputs a physical fitness confidence score with a value between 0 and 1, or maps it to a discrete physical fitness level (such as good, slightly fatigued, severely deformed). This physical fitness status value is then incorporated into the constraint information generation process by the first subsystem as one of the macro-semantic features, and serves as a regulating factor to notify the second subsystem to dynamically adjust the spatial distance threshold for judging the power accumulation phase and the drop shot phase.

[0039] S230. Encode the macroscopic semantic features to generate a semantic condition vector and a temporal transition probability correction matrix, and use the semantic condition vector and the temporal transition probability correction matrix as constraint information.

[0040] The semantic condition vector can be a low-dimensional vector generated by encoding the extracted macroscopic semantic features. The temporal transition probability correction matrix can be a matrix data structure whose dimension corresponds to the number of stages of the swing action. It is used to correct the transition probabilities between each action stage in the baseline state transition probability matrix, realizing the dynamic state transition logic under different tactical actions.

[0041] In this embodiment of the invention, one-hot, multi-hot, and other encoding rules can be used to classify and encode the obtained macroscopic semantic features. For example, the action type can be encoded as a three-dimensional vector, the stance posture as a two-dimensional vector, and the intention to spin the ball as a three-dimensional vector. Then, the encoded vectors corresponding to each feature are concatenated in a preset order to form a one-dimensional semantic condition vector. Furthermore, based on the extracted macroscopic semantic features and combined with the dimension and structure of the baseline state transition probability matrix, a temporal transition probability correction matrix with the same dimension as the baseline matrix is ​​generated. The element in the i-th row and j-th column of this correction matrix represents the probability correction amount for transitioning from stage Si to stage Sj, used to dynamically adjust the state transition probability between each action stage. The value of this correction amount follows a preset semantic-correction mapping rule. For example, under default conditions (such as topspin or normal swing), the value of each element in the correction matrix is ​​zero, or a non-zero adjustment amount with an absolute value less than a preset threshold, so that the baseline state transition probability matrix can be slightly corrected or not corrected subsequently. When the macroscopic semantic features indicate a specific tactical intention, the correction amount at the corresponding position is set to a specific value. For example, when the intended spin of the shot is a backspin or slice shot, the correction value of the transition probability from the power-up phase to the drop shot phase in the correction matrix is ​​set to negative infinity or a very large negative value, so that the transition probability of this path after subsequent matrix superposition is zero or less than a preset threshold, thereby blocking the transition path from the power-up phase to the drop shot phase during the decoding process; when the position on the court is a net volley, the correction value of the transition probability from the start phase to the contact phase in the temporal transition probability correction matrix is ​​set to a large positive value, so that the probability of directly transitioning from the start phase to the contact phase after superposition is significantly increased, in order to adapt to the short swing characteristics of the net volley action without a complete power-up process and drop shot process.

[0042] In one specific embodiment, multi-class one-hot encoding or multi-label multi-hot encoding can be used to numerically represent each macroscopic semantic feature, and the encoding results are concatenated into a one-dimensional tensor as a semantic condition vector. Specifically, for the action type feature, if it is identified as "forehand", it is encoded as [1,0,0], "one-handed backhand" is encoded as [0,1,0], and "two-handed backhand" is encoded as [0,0,1]; for the ball spin intention feature, if it is identified as "topspin", it is encoded as [1,0,0], "flat hit" is encoded as [0,1,0], and "backspin / slice" is encoded as [0,0,1]; for the stance feature, if it is identified as "open", it is encoded as [1,0], and "closed" is encoded as [0,1]. The above feature encoding segments are concatenated end to end in a preset fixed order to form a one-dimensional semantic condition vector, for example, [1,0,0,0,0,1,1,0,…]. The semantic condition vector is then sent to the second subsystem as part of the constraint information. It is fused with the multi-target spatial coordinate features extracted by the second subsystem and used as additional input features for the temporal deep prediction network, providing the network with prior guidance on macroscopic tactical semantics.

[0043] The temporal transition probability correction matrix can be a 5×5 numerical matrix constructed based on the state transition probabilities of a Markov chain, corresponding to five action stages: S1 (start), S2 (charge), S3 (slam), S4 (ball contact), and S5 (follow-through). The system presets a baseline state transition probability matrix, where each element represents the initial transition probability between two stages. In a normal topspin swing scenario, the value of P(S1→S3) in the baseline matrix is ​​an extremely high probability, such as 0.9, indicating a very high probability of entering the slam phase after the charge phase; while the value of P(S1→S4) is an extremely low probability, such as 0.05, indicating an extremely low probability of directly jumping to the ball contact phase after the charge phase. When the macroscopic semantic feature of "backspin / slice shot" is identified from the sampled video frames, the system generates the corresponding temporal transition probability correction matrix. In this correction matrix, the elements corresponding to the S2→S3 transition path are forcibly encoded as negative infinity or a very large negative value (to make the superimposed transition probability approach zero), while the elements corresponding to the S2→S4 transition path are set to higher correction values ​​(e.g., increasing the base value from 0.05 to 0.9). During subsequent decoding, this correction matrix is ​​superimposed on the base state transition probability matrix, making the value of P(S2→S3) in the corrected transition probability matrix zero or close to zero. This cuts off the path of the state machine searching for the drop phase from the algorithm's underlying layer, forcing the decoding path to flow directly from the charging phase to the ball contact phase, preventing decoding deadlock or incorrect segmentation caused by forcibly searching for a non-existent drop phase during the cutting action.

[0044] S240. Extract multi-target spatial coordinate information from each video frame of the target video.

[0045] In this embodiment of the invention, each video frame of the target video can be sequentially input into a preset target detection model to obtain the coordinates of human skeleton key points, sphere coordinates, and racket head coordinates in each video frame. The preset target detection model may include, but is not limited to: deep learning-based pose estimation models (e.g., MediaPipe, MMPose, OpenPose, etc.), target detection models based on deep learning or traditional visual features (e.g., YOLO series, color / shape feature-based detection models), and any combination of the above models.

[0046] S250. Align and stitch together the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head along the time axis to obtain multi-target spatial features.

[0047] Among them, multi-target spatial features can refer to composite feature vectors formed by aligning and stitching the three-dimensional spatial coordinates of three types of physical entities: human skeleton key points, sphere, and racket head. This feature vector integrates three complementary spatial information types—human posture, ball position, and racket position—within the same frame, and can comprehensively represent the motion state at the moment of the swing.

[0048] In this embodiment of the invention, after obtaining the coordinates of human skeleton key points, sphere coordinates, and racket head coordinates in each video frame of the target video, the three types of entity coordinates extracted in the same frame are matched and aligned according to the timestamp, based on the timestamp. If a certain type of coordinate is missing in a certain frame (e.g., the sphere is occluded, causing detection failure), it can be supplemented by linear interpolation or motion prediction of the preceding and following frames to ensure that each frame has complete three types of coordinate data. Then, the three types of entity coordinates of each aligned frame are concatenated into a one-dimensional feature vector according to a preset splicing order (e.g., first arranging the coordinates of each joint of the human body, then arranging the sphere coordinates, and finally arranging the racket head coordinates, etc.). This one-dimensional feature vector is the multi-target spatial feature of the frame.

[0049] S260. The semantic condition vector is fused with the multi-objective spatial features to obtain the fused feature tensor.

[0050] Among them, the fusion feature tensor can refer to the multi-dimensional array structure obtained by fusing semantic condition vectors and multi-object spatial features. It contains both spatial dimension motion geometry information and channel dimension semantic regulation information, while maintaining the inter-frame sequence relationship in the temporal dimension, and can be used as input to a temporal depth prediction network.

[0051] In this embodiment of the invention, the macroscopic tactical semantic information carried by the semantic condition vector can be deeply fused with the low-level spatial motion information carried by the multi-target spatial features. The fusion method can include, but is not limited to, feature splicing in the channel dimension, element-wise weighted control based on attention mechanism, etc., aiming to enable the two types of heterogeneous information to achieve complementary enhancement in a unified feature space, and finally obtain the fused feature tensor after deep fusion.

[0052] S270. Input the fused feature tensor into the preset temporal depth prediction network to obtain the original probability distribution of each video frame in the target video belonging to each action stage.

[0053] The pre-built temporal deep prediction network can refer to a pre-constructed and trained deep neural network model with temporal modeling capabilities, used to analyze the input video frame sequence frame by frame, extract the spatiotemporal motion patterns, and output the probability of action stage attribution for each frame. For example, the pre-built temporal deep prediction network can include, but is not limited to, Transformer, ST-GCN models, etc. It should be noted that, with the continuous evolution of future model technologies, the pre-built temporal deep prediction network can also adopt deep learning model structures based on spatiotemporal convolution, recurrent neural networks, or any other deep learning model structure with temporal sequence modeling capabilities; these are all within the scope of protection of this application.

[0054] The original probability distribution can refer to the probability vector output by the preset temporal depth prediction network for a single video frame. Each element in the vector corresponds to the confidence value of a frame belonging to a preset action stage, and the sum of all elements is 1.

[0055] In this embodiment of the invention, the aforementioned constructed fusion feature tensor can be used as input and fed into a pre-constructed and trained preset temporal depth prediction network. This network has the ability to model temporal dependencies, capture the dynamic pattern of action evolution between consecutive video frames, extract the spatiotemporal pattern related to the division of the swing action stage from each frame, and output the original probability distribution (probability vector) corresponding to each frame. The original probability distribution is equal to the number of action stages, and the value of each element in the probability distribution represents the confidence that the frame belongs to the corresponding action stage.

[0056] S280. Based on the original probability distribution and the time-series transition probability correction matrix, determine the time information corresponding to multiple action stages of the swing action.

[0057] In this embodiment of the invention, a pre-constructed baseline state transition rule conforming to the physical logic and unidirectional flow characteristics of the ball swing action can be obtained. This rule clearly defines the legal and prohibited flow paths between each predetermined stage of the swing action, as well as the basic flow probability corresponding to each legal path, providing a basic constraint framework for temporal segmentation that conforms to common sense about motion. Next, a temporal transition probability correction matrix generated based on macroscopic semantic features is applied to the baseline state transition rule to dynamically adjust the stage flow probabilities in the baseline rule. For example, the flow probability of legal flow paths conforming to the semantics of the current hitting scenario is strengthened, while the flow probability of prohibited flow paths that do not conform to the semantics of the current scenario (such as the "power-up stage → drop stage" path in the case of a spin slice shot) is zeroed or significantly weakened, ultimately resulting in a corrected temporal transition rule adapted to the current hitting scenario. Furthermore, using the original probability distribution as observational evidence and the corrected temporal transition rule as the constraint condition for the jumps between each action stage, a preset sequence decoding algorithm is invoked to perform a global optimal path search, assigning a unique action stage label to each video frame in the target video. Finally, the frame numbers where these labels changed are located and converted in conjunction with the frame rate of the target video to accurately determine the time information corresponding to each action phase, including start-up, charging, drop shot, ball contact, and follow-through.

[0058] The technical solution of this invention extracts macroscopic semantic features from downsampled video frames and encodes them to generate semantic condition vectors and temporal transition probability correction matrices. Simultaneously, it extracts three types of coordinates—human skeleton key points, sphere, and racket head—from each frame of the target video and concatenates them into multi-object spatial features. The semantic condition vectors and multi-object spatial features are then fused and input into a temporal depth prediction network. This allows the network to perform frame-by-frame action stage discrimination based on spatial motion features, while also being dynamically guided by macroscopic semantic features to adjust attention weights. This enables the network to adaptively adjust the level of attention to specific action patterns in complex and ever-changing tactical scenarios. Based on the collaborative decoding of the original probability distribution and the temporal transition probability correction matrix, it can dynamically correct the stage transition path according to the player's hitting intention (such as topspin or slice), avoiding stage misjudgments or deadlocks caused by fixed transition rules. Ultimately, this achieves high-precision, highly generalizable temporal segmentation of the swing action stage.

[0059] Furthermore, based on the above embodiments of the invention, the preset temporal depth prediction network integrates an attention mechanism to fuse semantic conditional vectors with multi-object spatial features to obtain a fused feature tensor, including: Based on semantic conditional vectors, attention weights corresponding to each spatial dimension in multi-objective spatial features are generated through an attention mechanism. The multi-object spatial features are weighted element-wise based on attention weights to obtain the fused feature tensor.

[0060] The attention mechanism can refer to an algorithm module integrated into a pre-defined temporal deep prediction network. Based on input control signals (such as semantic condition vectors), it dynamically generates differentiated weight coefficients for each spatial dimension of the multi-target spatial features. This guides the network to allocate different levels of attention to different parts of the input features, strengthening features relevant to the current swing segmentation task and suppressing irrelevant interference features to adapt to different swing tactical scenarios. Spatial dimension refers to the dimension corresponding to the time step, entity object, or spatial coordinates in the multi-target spatial features, used to represent the spatial motion information of different targets at different times. Attention weights are a set of numerical coefficients generated by the attention mechanism based on the semantic condition vector. Each coefficient corresponds to a spatial dimension in the multi-target spatial features. The magnitude of this weight determines the importance of the corresponding feature component in subsequent network calculations; a high weight indicates that the feature component is given priority, while a low weight indicates that the feature component is suppressed.

[0061] In this embodiment of the invention, a pre-encoded semantic condition vector can be input into the attention mechanism module of a preset temporal depth prediction network as global guiding information for attention weight generation. This semantic condition vector carries the global macro-semantics of the swing action extracted from low-frame-rate sampled video frames, covering core tactical features such as ball spin intention, racket hand, stance, and action type. Then, the attention mechanism module can analyze the correlation between each spatial dimension of the multi-target spatial features and the current tactical intention based on the macro-semantic information carried by the semantic condition vector, generating a corresponding attention weight value for each spatial dimension. For example, when the semantic condition vector indicates that the current action is a net volley, the attention mechanism module automatically assigns a lower weight value to the feature dimension representing a large backward swing and a higher weight value to the feature dimension representing a short forward push. Then, based on the generated attention weights, the multi-target spatial features are weighted element-wise. In this way, the original feature values ​​of spatial dimensions corresponding to high weights are preserved or enhanced, while the original feature values ​​of spatial dimensions corresponding to low weights are suppressed or attenuated. Finally, the multi-target spatial features after dynamic reweighting are used as a fusion feature tensor that integrates global macroscopic semantic constraints and microscopic spatial motion details. This tensor is then output to the subsequent network layer of the preset temporal depth prediction network for probability prediction of each stage of the subsequent swing action.

[0062] In one embodiment of the present invention, the preset temporal depth prediction network is a Transformer network, and the attention mechanism is a cross-attention mechanism. Specifically, the semantic condition vector can be linearly mapped to serve as the query vector, and the multi-target spatial features can be linearly mapped to serve as the key vector and value vector; the attention weight distribution is obtained by calculating the similarity between the query vector and the key vector. When the macroscopic semantic features contained in the semantic condition vector indicate a front-end interception, the attention weight distribution automatically reduces the weight corresponding to the video frames exhibiting large backswing motion features in the time window, thereby weakening the response during the build-up phase.

[0063] In another embodiment of the invention, the preset temporal depth prediction network is a spatiotemporal graph convolutional network that incorporates a channel attention mechanism. Specifically, the semantic condition vector can be mapped to a channel weight vector through a fully connected layer. The dimension of this channel weight vector is the same as the number of channels of the node features in the spatiotemporal graph convolutional network. The channel weight vector is multiplied channel-by-channel with the node feature matrix to adjust the activation values ​​of each feature channel through the semantic condition vector. When the macroscopic semantic feature indicates a front-line intercept, the activation values ​​of the feature channels corresponding to large displacements in the shoulder or wrist are suppressed.

[0064] It should be understood that the attention mechanism in this embodiment is not limited to the two implementation methods described above. Any mechanism that can apply differentiated weights to different spatial dimensions of multi-objective spatial features using semantic conditional vectors, such as self-attention mechanisms, spatial attention mechanisms, and hybrid attention mechanisms, can be used to implement this invention.

[0065] This embodiment integrates an attention mechanism into a temporal deep prediction network and utilizes semantic conditional vectors to regulate the generation of attention weights, achieving dynamic guidance of macro-tactical understanding for micro-motion feature extraction. This fusion approach enables the network to automatically focus on the matching motion pattern based on the tactical type of the current action when processing multi-target spatial features, while suppressing responses to irrelevant or interfering motion patterns. Compared to static fusion methods that simply concatenate semantic features with spatial coordinates, this scheme achieves refined regulation of spatial motion information by semantic constraints at the feature level through element-wise weighting of attention weights. This significantly improves the network's generalization ability and motion phase segmentation accuracy for swing actions of different player styles and tactical intentions.

[0066] Furthermore, based on the above embodiments of the invention, the time information corresponding to multiple action stages of the swing action is determined based on the original probability distribution and the time transition probability correction matrix, including: Obtain the preset baseline state transition probability matrix; the elements in the baseline state transition probability matrix represent the initial transition probability between two stages in multiple action stages; The corrected state transition probability matrix is ​​obtained by superimposing the time-series transition probability correction matrix with the baseline state transition probability matrix. The corrected state transition probability matrix is ​​used as the transition probability constraint. The preset sequence decoding algorithm is called according to the original probability distribution of each video frame in the target video to determine the action stage label corresponding to each video frame. Based on the changes in the action stage labels corresponding to each video frame and the frame rate of the target video, the time information corresponding to each action stage is determined.

[0067] The baseline state transition probability matrix can be a pre-defined and stored parameter matrix representing the general transition probabilities between different phases of a swing motion without external specific semantic constraints. Each row of this matrix represents the initial probability distribution of transitioning from the current phase to other phases. The transition probability constraint can be a probability parameter used in a pre-defined sequence decoding algorithm to constrain the rationality of phase transitions between adjacent frames. Numerically, it is taken from the modified state transition probability matrix and represents the permissible degree of transition from a certain phase in the current frame to a certain phase in the next frame. The pre-defined sequence decoding algorithm can be a computational method capable of finding a globally optimal state label path under time-series data and state transition constraints, such as, but not limited to, the Viterbi algorithm and dynamic programming-based sequence labeling algorithms. The action phase label can be a classification identifier assigned to each video frame in the target video, explicitly indicating which specific phase the frame belongs to: the start-up phase, the charging phase, the drop phase, the ball-touching phase, or the follow-through phase.

[0068] In this embodiment of the invention, the process of determining the time information corresponding to each phase of the swing motion specifically includes: Step 1: Load the baseline state transition probability matrix A that matches the swing motion logic. This matrix defines the default transition rules between each stage of the swing motion. Each element in the matrix... This represents the initial probability of transitioning to the j-th action stage in the next frame, given that the current frame is in the i-th action stage. These initial probabilities are set based on the kinematic laws of general swing actions. For example, the probability of transitioning from the power-up stage (S2) to the drop stage (S3) is relatively high, while the probability of directly jumping from the power-up stage to the follow-through stage (S5) is extremely low.

[0069] Step 2: Add the temporal transition probability correction matrix C generated based on macroscopic semantic features to the aforementioned baseline state transition probability matrix A element-wise to obtain the corrected state transition probability matrix. The temporal transition probability correction matrix C and the baseline state transition probability matrix A have the same size (both N×N, where N is the number of action stages), and their element values... This represents the adjustment amount for the transition probability at the corresponding stage. For transition paths that need to be suppressed (e.g., S2→S3 in the cutting ball scenario), the element at the corresponding position in the correction matrix C takes a negative value (e.g., negative infinity or a very large negative value); for transition paths that need to be enhanced, the value takes a positive value; for paths that do not require adjustment, the value takes zero. The summation yields the corrected state transition probability matrix, whose elements are... .

[0070] Step 3: Using the original probability distribution output by the preset temporal depth prediction network as the observation probability (i.e., the confidence level of each frame independently belonging to each stage), and the corrected state transition probability matrix as the transition probability constraint, the preset sequence decoding algorithms such as the Viterbi algorithm and dynamic programming algorithm are called to perform global optimal decoding on each video frame of the target video. Taking the Viterbi algorithm as an example, when calculating the total score of each candidate stage path, it considers: ① the observation probability that the current frame is predicted to be a certain stage; ② the corrected transition probability of transitioning from a certain stage in the previous frame to a certain stage in the current frame. The algorithm searches for the path with the highest total score among all possible stage paths and finally outputs the action stage label corresponding to each video frame in the target video.

[0071] Step 4: Traverse the action phase labels obtained in Step 3, and detect the boundary frame numbers where phase labels change between adjacent video frames; whenever a label changes (e.g., from S2 to S3), mark this as a phase transition point. Then, using the known frame rate of the target video (in frames per second), divide the frame number of each boundary frame by the frame rate to convert it into a timestamp in seconds. Specifically, for instantaneous event phases (such as the charging phase and the ball-touching phase), a corresponding single timestamp will be output; for continuous process phases (such as the start phase, the drop phase, and the follow-through phase), start and end time information consisting of the start and end times will be output.

[0072] This embodiment achieves dynamic control of the stage transition path during decoding by encoding macroscopic semantics such as action type and hitting intention into a temporal transition probability correction matrix and superimposing it with the baseline transition probability matrix. This mechanism enables the same segmentation model to adaptively adjust the allowance of stage transitions when facing different player styles and tactical intentions (such as the different requirements of topspin and slice shots for the drop phase), fundamentally solving the problems of poor generalization ability and easy temporal misjudgment in traditional fixed-rule or static decoding schemes in multiple scenarios. At the same time, the use of a sequence decoding algorithm effectively filters out local noise and jitter in frame-by-frame prediction, outputting smooth and physically reasonable stage segmentation results, significantly improving the accuracy and robustness of fine action stage extraction in complex motion scenarios.

[0073] Furthermore, based on the above embodiments of the invention, the multiple action stages include a ball-touching stage, and the above method further includes: Extract the frame number range corresponding to the ball contact phase; Determine whether there are target video frames within the frame sequence interval where the racket and the ball spatially overlap; If it exists, the time corresponding to the target video frame will be used as the time information of the ball-touching phase; If not, the three-dimensional motion trajectory of the ball corresponding to the ball coordinates and the three-dimensional motion trajectory of the racket head corresponding to the racket head coordinates within the frame sequence interval are fitted. The intersection point of the three-dimensional motion trajectory of the ball and the three-dimensional motion trajectory of the racket head in three-dimensional space is determined. Based on the time parameter corresponding to the intersection point, the corrected ball contact time is obtained by time interpolation, and the time information of the ball contact stage is updated with the corrected ball contact time.

[0074] The frame sequence interval refers to a continuous range of frame numbers extracted based on the action stage labels corresponding to each video frame in the target video. All video frames within this range are marked as the same action stage, representing the time window of that stage. Spatial overlap between the racket and the ball refers to a state in three-dimensional space where the distance between the racket head coordinates and the ball coordinates is less than a preset contact threshold, indicating that the racket and ball have made physical contact or are about to make contact. The three-dimensional trajectory of the ball refers to a spatial curve equation describing the change of the ball's position in three-dimensional space over time, obtained through mathematical fitting methods (such as polynomial fitting) based on the three-dimensional coordinate sequence of the ball from multiple video frames within the frame sequence interval. The three-dimensional trajectory of the racket head refers to a spatial curve equation describing the change of the racket head's position in three-dimensional space over time, obtained through mathematical fitting methods based on the three-dimensional coordinate sequence of the racket head from multiple video frames within the frame sequence interval. The three-dimensional spatial intersection point refers to the common intersection point of the three-dimensional trajectory of the ball and the three-dimensional trajectory of the racket head in three-dimensional space; the time corresponding to this intersection point is the contact time. Temporal interpolation refers to mapping the time parameters corresponding to the intersection points in three-dimensional space onto the time axis of discrete video frames, and obtaining an accurate ball-touching timestamp through interpolation calculation. It is used to solve the problem of missed ball-touching moments caused by dropped frames or excessively fast movements.

[0075] Specifically, this solution may also include the following ball touch time acquisition process: Step 1: Based on obtaining the action stage labels corresponding to each video frame in the target video, extract the frame number range of all consecutive video frames marked as the ball-touching stage. This frame number range constitutes the candidate range of the ball-touching stage on the time axis.

[0076] Step 2: Traverse each video frame within the frame number interval, obtain the corresponding ball coordinates and racket head coordinates, and calculate the Euclidean distance between them. When the distance is less than the preset contact threshold, determine that the video frame is the target video frame in which the racket and the ball overlap in space, and record the corresponding frame number.

[0077] Step 3: If at least one target video frame is detected within the frame sequence number range, the time corresponding to the overlapping frame with the highest confidence or the first occurrence is taken as the final time information of the ball-touching stage. This time can be calculated by dividing the frame sequence number of the target video frame by the frame rate of the target video.

[0078] Step 4: If no spatially overlapping target video frames are detected within the frame sequence number range, then perform the following physical interpolation error correction procedure: Step 4.1: Obtain the sphere coordinate sequence and racket head coordinate sequence corresponding to each video frame within the frame number interval. Fit the sphere's trajectory in three-dimensional space based on the sphere coordinate sequence to describe the change in the sphere's position over time; and fit the racket head's trajectory in three-dimensional space based on the racket head coordinate sequence to describe the change in the racket head's position over time. The fitting method can be polynomial fitting or parameter estimation based on a physical model (such as uniform linear motion, parabolic motion, etc.). This embodiment does not impose any restrictions on this.

[0079] Step 4.2: Based on the fitted three-dimensional trajectory of the ball and the three-dimensional trajectory of the racket head, determine the intersection point of the two spatial trajectories in three-dimensional space. If the two trajectories do not strictly intersect at a single point, calculate the moment that minimizes the Euclidean distance between the ball's position and the racket head's position, and use this as the time parameter corresponding to the approximate intersection point.

[0080] Step 4.3: The intersection of the two spatial trajectories in three-dimensional space corresponds to a continuous time parameter, representing the theoretically precise moment when the racket and ball make physical contact. Since the target video is composed of discrete video frames, this moment may not fall exactly on the sampling time of a single frame. Therefore, this time parameter is used as input, time mapping is performed based on the frame rate of the target video, and the precise timestamp corresponding to this time parameter on the video time axis is calculated through linear interpolation. This timestamp is the corrected ball contact time.

[0081] Step 4.4: Replace the ball contact time information initially obtained based on the original label sequence with the corrected ball contact time as the final output of the ball contact stage.

[0082] This embodiment effectively solves the problem of failing to capture the moment of ball contact in high-speed swing scenarios due to video frame rate limitations or poor shooting conditions by combining visual frame detection with physical trajectory deduction. Specifically, when the video frame clearly records the ball contact, the visual detection result is used first to ensure maximum efficiency and direct accuracy. When visual detection fails, the three-dimensional spatial intersection point is calculated using the movement trajectories of the ball and racket head in the short time before and after contact. The moment of contact is deduced through physical laws, changing the acquisition of the ball contact time from "relying on luck to capture" to "relying on physical laws." The output accuracy can reach the microsecond level, significantly improving the robustness and accuracy of the time information during the ball contact phase, thereby providing a more reliable data foundation for subsequent sports technique analysis, action correction, or tactical statistics.

[0083] Furthermore, based on the above embodiments of the invention, constraint information is applied to the swing motion segmentation process through an asynchronous update mechanism, which includes: A globally shared state board is maintained in memory; the globally shared state board is used to store constraint information. When processing each video frame of the target video, read the constraint information stored in the global shared state panel, and use the constraint information to guide the swing action segmentation of the current video frame; After generating new constraint information based on the sampled video frames, the new constraint information is written to the global shared state board. For the same swing cycle, the constraint information in the globally shared state board remains locked after being written until it is restored to the preset baseline state after the current swing cycle ends.

[0084] In this embodiment, the global shared state board can refer to a data structure area allocated in system memory to store currently effective constraint information, allowing the first subsystem to write and the second subsystem to read, thereby realizing data transfer and state synchronization between two asynchronous processing systems. The asynchronous update mechanism can refer to a data interaction method where the first and second subsystems exchange data through a shared storage area, and write and read operations do not require synchronous waiting; wherein, the second subsystem continuously runs with currently available data, and the first subsystem updates its data independently after completing calculations, without blocking each other. The swing action cycle can refer to a complete action process starting from the initiation phase of a swing action, through the power-building phase, the drop phase, the ball-contact phase, and ending at the end of the follow-through phase. It is understood that there is usually a brief preparation or movement interval between adjacent swing actions. The preset baseline state can refer to the default state of the global shared state board, corresponding to the general action segmentation setting when no specific tactical intention is recognized, used to ensure the normal operation of the basic segmentation function when there is no specific semantic guidance.

[0085] In this embodiment of the invention, constraint information can be transmitted to the action segmentation process through an asynchronous update mechanism, so as to guide and regulate the subsequent action stage segmentation using the constraint information. The asynchronous update mechanism specifically includes: Step 1: When the system starts up, allocate a shared storage area in memory as a global shared state board. This state board can be used to store the latest available constraint information and write preset baseline constraint parameters in the initial state. These baseline constraint parameters correspond to general action segmentation settings that do not distinguish between specific tactical intentions.

[0086] Step 2: During the processing of the target video, when the second subsystem needs to segment a video frame for a swing motion, it first accesses the globally shared state panel, reads the currently stored constraint information, and processes the current video frame based on this constraint information. It can be understood that regardless of which frame the current video frame corresponds to, the second subsystem always uses the latest available constraint information in the state panel as its guide, without blocking and waiting for the processing results from the first system.

[0087] Step 3: After the first subsystem extracts the macroscopic semantic features from the target video and encodes them to generate new constraint information, it can write them to the global shared state board to overwrite the previously stored original constraint information.

[0088] Step 4: For the same continuous swing action cycle, once the first subsystem generates and writes the corresponding constraint information based on the macroscopic semantic features of the current action, this constraint information remains locked in the global shared state panel and is not overwritten by other constraint information that may be generated subsequently, until the end of the current swing action cycle. The end of the swing action cycle can be detected by the second subsystem as the end of the follow-through phase or triggered by a preset timeout condition.

[0089] Step 5: When the current swing action cycle is detected to have ended, the system resets the constraint information stored in the global shared status board to the preset baseline state in order to prepare for the next swing action cycle and avoid the specific tactical constraints of the previous cycle from causing undue interference to subsequent different actions.

[0090] This embodiment uses an asynchronous update mechanism to ensure that the semantic reasoning delay of the first subsystem does not block the frame-by-frame processing of the second subsystem, thus guaranteeing the real-time analysis capability of high frame rate videos. At the same time, the periodic locking strategy ensures that the segmentation logic of the same swing action is consistent before and after, avoiding the change of action stage labels or logical conflicts caused by mid-process updates of semantic understanding. Finally, under the condition of limited computing resources, it achieves swing action temporal segmentation with both high accuracy and high robustness.

[0091] Example 3 Figure 3This is a flowchart of a swing motion segmentation method provided in Embodiment 3 of the present invention. It is further optimized and extended based on the above embodiments and can be combined with various optional technical solutions in the above embodiments. For example... Figure 3 As shown in the figure, the swing motion segmentation method provided in this embodiment is a further refinement of the swing motion segmentation process using a multimodal large language model and a rule-based state machine. The method specifically includes the following steps: S310. Downsample the target video to obtain sampled video frames.

[0092] S320. Input the sampled video frames into a preset multimodal large language model to obtain macroscopic semantic features.

[0093] S330. Query the preset semantic-parameter mapping knowledge base, obtain the judgment control parameters corresponding to the macro semantic features, and use the judgment control parameters as constraint information.

[0094] The semantic-parameter mapping knowledge base can refer to a pre-built knowledge base used to establish the mapping relationship between macroscopic semantic features and decision control parameters. The construction of this knowledge base may be based on tennis professional teaching data, sports biomechanical analysis results, and expert experience rules. Decision control parameters can refer to values ​​or instructions directly used to configure or adjust state machine decision conditions and stage detection switches. Specifically, they may include at least one of rotation angle decision condition parameters, time window decision condition parameters, speed decision condition parameters, and stage detection switch instructions. Rotation angle decision condition parameters can refer to parameters used to set the trunk rotation angle or joint angle thresholds that the state machine needs to meet when determining a specific action stage. For example, it may include the minimum trunk rotation angle value required during the power-accumulation stage. Time window decision condition parameters can refer to parameters used to set the maximum or minimum time interval allowed when the state machine jumps through the decision stage. For example, it may include the maximum allowed time window value from the start-up stage to the ball-touching stage. Speed ​​decision condition parameters can refer to parameters used to set the speed or angular velocity thresholds that the state machine needs to meet when entering or exiting the decision stage. For example, it may include the lower limit of the wrist-racket head angle angular velocity required for the start-up stage decision. A stage detection switch instruction can refer to an instruction-type parameter used to control the opening or closing of the detection function for a specific action stage in a state machine. For example, it can include closing the detection function for the drop shot stage when the ball is identified as a slice shot.

[0095] In this embodiment of the invention, the system can pre-build and maintain a semantic-parameter mapping knowledge base, which stores the correspondence between macroscopic semantic feature combinations and decision control parameters in a structured form. Thus, after extracting the macroscopic semantic features of the current swing, these features can be used as query conditions to search and match within the knowledge base, thereby obtaining the corresponding decision control parameters, which are then used as constraint information. Specifically, the query operation can be implemented by combining various semantic tags from the macroscopic semantic features (e.g., action type "forehand," stance "closed," and spin intention "topspin") to form a query key, and then searching for the corresponding decision control parameters in the mapping table of the knowledge base. These decision control parameters may include at least one of the following: rotation angle decision condition parameters, time window decision condition parameters, speed decision condition parameters, and stage detection switch instructions.

[0096] S340. Configure the judgment conditions and stage detection switches of the preset state machine according to the judgment control parameters.

[0097] The preset state machine can refer to a finite state machine model with a finite number of states (corresponding to each action stage), sequential transitions between states, and configurable entry and exit conditions for each state. The judgment condition can refer to the quantized rules used by the state machine during frame-by-frame judgment to determine whether the transition from the current action stage to the next action stage is satisfied. These judgment conditions may include, but are not limited to: confirming entry into the start-up stage when the included angular velocity exceeds a velocity judgment threshold for multiple consecutive frames, or confirming the reaching of the power accumulation extreme point when the racket head velocity vector direction reverses and the velocity approaches zero. The stage detection switch can refer to a Boolean instruction or enable signal used to control the enabling or disabling of a specific action stage detection module in the state machine; for example, when the first subsystem recognizes the intention to slice the ball, it can issue a "disable racket drop stage detection" switch instruction, causing the state machine to skip the S3 judgment logic and directly transition from the power accumulation stage to the ball contact stage.

[0098] In this embodiment of the invention, the queried judgment control parameters can be injected into a pre-built rule-based state machine, enabling its judgment logic to dynamically adapt to different players and different tactical scenarios. The configuration method of the state machine may include, but is not limited to: ① writing the acquired numerical parameters into the comparator threshold register of the corresponding judgment node in the state machine. For example: writing the minimum wrist-racket head angle angular velocity threshold required for the start-up phase into the start-up phase judgment module, replacing the default angular velocity threshold; writing the maximum allowable time interval from the end of the power-charging phase to the start of the ball-touching phase into the ball-touching phase estimation module, limiting the allowable time range for entering the ball-touching phase after power-charging / racket drop; when the macroscopic semantic feature indicates a "closed stance," lowering the torso rotation angle judgment threshold by a preset amount (e.g., reducing it by 15 degrees) and writing it into the power-charging phase judgment module, etc. ② Configure the state transition control register inside the state machine according to the phase detection switch instruction. For example, when the instruction to "turn off the drop phase detection" is received, the enable signal of the S3 detection module is set to low level, and the module is bypassed in the subsequent frame-by-frame judgment; when "net interception" is detected, the timeout threshold of the follow-up phase detection module is shortened to adapt to the fast retraction characteristics of the interception action, etc.

[0099] After configuring all the judgment conditions and switches, the state machine can be initialized and verified to ensure that the threshold parameters and enable states of all judgment modules have been updated. The state machine is then reset to standby mode to wait for the input of the target video frame sequence to begin frame-by-frame judgment.

[0100] S350: Call the preset state machine to determine the action stage of the target video frame by frame according to the multi-target spatial coordinate information, so as to determine the action stage label corresponding to each video frame in the target video.

[0101] In this embodiment of the invention, the multi-target spatial coordinate information of each frame can be sequentially input into a preset state machine, and a series of rule calculations dynamically controlled by parameters can be executed. For example, it can determine the starting stage (S1) by the wrist-racket head angle angular velocity, capture the power accumulation extreme point by the racket head speed reversal (S2), identify the drop stage by comparing the racket head and wrist height and acceleration (S3), estimate the ball contact time by relying on the timing window (S4), and finally confirm the follow-through stage by comprehensively considering multiple conditions such as ball speed and racket speed (S5). In this way, each frame in the video is given a precise action stage label.

[0102] Furthermore, based on the above embodiments of the invention, a preset state machine is invoked to perform frame-by-frame action stage determination on the target video according to the multi-target spatial coordinate information, including: The angular velocity of the angle between the player's wrist and the racket head is determined based on the coordinates of key points of the human skeleton and the coordinates of the racket head in at least two video frames in the target video. When the angular velocity of the angle exceeds the velocity judgment condition in multiple consecutive frames, the start-up phase is confirmed. The racket head motion speed and velocity vector direction are determined based on the racket head coordinates of at least two video frames in the target video. When the racket head motion speed is detected to be less than a preset speed threshold and the velocity vector direction is reversed, the end time of the power accumulation phase is confirmed. After the power-up phase ends, if the phase detection switch command indicates that the racket drop phase detection is turned off, the determination of the racket drop phase is skipped and the inference of the ball contact phase is directly executed. Otherwise, the racket head height, racket wrist height and racket head acceleration are determined based on the coordinates of the human skeleton key points and the racket head coordinates of at least two video frames in the target video. When the racket head height is detected to be lower than the racket wrist height and the racket head acceleration is a continuous downward acceleration, the racket drop phase is confirmed. After the power-accumulation phase or the drop phase, the ball contact phase is presumed based on the time window determination condition parameters; if a video frame of the racket and the ball overlapping in space is detected before the presumed ball contact phase, the time corresponding to the video frame is used as the time information of the ball contact phase. The ball velocity vector and racket head three-dimensional acceleration are determined based on the ball coordinates and racket head coordinates of at least two video frames in the target video. When it is detected that the ball velocity vector is away from the racket head, the racket head three-dimensional acceleration is negative, and the racket head three-dimensional acceleration exceeds the preset rotation judgment condition, the follow-through phase is confirmed.

[0103] In this embodiment of the invention, the process of determining the frame-by-frame action stage of the target video using a rule-based state machine specifically includes: Step 1: The preset state machine starts from the first frame of the target video or the frame after the end of the previous swing cycle, and reads the coordinates of key points of the human skeleton and the racket head from the multi-target spatial coordinate information frame by frame. Then, based on the wrist joint coordinates and racket head coordinates on the racket-holding hand side, it calculates the direction vector of the wrist-racket head line in three-dimensional space, and obtains the angular velocity of the player's wrist and racket head by dividing the angle change of this direction vector in adjacent frames by the inter-frame time interval. Then, it compares the angular velocity of multiple consecutive frames with the velocity determination condition. When it detects that the angular velocity of multiple consecutive frames (e.g., three consecutive frames) all exceed the threshold specified by the velocity determination condition, the preset state machine confirms that the player has entered the swing initiation phase from a static or ready posture, and marks the corresponding frame as the start of the initiation phase. If no power-gathering phase characteristics are detected within the specified time window, the state rollback is performed, and the initiation phase mark is canceled.

[0104] Step 2: After confirming entry into the startup phase, the preset state machine continues to read the racket head coordinates frame by frame and calculates the instantaneous velocity (scalar) and velocity vector direction of the racket head based on the changes in the three-dimensional coordinates of the racket head in adjacent frames. When it is detected that the racket head velocity drops below the preset velocity threshold (i.e., the velocity approaches zero) and the velocity vector direction reverses (i.e., the racket head changes from backward retreating to forward swinging), the preset state machine confirms that the player has reached the power-charging extreme point and marks the current frame as the end time of the power-charging phase (i.e., the trigger time).

[0105] Step 3: After the end of the power-up phase, the preset state machine can check the phase detection switch instruction in the judgment control parameters. If the instruction indicates that the racket drop phase detection is turned off (e.g., macroscopic semantic features identify it as a backspin slice ball), the state machine skips the racket drop phase judgment process and directly enters the ball contact phase estimation process. If the phase detection switch instruction does not indicate that the racket drop phase detection is turned off, the following normal racket drop judgment process is executed: The preset state machine calculates the vertical height of the racket head, the vertical height of the racket-holding wrist, and the vertical acceleration of the racket head frame by frame based on the coordinates of the racket-holding wrist joint and the racket head in the key point coordinates of the human skeleton. If it is detected that the racket head height is lower than the racket-holding wrist height and the racket head has a continuous downward acceleration (i.e., the direction of the racket head's vertical acceleration is downward and this direction is maintained for multiple consecutive frames), the racket drop phase is confirmed to have been entered.

[0106] Step 4: After the end of the power-up phase (or after the end of the racket drop phase if one exists), the preset state machine defines an expected ball-touching time window based on the time window judgment condition parameter in the judgment control parameters, and presumes the entry into the ball-touching phase. Then, within this expected ball-touching time window, it checks whether there is a target video frame where the racket and ball space overlap. If so, the time corresponding to that frame is directly used as the time information for the ball-touching phase. If not, the precise ball-touching time is calculated using the physical interpolation error correction process described in the previous embodiment. Specifically: extract the frame number interval corresponding to the ball-touching phase; determine whether there is a target video frame where the racket and ball space overlap within the frame number interval; if so, use the time corresponding to the target video frame as the time information for the ball-touching phase; if not, fit the three-dimensional motion trajectory of the ball corresponding to the ball coordinates and the three-dimensional motion trajectory of the racket head corresponding to the racket head coordinates within the frame number interval, determine the intersection point of the three-dimensional motion trajectory of the ball and the three-dimensional motion trajectory of the racket head in three-dimensional space, obtain the corrected ball-touching time through time interpolation based on the time parameter corresponding to the intersection point, and update the time information for the ball-touching phase with the corrected ball-touching time.

[0107] Step 5: After the contact phase time is determined, the preset state machine calculates the velocity vector direction of the ball and the three-dimensional acceleration of the racket head based on the ball coordinates and racket head coordinates. When the following three conditions are simultaneously met, the follow-through phase is confirmed: ① The ball velocity vector is away from the racket head, that is, the direction of the ball's movement is away from the racket head, indicating that the ball has been hit; ② The three-dimensional acceleration of the racket head is negative, that is, the resultant acceleration of the racket head in three-dimensional space is negative, indicating that the racket head is decelerating; ③ The three-dimensional acceleration of the racket head exceeds the preset spin judgment condition, that is, the absolute value of the racket head acceleration exceeds the preset spin judgment threshold issued by the slow system based on the ball's spin intention.

[0108] S360: Determine the time information corresponding to each action stage based on the changes in the action stage labels corresponding to each video frame and the frame rate of the target video.

[0109] It should be understood that the constraint information in this embodiment can also be applied to the swing action segmentation process through the asynchronous update mechanism described in the previous embodiments. The specific implementation process can be referred to the above embodiments, and will not be repeated here.

[0110] The technical solution of this invention involves downsampling the target video and extracting macroscopic semantic features using a multimodal large language model. Then, it queries a semantic-parameter mapping knowledge base to transform tactical intentions into quantifiable and executable judgment and control parameters. This allows a state machine based on fixed rules to dynamically adjust the judgment threshold and stage detection switch according to scene factors such as the current action type, stance, and hitting intention. This fundamentally overcomes the shortcomings of traditional fixed-threshold state machines, which cannot recognize tactical intentions and use the same standard to handle all scenarios, resulting in poor generalization ability. Simultaneously, the state machine integrates three types of coordinate information—human skeleton key points, the ball, and the racket head—for frame-by-frame judgment. This enables precise capture of the small displacement and brief duration of the drop phase during the swing, as well as the high-speed instantaneous contact moment, significantly improving the accuracy, robustness, and scene adaptability of swing action stage segmentation in complex and variable scenarios.

[0111] Example 4 Figure 4 This is a schematic diagram of a swing motion segmentation system provided in Embodiment 4 of the present invention. Figure 4 As shown, the system includes a first subsystem 41 and a second subsystem 42 that operate asynchronously, wherein, The first subsystem 41 is used to downsample the target video to obtain sampled video frames, extract macroscopic semantic features from the sampled video frames, generate constraint information based on the macroscopic semantic features, and send the constraint information to the second subsystem 42. The second subsystem 42, connected to the first subsystem 41, is used to receive constraint information, extract multi-target spatial coordinate information from each video frame of the target video, and perform swing motion segmentation on the target video based on the constraint information and multi-target spatial coordinate information, and output the time information corresponding to multiple action stages of the swing motion; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head.

[0112] In this embodiment, the first subsystem, also known as the slow system, is primarily used to downsample the target video to obtain sampled video frames. It then extracts macroscopic tactical semantic features from these frames using a multimodal large language model, and generates constraint information based on these macroscopic tactical semantic features to regulate the action segmentation process of the second subsystem. The macroscopic tactical semantic features may include at least one of the following: action type, racket hand, hitting direction, stance, hitting spin intention, on-court position, and player physical condition. The constraint information, depending on the implementation of the second subsystem, can be a combination of a semantic condition vector and a temporal transition probability correction matrix, or it can be a set of decision control parameters. The first subsystem performs sparse frame analysis at a sampling frame rate lower than the original frame rate of the target video, focusing on maintaining relatively stable high-level semantics within a second-level time span, providing prior guidance for scene adaptation for the second subsystem.

[0113] The second subsystem, also known as the fast system, is primarily used to extract multi-target spatial coordinate information from the original frames of the target video. Guided by the constraint information generated by the first subsystem, it performs frame-by-frame temporal motion stage segmentation of the target video, outputting the time information corresponding to the initiation stage, power-gathering stage, drop stage, ball-contact stage, and follow-through stage of the swing motion. The multi-target spatial coordinate information includes at least the coordinates of key points on the human skeleton, the coordinates of the sphere, and the coordinates of the racket head. The second subsystem directly processes each frame or most frames of the target video, capturing the spatial motion features of the human body, the sphere, and the racket with millisecond-level temporal resolution. Depending on the specific implementation architecture, the second subsystem can employ a temporal deep prediction network with an integrated attention mechanism, adjusting the network's attention weights for multi-target spatial features through semantic conditional vectors and constraining the decoding process based on a temporal transition probability correction matrix; or it can employ a rule-based state machine, dynamically adjusting the judgment threshold and stage transition path through judgment control parameters issued by the slow system.

[0114] Based on the first and second subsystems provided in this embodiment, the swing motion segmentation method described in any embodiment of the present invention can be executed collaboratively. The specific implementation process can be referred to the above embodiments, and will not be repeated here.

[0115] Based on the above fast-slow thinking dual-system architecture, Figure 5 This is a flowchart of a swing motion segmentation method provided in Embodiment 4 of the present invention. Figure 5 As shown, the "slow system + fast system" architecture adopted in this embodiment first samples the user training video at a low frame rate. Then, the multimodal large language model (VLM) of the slow system extracts macroscopic tactical semantic features from the sampled frames, such as action type, racket hand, stance, ball spin intention, on-court position, and fatigue level. These features are then semantically encoded to generate a semantic conditional vector and a temporal transition probability correction matrix. Simultaneously, the target detection model of the fast system extracts the coordinates of the 3D skeleton key points of the human body, the spatial coordinates and trajectory of the tennis ball, and the 3D coordinates of the racket head from the target video frame by frame. These three coordinates are then aligned and concatenated and fused with the semantic conditional vector. The result is input into a temporal deep prediction network (such as ST-GCN or Transformer) with an integrated attention mechanism. The network outputs the original probability distribution of each frame belonging to the five stages of initiation, power generation, ball impact, and follow-through. Finally, in the decoding stage, the original probability distribution is combined with the temporal transition probability correction matrix. The optimal action stage label sequence is determined by Viterbi decoding, and the dynamic constraint process in stage S3 and the physical interpolation error correction process in stage S4 are executed. Finally, the accurate action timestamps of each stage are output, thereby achieving high-precision and robust temporal segmentation of high-speed swing actions.

[0116] The S3 stage dynamic constraint process specifically includes: when the slow system identifies a backspin slice ball, the correction matrix forces the transition probability from the power-up phase to the impact phase to zero, allowing the state to flow directly from the power-up phase to the contact phase, preventing the decoder from deadlocking when searching for the impact phase. The S4 stage physical interpolation error correction process specifically includes: if no video frame of spatial overlap between the racket and the ball is detected in the high-probability interval of the contact phase, the three-dimensional trajectory equation of the ball and the three-dimensional trajectory equation of the racket head in that interval are fitted, the intersection point of the two trajectories in three-dimensional space is calculated, and the microsecond-level actual contact time is back-calculated through time interpolation, forcibly correcting the network output.

[0117] This embodiment achieves high-precision motion segmentation for complex and ever-changing tennis training scenarios through deep fusion of macroscopic semantic understanding from a slow system and microscopic prediction from a temporal deep prediction network from a fast system. The semantic conditional vector generated by the slow system is embedded within the network's internal attention control mechanism, enabling the network to adaptively adjust the attention weights for features such as large backswings based on current tactical intentions (e.g., net volleys, baseline rallies), overcoming the generalization flaw of traditional posture models that "use the same yardstick to measure all actions." The temporal transition probability correction matrix dynamically adjusts the transition path during the decoding phase, precisely adapting to special actions without a drop shot, such as slice shots. The S4 physical trajectory intersection interpolation method transforms the capture of the contact moment from "relying on the shot to the contact frame" to "calculating based on physical laws," fundamentally solving the problem of misjudgment of the contact point caused by visual missed shots during high-speed motion.

[0118] Figure 6This is a flowchart of another swing motion segmentation method provided in Embodiment 4 of the present invention. Figure 6 As shown, the "slow system + fast system" architecture adopted in this embodiment first samples the user training video at a low frame rate. Then, the multimodal large language model (VLM) of the slow system extracts macroscopic tactical semantic features. The parameter generation and rule translation module queries the dynamic parameter library and converts the semantic features into judgment control parameters, including threshold parameters, time windows, and detection switch instructions, which are then sent to the fast system. The target detection model in the fast system extracts the spatial coordinates of the human skeleton key points, tennis ball, and racket head frame by frame. Then, the state machine executes rule-based logical judgments frame by frame according to the judgment control parameters issued by the slow system: S1 is judged by monitoring the wrist-racket head angular velocity and using a fake start retreat; S2 is judged by forcibly capturing the extreme point of racket head retreat; S3 is judged by monitoring the racket head height and downward acceleration, and selectively skips the stage detection switch instruction (such as the cut exemption instruction) to turn off the racket drop stage detection according to the instruction; S4 is located at the moment of ball contact by physical trajectory intersection interpolation; and S5 is judged by the three joint conditions of ball release, racket head deceleration, and torso rotation. This solution maintains the low computational overhead and strong interpretability of the rule system, and solves the technical pain point of poor generalization ability of traditional fixed rule state machines in complex and ever-changing scenarios by dynamically adjusting the threshold driven by the semantics of the slow system, thus realizing adaptive action segmentation under the rule scheme.

[0119] Based on the core idea of ​​this application, its scope of protection is not limited to the specific implementation method within the fast system (whether it adopts a temporal deep prediction network or a rule-based state machine), but covers any technical solution that adopts the overall architecture of "fast-slow thinking dual system". As long as the system architecture contains the following elements, it should fall within the scope of protection of this application: the first subsystem (slow system) extracts macroscopic tactical semantic features from low frame rate sampled frames and generates constraint information for dynamically adjusting the action segmentation process; the second subsystem (fast system) extracts multi-target spatial coordinate information that integrates human skeleton key points, sphere and racket head from high frame rate original video frames; and the second subsystem performs frame-by-frame temporal action stage segmentation of the target video under the guidance of the constraint information, and finally outputs the time information of multiple predetermined stages of the swing action. Regardless of whether the fast system uses neural network prediction, rule-based state machine decision-making, or a combination of both, regardless of whether the specific form of the constraint information is a semantic condition vector and transition probability correction matrix, or a decision control parameter and detection switch instruction, and regardless of whether the slow system uses a multimodal large language model or other semantic understanding model, as long as it embodies the dual-system collaborative architecture of "slow system macro-semantic understanding guiding fast system micro-action segmentation", it constitutes an implementation of this application.

[0120] It is worth noting that the physical deployment of the first subsystem (slow system) and the second subsystem (fast system) in this embodiment can be flexibly selected according to the actual application scenario. Specifically, they can be integrated into the same computing device, such as deployed on the same high-performance server or edge computing device, and the transmission of constraint information and state synchronization can be achieved through internal inter-process communication or shared memory mechanisms; or they can be configured in two independent devices, for example, the first subsystem can be deployed on a cloud server to utilize its powerful large model inference capabilities, and the second subsystem can be deployed on a computing device close to the edge of the video acquisition source to achieve low-latency real-time multi-target coordinate extraction and action segmentation, with the two asynchronously transmitting constraint information and state data through a network connection. This embodiment does not limit the specific deployment form of the first and second subsystems, as long as they can work together to complete the swing action segmentation method.

[0121] Example 5 Figure 7 This is a schematic diagram of a racket swing motion segmentation device provided in Embodiment 5 of the present invention. Figure 7 As shown, the device includes: The downsampling module 51 is used to downsample the target video to obtain sampled video frames; The semantic extraction module 52 is used to extract macroscopic semantic features from the sampled video frames and generate constraint information based on the macroscopic semantic features; The coordinate extraction module 53 is used to extract multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head; The action segmentation module 54 is used to segment the swing action of the target video based on constraint information and multi-target spatial coordinate information, and output the time information corresponding to multiple action stages of the swing action.

[0122] Furthermore, based on the above embodiments of the invention, the downsampling module 51 includes at least one of the following: The first sampling unit is used to extract video frames from the target video at preset time intervals as sampled video frames. The second sampling unit is used to extract a preset number of video frames before and after the time of the preset trigger event from the target video when a preset trigger event is detected, and use them as sampled video frames.

[0123] Furthermore, based on the above embodiments of the invention, the semantic extraction module 52 includes: The semantic extraction unit is used to input the sampled video frames into a preset multimodal large language model to obtain macro-semantic features; the macro-semantic features include at least one of the following: action type, racket hand, hitting direction, stance, hitting spin intention, on-court position, and player's physical condition; The first constraint information determination unit is used to encode macroscopic semantic features, generate semantic condition vectors and temporal transition probability correction matrices, and use the semantic condition vectors and temporal transition probability correction matrices as constraint information.

[0124] Furthermore, based on the above embodiments of the invention, the coordinate extraction module 53 includes: The coordinate extraction unit is used to sequentially input each video frame of the target video into the preset target detection model to obtain the coordinates of the human skeleton key points, the sphere coordinates, and the racket head coordinates in each video frame.

[0125] Furthermore, based on the above embodiments of the invention, the action segmentation module 54 includes: The coordinate stitching unit is used to align and stitch together the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head on the time axis to obtain multi-target spatial features. The feature fusion unit is used to fuse semantic condition vectors with multi-objective spatial features to obtain a fused feature tensor. The first action stage determination unit is used to input the fused feature tensor into a preset temporal depth prediction network to obtain the original probability distribution of each video frame in the target video belonging to each action stage; the action stages include the initiation stage, the charging stage, the drop stage, the ball contact stage, and the follow-through stage; The first time information determination unit is used to determine the time information corresponding to multiple action stages of the swing action based on the original probability distribution and the time sequence transition probability correction matrix.

[0126] Furthermore, based on the above embodiments, the preset temporal depth prediction network integrates an attention mechanism, and the feature fusion unit is specifically used for: Based on semantic conditional vectors, attention weights corresponding to each spatial dimension in multi-objective spatial features are generated through an attention mechanism. The multi-object spatial features are weighted element-wise based on attention weights to obtain the fused feature tensor.

[0127] Furthermore, based on the above embodiments of the invention, the first time information determination unit is specifically used for: Obtain the preset baseline state transition probability matrix; the elements in the baseline state transition probability matrix represent the initial transition probability between two stages in multiple action stages; The corrected state transition probability matrix is ​​obtained by superimposing the time-series transition probability correction matrix with the baseline state transition probability matrix. The corrected state transition probability matrix is ​​used as the transition probability constraint. The preset sequence decoding algorithm is called according to the original probability distribution of each video frame in the target video to determine the action stage label corresponding to each video frame. Based on the changes in the action stage labels corresponding to each video frame and the frame rate of the target video, the time information corresponding to each action stage is determined.

[0128] Furthermore, based on the above embodiments of the invention, the semantic extraction module 52 includes: The semantic extraction unit is used to input the sampled video frames into a preset multimodal large language model to obtain macro-semantic features; the macro-semantic features include at least one of the following: action type, racket hand, hitting direction, stance, hitting spin intention, on-court position, and player's physical condition; The second constraint information determination unit is used to query a preset semantic-parameter mapping knowledge base, obtain the judgment control parameters corresponding to the macroscopic semantic features, and use the judgment control parameters as constraint information; wherein, the judgment control parameters include at least one of the following: rotation angle judgment condition parameters, time window judgment condition parameters, speed judgment condition parameters, and stage detection switch instructions.

[0129] Furthermore, based on the above embodiments of the invention, the action segmentation module 54 includes: The state machine configuration unit is used to configure the judgment conditions and stage detection switches of the preset state machine according to the judgment control parameters. The second action stage determination unit is used to call a preset state machine to perform frame-by-frame action stage determination on the target video according to the multi-target spatial coordinate information, so as to determine the action stage label corresponding to each video frame in the target video. The second time information determination unit is used to determine the time information corresponding to each action stage according to the changes in the action stage labels corresponding to each video frame and the frame rate of the target video; the action stages include the start-up stage, the charging stage, the drop stage, the ball contact stage, and the follow-through stage.

[0130] Furthermore, based on the above embodiments of the invention, the second action stage determining unit is specifically used for: The angular velocity of the angle between the player's wrist and the racket head is determined based on the coordinates of key points of the human skeleton and the coordinates of the racket head in at least two video frames in the target video. When the angular velocity of the angle exceeds the velocity judgment condition in multiple consecutive frames, the start-up phase is confirmed. The racket head motion speed and velocity vector direction are determined based on the racket head coordinates of at least two video frames in the target video. When the racket head motion speed is detected to be less than a preset speed threshold and the velocity vector direction is reversed, the end time of the power accumulation phase is confirmed. After the power-up phase ends, if the phase detection switch command indicates that the racket drop phase detection is turned off, the determination of the racket drop phase is skipped and the inference of the ball contact phase is directly executed. Otherwise, the racket head height, racket wrist height and racket head acceleration are determined based on the coordinates of the human skeleton key points and the racket head coordinates of at least two video frames in the target video. When the racket head height is detected to be lower than the racket wrist height and the racket head acceleration is a continuous downward acceleration, the racket drop phase is confirmed. After the power-accumulation phase or the drop phase, the ball contact phase is presumed based on the time window determination condition parameters; if a video frame of the racket and the ball overlapping in space is detected before the presumed ball contact phase, the time corresponding to the video frame is used as the time information of the ball contact phase. The ball velocity vector and racket head three-dimensional acceleration are determined based on the ball coordinates and racket head coordinates of at least two video frames in the target video. When it is detected that the ball velocity vector is away from the racket head, the racket head three-dimensional acceleration is negative, and the racket head three-dimensional acceleration exceeds the preset rotation judgment condition, the follow-through phase is confirmed.

[0131] Furthermore, based on the above embodiments of the invention, the multiple action phases include a ball contact phase, and the racket swing action segmentation device further includes a ball contact phase time correction module, specifically used for: Extract the frame number range corresponding to the ball contact phase; Determine whether there are target video frames within the frame sequence interval where the racket and the ball spatially overlap; If it exists, the time corresponding to the target video frame will be used as the time information of the ball-touching phase; If not, the three-dimensional motion trajectory of the ball corresponding to the ball coordinates and the three-dimensional motion trajectory of the racket head corresponding to the racket head coordinates within the frame sequence interval are fitted. The intersection point of the three-dimensional motion trajectory of the ball and the three-dimensional motion trajectory of the racket head in three-dimensional space is determined. Based on the time parameter corresponding to the intersection point, the corrected ball contact time is obtained by time interpolation, and the time information of the ball contact stage is updated with the corrected ball contact time.

[0132] Furthermore, based on the above embodiments of the invention, constraint information is applied to the swing motion segmentation process through an asynchronous update mechanism, which includes: A globally shared state board is maintained in memory; the globally shared state board is used to store constraint information. When processing each video frame of the target video, read the constraint information stored in the global shared state panel, and use the constraint information to guide the swing action segmentation of the current video frame; After generating new constraint information based on the sampled video frames, the new constraint information is written to the global shared state board. For the same swing cycle, the constraint information in the globally shared state board remains locked after being written until it is restored to the preset baseline state after the current swing cycle ends.

[0133] The swing motion segmentation device provided in this embodiment of the invention can execute the swing motion segmentation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0134] Example 6 Figure 8 A schematic diagram of an electronic device 60 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0135] like Figure 8 As shown, the electronic device 60 includes at least one processor 61 and a memory, such as a read-only memory (ROM) 62 and a random access memory (RAM) 63, communicatively connected to the at least one processor 61. The memory stores computer programs executable by the at least one processor. The processor 61 can perform various appropriate actions and processes based on the computer program stored in the ROM 62 or loaded from storage unit 68 into the RAM 63. The RAM 63 can also store various programs and data required for the operation of the electronic device 60. The processor 61, ROM 62, and RAM 63 are interconnected via a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0136] Multiple components in electronic device 60 are connected to I / O interface 65, including: input unit 66, such as keyboard, mouse, etc.; output unit 67, such as various types of monitors, speakers, etc.; storage unit 68, such as disk, optical disk, etc.; and communication unit 69, such as network card, modem, wireless transceiver, etc. Communication unit 69 allows electronic device 60 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0137] Processor 61 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 61 performs the various methods and processes described above, such as the swing motion segmentation method.

[0138] In some embodiments, the swing motion segmentation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 68. In some embodiments, part or all of the computer program may be loaded into and / or installed on electronic device 60 via ROM 62 and / or communication unit 69. When the computer program is loaded into RAM 63 and executed by processor 61, one or more steps of the swing motion segmentation method described above may be performed. Alternatively, in other embodiments, processor 61 may be configured to perform the swing motion segmentation method by any other suitable means (e.g., by means of firmware).

[0139] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] In some embodiments, the swing motion segmentation method may be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the swing motion segmentation method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program may be executed entirely on a machine, partially on a machine, partially on a remote machine as a standalone software package, or entirely on a remote machine or server.

[0141] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0143] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0144] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0145] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0146] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for segmenting racket swing motions, characterized in that, The method includes: The target video is downsampled to obtain sampled video frames; Macro-level semantic features are extracted from the sampled video frames, and constraint information is generated based on the macro-level semantic features; the macro-level semantic features are semantic information extracted from the video frame that describes the high-level tactical intent and overall motion attributes of the swing action; Extract multi-target spatial coordinate information from each video frame of the target video; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head; Based on the constraint information and the multi-target spatial coordinate information, the target video is segmented into swing motions, and the time information corresponding to multiple action stages of the swing motion is output.

2. The method according to claim 1, characterized in that, The downsampling of the target video to obtain sampled video frames includes at least one of the following: Video frames are extracted from the target video at preset time intervals and used as the sampled video frames; When a preset trigger event is detected, a preset number of video frames are extracted from the target video before and after the time when the preset trigger event occurs, and these are used as the sampled video frames.

3. The method according to claim 1, characterized in that, The step of extracting macroscopic semantic features from the sampled video frames and generating constraint information based on the macroscopic semantic features includes: The sampled video frames are input into a preset multimodal large language model to obtain the macro-semantic features; the macro-semantic features include at least one of the following: action type, racket hand, hitting direction, stance, hitting spin intention, on-court position, and player's physical condition; The macroscopic semantic features are encoded to generate a semantic condition vector and a temporal transition probability correction matrix, and the semantic condition vector and the temporal transition probability correction matrix are used as the constraint information. The step of encoding the macroscopic semantic features to generate a semantic condition vector and a temporal transition probability correction matrix includes: Each of the macro-semantic features is classified and encoded using either One-hot encoding or Multi-hot encoding. The encoded vectors corresponding to each of the macro-semantic features are then concatenated in a preset order to obtain a one-dimensional semantic condition vector. Based on the macro-semantic features and the preset semantic-correction mapping rules, the transition probability correction values ​​between each action stage are determined, and the temporal transition probability correction matrix with the same dimension as the preset baseline state transition probability matrix is ​​generated according to the transition probability correction values.

4. The method according to claim 1, characterized in that, Extracting multi-target spatial coordinate information from each video frame of the target video includes: Each video frame of the target video is sequentially input into a preset target detection model to obtain the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head in each video frame.

5. The method according to claim 3, characterized in that, The step of segmenting the target video based on the constraint information and the multi-target spatial coordinate information, and outputting the time information corresponding to multiple action stages of the swing action, includes: The coordinates of the key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head are aligned and stitched together along the time axis to obtain multi-target spatial features; The semantic condition vector is fused with the multi-objective spatial features to obtain a fused feature tensor; The fused feature tensor is input into a preset temporal depth prediction network to obtain the original probability distribution of each video frame in the target video belonging to each action stage; the action stages include the initiation stage, the charging stage, the drop stage, the ball contact stage, and the follow-through stage; Based on the original probability distribution and the time transition probability correction matrix, the time information corresponding to multiple action stages of the swing action is determined.

6. The method according to claim 5, characterized in that, The preset temporal depth prediction network integrates an attention mechanism. The step of fusing the semantic conditional vector with the multi-objective spatial features to obtain a fused feature tensor includes: Based on the semantic condition vector, attention weights corresponding to each spatial dimension in the multi-object spatial features are generated through the attention mechanism. The multi-target spatial features are weighted element-wise based on the attention weights to obtain the fused feature tensor.

7. The method according to claim 5, characterized in that, The determination of the time information corresponding to multiple action stages of the swing action based on the original probability distribution and the time transition probability correction matrix includes: Obtain a preset baseline state transition probability matrix; the elements in the baseline state transition probability matrix represent the initial transition probability between two stages in the plurality of action stages; The time-series transition probability correction matrix is ​​superimposed with the baseline state transition probability matrix to obtain the corrected state transition probability matrix; Using the corrected state transition probability matrix as a transition probability constraint, a preset sequence decoding algorithm is called according to the original probability distribution of each video frame in the target video to determine the action stage label corresponding to each video frame; Based on the changes in the action stage labels corresponding to each video frame and the frame rate of the target video, the time information corresponding to each action stage is determined.

8. The method according to claim 1, characterized in that, The step of extracting macroscopic semantic features from the sampled video frames and generating constraint information based on the macroscopic semantic features includes: The sampled video frames are input into a preset multimodal large language model to obtain the macroscopic semantic features; the macroscopic semantic features include at least one of the following: action type, racket hand, hitting direction, stance, hitting spin intention, on-court position, and player's physical condition; The system queries a preset semantic-parameter mapping knowledge base to obtain the judgment control parameters corresponding to the macroscopic semantic features, and uses the judgment control parameters as the constraint information; wherein, the judgment control parameters include at least one of the following: rotation angle judgment condition parameters, time window judgment condition parameters, speed judgment condition parameters, and stage detection switch instructions.

9. The method according to claim 8, characterized in that, The step of segmenting the target video based on the constraint information and the multi-target spatial coordinate information, and outputting the time information corresponding to multiple action stages of the swing action, includes: Configure the judgment conditions and stage detection switches of the preset state machine according to the judgment control parameters; The preset state machine is invoked to determine the action stage of the target video frame by frame according to the multi-target spatial coordinate information, so as to determine the action stage label corresponding to each video frame in the target video; Based on the changes in the action stage labels corresponding to each video frame and the frame rate of the target video, the time information corresponding to each action stage is determined; the action stage includes the start-up stage, the charging stage, the drop stage, the ball contact stage, and the follow-through stage.

10. The method according to claim 9, characterized in that, The step of calling the preset state machine to perform frame-by-frame action stage determination on the target video according to the multi-target spatial coordinate information includes: The angular velocity of the angle between the player's wrist and the racket head is determined based on the coordinates of the key points of the human skeleton and the coordinates of the racket head in at least two video frames in the target video. When the angular velocity of the angle exceeds the speed determination condition for multiple consecutive frames, the player is confirmed to enter the startup phase. The racket head movement speed and velocity vector direction are determined based on the racket head coordinates of at least two video frames in the target video. When the racket head movement speed is detected to be less than a preset speed threshold and the velocity vector direction is reversed, the end time of the power accumulation phase is confirmed. After the power-accumulation phase ends, if the phase detection switch command indicates that the racket drop phase detection is turned off, the determination of the racket drop phase is skipped, and the inference of the ball contact phase is directly executed. Otherwise, the racket head height, racket wrist height, and racket head acceleration are determined based on the coordinates of the human skeleton key points and the racket head coordinates of at least two video frames in the target video. When the racket head height is detected to be lower than the racket wrist height and the racket head acceleration is a continuous downward acceleration, the racket drop phase is confirmed. After the power-accumulation phase or the racket drop phase ends, the ball contact phase is presumed based on the time window determination condition parameters; wherein, if a video frame of spatial overlap between the racket and the ball is detected before the presumed ball contact phase, the time corresponding to the video frame is used as the time information of the ball contact phase; Based on the ball coordinates and racket head coordinates of at least two video frames in the target video, the ball velocity vector and racket head three-dimensional acceleration are determined. When it is detected that the ball velocity vector is away from the racket head, the racket head three-dimensional acceleration is negative, and the racket head three-dimensional acceleration exceeds the preset rotation judgment condition, the follow-through phase is confirmed.

11. The method according to claim 1, characterized in that, The multiple action phases include a ball-touching phase, and the method further includes: Extract the frame number range corresponding to the ball-touching phase; Determine whether there is a target video frame within the frame number interval where the racket and the ball spatially overlap; If it exists, the time corresponding to the target video frame is used as the time information of the ball-touching phase; If not, fit the three-dimensional motion trajectory of the ball corresponding to the ball coordinates and the three-dimensional motion trajectory of the racket head corresponding to the racket head coordinates within the frame number interval, determine the intersection point of the three-dimensional motion trajectory of the ball and the three-dimensional motion trajectory of the racket head in three-dimensional space, obtain the corrected ball contact time through time interpolation based on the time parameter corresponding to the intersection point, and update the time information of the ball contact stage with the corrected ball contact time.

12. The method according to claim 1, characterized in that, The constraint information is applied to the swing motion segmentation process through an asynchronous update mechanism, which includes: A globally shared state board is maintained in memory; the globally shared state board is used to store the constraint information. When processing each video frame of the target video, the constraint information stored in the global shared state panel is read, and the current video frame is guided to perform swing action segmentation based on the constraint information; After generating new constraint information based on the sampled video frames, the new constraint information is written into the global shared state board. Specifically, for the same swing action cycle, the constraint information in the global shared state board remains locked after being written until it is restored to the preset baseline state after the current swing action cycle ends.

13. A swing motion segmentation system, characterized in that, For executing the swing motion segmentation method according to any one of claims 1-12, the system comprises a first subsystem and a second subsystem operating asynchronously, wherein... The first subsystem is used to downsample the target video to obtain sampled video frames, extract macroscopic semantic features from the sampled video frames, generate constraint information based on the macroscopic semantic features, and send the constraint information to the second subsystem; the macroscopic semantic features are semantic information extracted from the video frame that describes the high-level tactical intent and overall motion attributes of the swing action; The second subsystem, connected to the first subsystem, is used to receive the constraint information, extract multi-target spatial coordinate information from each video frame of the target video, and perform swing motion segmentation on the target video based on the constraint information and the multi-target spatial coordinate information, and output the time information corresponding to multiple action stages of the swing motion; the multi-target spatial coordinate information includes at least the coordinates of key points of the human skeleton, the coordinates of the sphere, and the coordinates of the racket head.

14. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the swing motion segmentation method according to any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the swing motion segmentation method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Sequential action segmentation method of athlete skeleton points based on computer vision

    CN120032293A

  • Tongue picture image segmentation method and system and storage medium

    CN120235896A