A computer vision-based intelligent tennis serve training method and system

The computer vision-based intelligent tennis serve training system utilizes laser guidance and AI biomechanical analysis to decouple the serve motion into sub-techniques, providing real-time, quantitative, and visual guidance. This solves the problems of untimely feedback, subjective analysis, and uncontrollable training variables in existing technologies, achieving efficient skill improvement and popularization.

CN122624879APending Publication Date: 2026-08-25ALIBABA SPORTS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610891789.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies lack training solutions in tennis serve training that balance professionalism, real-time performance, non-invasiveness, intuitiveness, and affordability, leading to problems such as untimely feedback, subjective analysis, uncontrollable training variables, and high costs.

Method used

The intelligent tennis serve training system, based on computer vision, uses laser guidance, 3D depth imaging, and AI biomechanical analysis to decouple the serve action into sub-technical actions, providing real-time, quantitative, and visualized professional guidance, and achieving non-invasive full-body motion data acquisition and diagnosis.

Benefits of technology

It achieves efficient skill enhancement by breaking away from inefficient repetition through instant feedback and personalized training, providing scientific and accurate motion analysis, reducing training costs, and improving training efficiency and accessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122624879A_ABST
    Figure CN122624879A_ABST
Patent Text Reader

Abstract

The application provides a computer vision-based intelligent tennis serving training method and system, which comprises the following steps: calibrating laser guiding parameters of a guiding device according to key physiological data of a trainer; decoupling a tennis serving action into multiple sub-technical action training items, and displaying each sub-technical action training item for the trainer to select; extracting target control parameters matched with the sub-technical action training item selected by the trainer from the laser guiding parameters, and controlling the guiding device to project laser guidance according to the target control parameters to guide the trainer to perform the corresponding sub-technical action according to the guidance; acquiring image sequence data of the trainer when performing the sub-technical action, constructing a human motion dynamic model according to the image sequence data, and determining key action parameters of the trainer; and after the trainer performs the sub-technical action, generating a motion video superimposed and rendered with the human motion dynamic model and the key action parameters for analysis result display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sports training and analysis technology, and in particular to an intelligent tennis serve training method and system based on computer vision. Background Technology

[0002] Tennis, as a popular sport, involves complex techniques and a long learning curve. Among all the technical aspects, the serve is widely recognized as the most difficult to master, involving a series of highly coordinated and complex movements such as stance, ball toss, push-off, body rotation, racket swing, and shot. The quality of the serve directly determines the course of the match; therefore, efficient serve training is crucial for improving a player's skill level.

[0003] Currently, the mainstream tennis serve training methods and their technical status mainly include the following: First, traditional manual instruction and extensive repetitive practice. This is the most basic and common training method. Trainees develop muscle memory through hundreds or thousands of repetitions under the verbal guidance of a coach. This method has the following drawbacks: Drawback 1: Subjectivity and delay in feedback. A coach's instruction relies heavily on their personal experience, limiting their observation angle and making it difficult to accurately and objectively quantify subtle errors in high-speed movements (such as a fraction of a second of incorrect shoulder rotation timing or a few centimeters of ball throwing deviation). Furthermore, feedback is usually delayed; trainees receive evaluation only after completing a set of movements, failing to obtain corrective information at the moment of execution. This leads to repeated reinforcement of incorrect movements, making correction difficult. Drawback 2: High cost and scarce resources. Hiring a professional coach for one-on-one instruction is expensive, a huge financial burden for most amateur enthusiasts. Excellent coaching resources are already scarce, resulting in extremely low accessibility for professional instruction.

[0004] Second, using ordinary video equipment for video playback analysis. With the widespread use of smartphones and portable cameras, many trainees record their own training videos and observe and reflect on their movements through slow-motion playback. This training method has the following drawbacks: Drawback 1: Lack of real-time and interactivity. The feedback loop of this method is very long. Trainees need to complete the training first and then spend time watching and comparing the video. The whole process is a complete separation of "training" and "analysis". When problems are discovered, the training enthusiasm and physical state may have changed, making it impossible to make targeted adjustments immediately, resulting in low training efficiency. Drawback 2: Single analytical dimension and lack of professional benchmarks. Ordinary 2D videos can only provide a single-plane perspective and cannot present the three-dimensional spatial relationship of the movement. For example, key biomechanical information such as the body's rotation range and the angle of hip-shoulder separation is completely lost. More importantly, ordinary enthusiasts lack a professional knowledge system. Even if they see their own movements, they cannot accurately judge where the problems are, let alone know where the gap is between their movements and those of professional athletes. The analysis is superficial, and the improvement effect is limited. Drawback 3: Unable to solve the problem of "constant variables" in training, leading to inefficient repetition. The tennis serve is a multi-variable coupled action, where the "stability of the ball toss" is the benchmark for all subsequent movements (pushing off, rotating the body, and swinging the racket). The core challenge for beginners lies in the inconsistency of the ball toss position each time. This means that even if a problem is identified in a particular shot through video replay (e.g., the hitting point is too low), it's difficult for the trainee to replicate the previous scenario for targeted correction in the next practice session because the ball toss position may have changed. The "input" (ball toss) in each practice session is unstable, making it difficult to correct and solidify the "output" (striking motion). This training model, unable to control key variables, is the root cause of why trainees need hundreds or thousands of blind practice sessions to find even a semblance of a feel for the game, resulting in a significant waste of time and energy.

[0005] Third, motion analysis systems based on wearable sensors. Some smart rackets or wearable devices (such as armbands and wristbands) integrating sensors like accelerometers and gyroscopes have appeared on the market to collect motion data, such as swing speed and racket face angle. Disadvantage 1: Intrusive experience and data limitations. These devices require trainees to wear or use specific equipment, altering their original exercise habits and creating a foreign body sensation, potentially affecting the natural flow of movements. Furthermore, their data collection is limited to the local limb where the sensor is located (such as the arm or racket), failing to acquire full-body, interconnected data, thus unable to analyze the fundamental nature of the serve—the complete kinetic chain starting from the legs. For example, it cannot determine whether the power generation sequence is correct (whether the legs drive the waist or the arm is actively generating power). Disadvantage 2: Lack of intuitive feedback on movement form. These devices provide abstract numbers (such as a swing speed of 90 km / h) but cannot tell the trainee "why" they didn't hit well. They cannot intuitively show "your toss position was 10 centimeters to the left" or "your knee flexion angle was insufficient," lacking visual diagnosis and guidance on the movement form itself.

[0006] Fourth, professional sports science laboratory equipment. Top sports academies or research institutions use optical motion capture systems (such as Vicon and OptiTrack) with multiple high-speed infrared cameras and mechanical platforms (force gauges) to perform extremely precise three-dimensional biomechanical analysis of trainees. Disadvantages: Extremely expensive, environmentally restricted, and complex to operate. These systems cost hundreds of thousands or even millions of dollars, require specialized indoor spaces free from sunlight interference, and involve attaching numerous reflective markers to the trainee's body. The entire testing process requires specialized technicians and is completely unsuitable for everyday training scenarios; its purpose is "research" rather than "training."

[0007] In summary, there is a significant market gap and technical pain point in the field of tennis serve training: the lack of a training solution that can balance professionalism, real-time performance, non-invasiveness, intuitiveness, and affordability. Summary of the Invention

[0008] In view of the above problems, the present invention is proposed to provide a computer vision-based intelligent tennis serve training method and system that solves or at least partially solves the above technical problems.

[0009] One aspect of the present invention provides a computer vision-based intelligent tennis serve training method, the method comprising: Acquire key physiological data of the trainee, and automatically calibrate the laser guidance parameters of the guidance device based on the key physiological data; The tennis serve motion is decoupled into multiple sub-technical training items, and each sub-technical training item is displayed based on a preset human-computer interaction interface for trainees to choose from. The target control parameters that match the sub-technical movement training item selected by the trainee are extracted from the laser guidance parameters. The control device projects laser guidance according to the target control parameters to guide the trainee to perform the corresponding sub-technical movement in accordance with the guidance. Acquire image sequence data of the trainee performing sub-technique actions, construct a human motion dynamic model based on the image sequence data, and determine the trainee's key action parameters; After the trainee completes the sub-technical action, a motion video with a superimposed rendering of the human motion dynamic model and the key action parameters is generated for the analysis results to be displayed.

[0010] Furthermore, the method also includes: The human motion dynamic model is compared with a preset expert motion model. Based on the motion comparison results and the key motion parameters, problems in the trainee's sub-technical motion training are diagnosed, and the diagnosis results and improvement suggestions are highlighted and / or broadcast by voice.

[0011] Furthermore, a human motion dynamic model is constructed based on the image sequence data, and key movement parameters of the trainee are determined, including: Identify key human joints in each image frame of the image sequence data to extract three-dimensional skeletal points, and construct a human motion dynamic model based on the temporal changes of each three-dimensional skeletal point. Based on the extraction results of each three-dimensional skeletal point, the hip-shoulder separation angle formed by hip rotation and shoulder rotation during the movement cycle is calculated, and the trunk twisting amplitude and the body's power transmission efficiency are determined according to the hip-shoulder separation angle. Calculate the angle changes and angular velocity timing of each three-dimensional skeletal point during the movement process. Analyze the order of force exertion and force transmission efficiency of each joint from the lower limbs and trunk to the upper limbs based on the angle changes and angular velocity timing. Determine whether the order of force exertion of each limb is correct based on the order of force exertion and force transmission efficiency of each joint.

[0012] Furthermore, constructing a human motion dynamic model based on the image sequence data and determining the trainee's key movement parameters also includes: Identify the tennis ball in each image frame of the image sequence data, fit the three-dimensional motion trajectory of the ball based on the continuous motion coordinates of the tennis ball, compare the three-dimensional motion trajectory of the ball with the preset standard trajectory, and obtain the trajectory spatial offset and the spatial error of the highest point of the ball throw. The system identifies the three-dimensional spatial coordinates of the tennis ball and the hitting tool at the moment of impact, locates the actual hitting position, and compares the actual hitting position with the standard ideal hitting point to calculate the deviation of the hitting point position.

[0013] Furthermore, the method also includes: After completing a set of training exercises for the selected sub-technical movements, a comprehensive report is generated, including a capability radar chart, progress curves for key movement parameters, and a summary of problems. The system also automatically recommends the training mode for the next stage.

[0014] Another aspect of the present invention provides an intelligent tennis serve training system based on computer vision, the system comprising a guidance and ball supply subsystem, a motion execution and acquisition subsystem, a human-computer interaction and feedback unit, and a central processing and analysis unit; The ball guidance and feeding subsystem is used to project laser guidance marks into the training area according to the control instructions of the central processing and analysis unit, and to store and feed tennis balls. The human-computer interaction and feedback unit is used to display multiple sub-technical movement training items decoupled from the tennis serve motion for trainees to choose from. The motion execution and acquisition subsystem is used to capture image sequence data of the trainee performing the selected sub-technique in a non-contact manner, and transmit the acquired image sequence data to the central processing and analysis unit in real time. A central processing and analysis unit is used to execute the computer vision-based intelligent tennis serve training method as described in any one of claims 1-4; The human-computer interaction and feedback unit is also used to receive and display motion videos sent by the central processing and analysis unit, which are overlaid with a human motion dynamic model and the key motion parameters.

[0015] Furthermore, the central processing and analysis unit is also used to compare the human motion dynamic model with the preset expert motion model, combine the motion comparison results and the key motion parameters to diagnose problems in the trainee's sub-technical motion training, and highlight and / or broadcast the diagnosis results and improvement suggestions.

[0016] Furthermore, the guiding and ball-supplying subsystem includes: The laser projection module includes a miniature laser projector for projecting a virtual hitting window and / or a position guide cursor onto the training area according to control instructions from the central processing and analysis unit. The storage and gravity ball supply module includes a ball storage compartment and a ball delivery slide, which is used to transport tennis balls in the ball storage compartment to the release buffer through the ball delivery slide using the principle of gravity; A buffer to be released, used to load and position tennis balls for free-fall training from the ball release chute; A tennis ball release device is used to release a tennis ball from a release buffer according to control instructions from a central processing and analysis unit.

[0017] Furthermore, the motion execution and acquisition subsystem includes at least one 3D depth camera and at least one high-speed camera, used to capture image sequence data when the trainee performs the selected sub-technique action, and to transmit the acquired image sequence data to the central processing and analysis unit in real time.

[0018] Furthermore, the human-computer interaction and feedback unit includes: The visualization module, implemented using a touch screen, is used to display multiple sub-technical movement training items decoupled from the tennis serve motion, and to monitor the trainee's touch selection commands; as well as to display motion videos sent by the central processing and analysis unit, which are overlaid with and rendered with the human motion dynamic model and the key motion parameters. The voice broadcast module is used to broadcast diagnostic results and improvement suggestions via voice.

[0019] The computer vision-based intelligent tennis serve training method and system provided in this invention can automatically calibrate the laser guidance parameters of the guidance device according to the trainee's key physiological data to achieve personalized guidance training for the trainee. By "decoupling" complex movements and controlling and guiding key variables such as ball toss position and hitting point, it breaks the vicious cycle of inefficient repetition and achieves truly targeted and efficient key movement practice. It can also achieve non-invasive, whole-body motion data collection of the trainee's complete technical movements, ensuring more comprehensive and accurate training analysis and achieving efficient skill improvement.

[0020] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings: Figure 1 This is a structural block diagram of a computer vision-based intelligent tennis serve training system according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the physical structure of the computer vision-based intelligent tennis serve training system according to an embodiment of the present invention. Figure 3This is a flowchart of a computer vision-based intelligent tennis serve training method according to an embodiment of the present invention. Detailed Implementation

[0022] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0023] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined.

[0024] To overcome the numerous shortcomings of existing technologies, such as untimely feedback, subjective analysis, uncontrollable training variables, and high costs, this invention provides an intelligent tennis serve training system based on computer vision. This invention utilizes non-invasive visual acquisition technology combined with an artificial intelligence biomechanical analysis engine to "decouple" the complex serve motion into independently trainable sub-motions, providing trainees with real-time, quantitative, and visualized professional guidance. This breaks through the dilemma of inefficient and repetitive training, achieving highly efficient skill improvement.

[0025] like Figure 1 As shown, the intelligent tennis serve training system based on computer vision proposed in this embodiment of the invention includes a guidance and ball supply subsystem 10, a motion execution and acquisition subsystem 20, a central processing and analysis unit 30, and a human-computer interaction and feedback unit 40, wherein: The ball guidance and feeding subsystem 10 is used to project laser guidance marks in the training area according to the control instructions of the central processing and analysis unit, and to store and feed tennis balls. The human-computer interaction and feedback unit 40 is used to display multiple sub-technical movement training items obtained by decoupling the tennis serve motion for trainees to choose from. The motion execution and acquisition subsystem 20 is used to capture image sequence data of the trainee performing the selected sub-technique in a non-contact manner, and transmit the acquired image sequence data to the central processing and analysis unit in real time. The central processing and analysis unit 30 is used to execute the steps of the computer vision-based intelligent tennis serve training method; The human-computer interaction and feedback unit 40 is also used to receive and display a motion video sent by the central processing and analysis unit 30, which is overlaid with a human motion dynamic model and the key motion parameters.

[0026] In one specific embodiment of the present invention, the system further includes a physical structure body, which is a cantilever (L-shaped) or portal frame structure. The guiding and ball supply subsystem 10, the motion execution and acquisition subsystem 20, the central processing and analysis unit 30, and the human-computer interaction and feedback unit 40 are deployed on the physical structure body, as detailed below. Figure 2 As shown.

[0027] In this embodiment, the guiding and ball-feeding subsystem 10 provides trainees with clear action target references and a continuous supply of training balls, which is key to solving the problem of "uncontrollable training variables" in the prior art. This subsystem receives control commands from the central processing and analysis unit 30 and mainly includes a laser projection module 11, a storage and gravity ball-feeding module 12, a tennis ball release device 13, and a release buffer 14. The laser projection module 11 includes a miniature laser projector used to project a virtual hitting window and / or a positioning guide cursor onto the training area according to control instructions from the central processing and analysis unit. The laser projection module 11 is typically a miniature laser projector mounted on top of the physical structure, capable of projecting a virtual hitting window of adjustable size and height in the air, and projecting a positioning guide cursor on the ground.

[0028] The storage and gravity ball supply module 12 includes a ball storage chamber and a ball delivery slide, which is used to transport the tennis balls in the ball storage chamber to the release buffer through the ball delivery slide using the principle of gravity. The storage and gravity ball supply module 12 is usually an inclined ball delivery slide fixed on the support leg. It uses the principle of gravity to transport the tennis balls in the ball storage chamber to the release buffer (14). Release buffer 14 is used to load and position a tennis ball for free-fall training from the ball release slide. Release buffer 14 can be a small storage chamber located directly above the tennis ball release device 13 for preloading and positioning a tennis ball that will be used for free-fall training.

[0029] The tennis ball release device 13 is used to release the tennis ball from the release buffer according to the control command of the central processing and analysis unit. The tennis ball release device 13 can be an electromagnetic control gate installed below the release buffer 14, which can release the tennis ball in the buffer on command.

[0030] In this invention, the laser projection unit 11 provides a constant hitting target, ensuring a uniform standard for each swing practice by the trainee. The release buffer 14 and the tennis ball release device 13 work together to provide a "perfect ball toss" in the "hit point perception training" mode. When this mode is selected, the system automatically diverts a ball from the gravity ball supply module 12 to the release buffer 14 for preparation. This design "decouples" the complex serving motion, fundamentally solving the core pain point of inefficient training caused by unstable ball tosses.

[0031] Understandably, this system can be expanded to analyze all tennis techniques, including forehand, backhand, volley, overhead smash, and footwork, by adding software models.

[0032] The core algorithm and hardware architecture of this invention are universal and can be directly applied to badminton (analyzing the whipping motion of a smash), table tennis (analyzing the fine wrist movements of the ball on the table), golf (analyzing the swing plane and body rotation), baseball (analyzing the swing motion), etc.

[0033] In this embodiment of the invention, the motion execution and acquisition subsystem 20 includes key equipment for non-invasively capturing all information about the actions performed by the trainee (executor) 21, specifically including a visual acquisition unit 22. The visual acquisition unit 22 includes at least one 3D depth camera and at least one high-speed camera, used to capture image sequence data when the trainee performs the selected sub-technique action, and transmits the acquired image sequence data to the central processing and analysis unit in real time. Specifically, the visual acquisition unit 22 may consist of at least one 3D depth camera and one high-speed camera installed at key locations in the system structure (such as the top or side), capturing the complete movements of the trainee 21 in a non-contact manner and transmitting the acquired image data to the central processing and analysis unit 30 in real time.

[0034] In this invention, the motion execution and acquisition subsystem 20 replaces traditional subjective observation and invasive wearable sensors with high-precision, multi-dimensional visual acquisition, providing high-quality raw data input for subsequent objective and quantitative analysis.

[0035] In this embodiment of the invention, the central processing and analysis unit 30 is the "brain" of the entire system, responsible for real-time analysis and diagnosis of the collected data, and issuing control commands to other subsystems. Specifically, real-time data analysis and diagnosis can be achieved through artificial intelligence (AI) algorithms. The central processing and analysis unit 30 is a high-performance computing unit that receives data from the visual acquisition unit 22 and sends commands and data to the guidance and ball supply subsystem 10 and the human-computer interaction and feedback unit 40. It mainly includes: The AI ​​Biomechanical Analysis Engine 31 is the core algorithm module that can build a real-time 3D skeletal model of the trainee and perform professional biomechanical analysis.

[0036] The Expert Movement Database 32 stores a large number of standard movement models and key parameter ranges of professional athletes, serving as a benchmark for analysis and comparison.

[0037] In this invention, the central processing and analysis unit 30 completely overcomes the problem of traditional guidance lacking objective quantitative standards. It can transform abstract action "feelings" into precise data, such as ball throwing stability and kinetic chain timing, and compare them with expert models to achieve scientific and accurate problem diagnosis.

[0038] Furthermore, the central processing and analysis unit 30 is used to acquire the trainee's key physiological data, automatically calibrate the laser guidance parameters of the guidance device based on the key physiological data; decouple the tennis serve action into multiple sub-technical training items, and display each sub-technical training item for the trainee to choose from based on a preset human-computer interaction interface; extract target control parameters matching the sub-technical training item selected by the trainee from the laser guidance parameters, control the guidance device to project laser guidance according to the target control parameters, so as to guide the trainee to perform the corresponding sub-technical action; acquire image sequence data of the trainee performing the sub-technical action, construct a human motion dynamic model based on the image sequence data, and determine the trainee's key action parameters; after the trainee completes the sub-technical action, generate a motion video with superimposed rendering of the human motion dynamic model and the key action parameters for displaying the analysis results. The process involves constructing a human motion dynamic model based on the image sequence data and determining the trainee's key movement parameters. Specifically, this includes: identifying key joints in each image frame of the image sequence data for three-dimensional skeletal point extraction; constructing a human motion dynamic model based on the temporal changes of each three-dimensional skeletal point; calculating the hip-shoulder separation angle formed by hip and shoulder rotation during the movement cycle based on the extraction results of each three-dimensional skeletal point; determining whether the trunk twisting amplitude and body power transmission efficiency meet the standards based on the hip-shoulder separation angle; calculating the angle changes and angular velocity temporal sequences of each three-dimensional skeletal point during the movement process; and determining the angle changes and angular velocity temporal sequences based on the angular velocity temporal sequences. The system analyzes the sequence of force application and the efficiency of force transmission from the lower limbs, trunk to upper limbs, and determines whether the sequence of force application is correct based on the sequence of force application and the efficiency of force transmission. It identifies the tennis ball in each image frame of the image sequence data, fits the ball's three-dimensional trajectory based on the continuous motion coordinates of the ball, compares the three-dimensional trajectory with a preset standard trajectory to obtain the trajectory spatial offset and the spatial error of the highest point of the ball throw. It identifies the three-dimensional spatial coordinates of the tennis ball and the hitting tool at the moment of impact, locates the actual hitting position, and compares the actual hitting position with the standard ideal hitting point to calculate the hitting point position deviation.

[0039] Furthermore, the central processing and analysis unit 30 is also used to compare the human motion dynamic model with the preset expert motion model, combine the motion comparison results and the key motion parameters to diagnose problems in the trainee's sub-technical motion training, and highlight and / or broadcast the diagnosis results and improvement suggestions.

[0040] Furthermore, the central processing and analysis unit 30 is also used to generate a comprehensive report containing a capability radar chart, a progress curve of key movement parameters and a summary of problems after completing a set of training projects selected by trainees for sub-technical movements, and to automatically recommend the training mode for the next stage.

[0041] In this embodiment of the invention, the human-computer interaction and feedback unit 40, as the core interface of human-computer interaction, can provide different training modes for trainees to choose from, and can also present complex analysis results to trainees in an intuitive and easy-to-understand manner in real time. It mainly includes: The visualization display module 41, implemented using a touch screen, is used to display multiple sub-technical action training items decoupled from the tennis serve action, and to monitor the trainee's touch selection commands; as well as to display motion videos sent by the central processing and analysis unit, which are overlaid with and rendered with the human motion dynamic model and the key action parameters. The voice broadcast module 42 is used to broadcast the diagnostic results and improvement suggestions via voice.

[0042] This invention solves the problem of feedback delay in the prior art. Within seconds after the action is completed, it allows trainees to immediately "see" and "understand" their problems through video playback with superimposed data, on-screen comparison, and highlighting of problem areas, greatly shortening the feedback loop of learning and correction.

[0043] The computer vision-based intelligent tennis serve training system provided in this invention forms a real-time closed loop of "guidance-execution-capture-analysis-feedback". The trainee performs actions with the assistance of the guidance and ball supply subsystem 10, the motion execution and acquisition subsystem 20 captures information, the central processing and analysis unit 30 performs intelligent diagnosis, and finally, the human-computer interaction and feedback unit 40 provides immediate guidance. This achieves personalized guidance training for the trainee and enables non-invasive, whole-body motion data acquisition of the trainee's complete technical movements, ensuring more comprehensive and accurate training analysis and achieving efficient skill improvement.

[0044] Figure 3 A flowchart illustrating an intelligent tennis serve training method based on computer vision, provided by an embodiment of the present invention, is shown. Figure 3 As shown, the intelligent tennis serve training method based on computer vision proposed in this invention includes the following steps: S1. Acquire key physiological data of the trainee and automatically calibrate the laser guidance parameters of the guidance device based on the key physiological data; S2. Decouple the tennis serve motion into multiple sub-technical motion training items, and display each sub-technical motion training item based on a preset human-computer interaction interface for trainees to select. S3. Extract target control parameters from the laser guidance parameters that match the training item of the sub-technical movement selected by the trainee, and control the guidance device to project laser guidance according to the target control parameters so as to guide the trainee to perform the corresponding sub-technical movement in accordance with the guidance. S4. Obtain image sequence data of the trainee performing sub-technique actions, construct a human motion dynamic model based on the image sequence data, and determine the trainee's key action parameters; S5. After the trainee completes the sub-technical action, a motion video with superimposed rendering of the human motion dynamic model and the key action parameters is generated for display of the analysis results.

[0045] In this embodiment of the invention, step S1, acquiring the trainee's key physiological data and automatically calibrating the laser guidance parameters of the guidance device based on the key physiological data, specifically includes: when the system is started, and the trainee stands in the designated position and performs a specific action, the visual acquisition unit automatically measures the trainee's height, arm span, and other key physiological data, and creates or loads their personal profile, automatically calibrating the laser guidance parameters of the guidance device based on the individual's key physiological data. This invention achieves truly personalized guidance. Based on the user's physiological data, the system automatically calibrates subsequent laser guidance parameters (such as the height of the "virtual hitting window"), ensuring the scientific validity and effectiveness of the guidance target.

[0046] In this embodiment of the invention, step S2 decouples the tennis serve motion into multiple sub-technical training items. These sub-technical training items are displayed on a preset human-computer interaction interface for the trainee to choose from. Specifically, the trainee selects the focus of the training session through the interface, such as "ball toss stability training," "striking point perception training," or "comprehensive complete motion training." When "striking point perception training" is selected, the system automatically diverts a ball from the main ball supply module to the release buffer, preparing for the next automatic ball drop. This invention solves the problem of inefficient training caused by multi-variable coupling. By decomposing the complex serve motion into independent, controllable training modules, trainees can tackle technical difficulties one by one. For example, they can first ensure ball toss stability, then practice swing timing, and finally integrate the training. The learning path is clear, avoiding blind practice.

[0047] In this embodiment of the invention, step S3 involves extracting target control parameters from the laser guidance parameters that match the sub-technical movement training item selected by the trainee, and controlling the guidance device to project laser guidance according to the target control parameters. This includes: projecting laser guidance (such as a fixed hitting window) according to the selection in S2, and providing a training ball so that the trainee can follow the guidance and perform the corresponding technical movement. This invention creates a stable and repeatable frame of reference for each practice session, enabling the trainee to focus on improving the specific movement itself, rather than dealing with unstable external conditions.

[0048] In this embodiment of the invention, step S4, which involves constructing a human motion dynamic model based on the image sequence data and determining the trainee's key movement parameters, includes: Identify key human joints in each image frame of the image sequence data to extract three-dimensional skeletal points, and construct a human motion dynamic model based on the temporal changes of each three-dimensional skeletal point. Based on the extraction results of each three-dimensional skeletal point, the hip-shoulder separation angle formed by hip rotation and shoulder rotation during the movement cycle is calculated, and the trunk twisting amplitude and the body's power transmission efficiency are determined according to the hip-shoulder separation angle. Calculate the angle changes and angular velocity timing of each three-dimensional skeletal point during the movement process. Analyze the order of force exertion and force transmission efficiency of each joint from the lower limbs and trunk to the upper limbs based on the angle changes and angular velocity timing. Determine whether the order of force exertion of each limb is correct based on the order of force exertion and force transmission efficiency of each joint.

[0049] Furthermore, constructing a human motion dynamic model based on the image sequence data and determining the trainee's key movement parameters also includes: Identify the tennis ball in each image frame of the image sequence data, fit the three-dimensional motion trajectory of the ball based on the continuous motion coordinates of the tennis ball, compare the three-dimensional motion trajectory of the ball with the preset standard trajectory, and obtain the trajectory spatial offset and the spatial error of the highest point of the ball throw. The system identifies the three-dimensional spatial coordinates of the tennis ball and the hitting tool at the moment of impact, locates the actual hitting position, and compares the actual hitting position with the standard ideal hitting point to calculate the deviation of the hitting point position.

[0050] This invention captures image sequences via a visual acquisition unit while the trainee performs actions. Simultaneously, the processing unit's AI engine performs professional biomechanical analysis, transforming abstract movements into precise and objective data to provide a basis for scientific diagnosis. Key movement parameters include: The three-dimensional spatial error at the highest point of the ball toss: used to quantify the stability of the ball toss, measured in centimeters. This is crucial for evaluating the fundamentals of the serve.

[0051] Impact point position deviation: Used to quantify the positional deviation between the actual impact point position and the standard ideal impact point position.

[0052] Peak angular velocity timing of the kinetic chain: used to determine if the force application sequence is correct. Ideally, the peak angular velocities occur sequentially from the hip, torso, and shoulder. A disordered timing sequence directly indicates a problem with disjointed force application.

[0053] Hip-shoulder separation angle: The maximum angular difference between hip rotation and shoulder rotation, measured in degrees. This parameter is a core indicator for measuring whether the body rotation is sufficient and whether it can effectively store power.

[0054] In this embodiment of the invention, the method further includes: comparing the human motion dynamic model with a preset expert motion model, diagnosing problems in the trainee's sub-technique motion training based on the motion comparison results and the key motion parameters, and highlighting and / or broadcasting the diagnostic results and improvement suggestions. Specifically, within 1-2 seconds after the training motion is completed, the display unit automatically plays a playback video with superimposed skeletal lines and key parameter data, and can compare it with the built-in expert model motion on the same screen. Problems diagnosed by AI are highlighted, and the voice unit broadcasts the core improvement suggestions. This invention can provide immediate, multi-dimensional, and easy-to-understand feedback, solving the problems of feedback delay and subjectivity in traditional training, allowing trainees to immediately know "where they went wrong," "where they are inferior to experts," and "how to improve."

[0055] In this embodiment of the invention, the method further includes: after completing a set of training exercises for sub-technical movements selected by the trainee, generating a comprehensive report containing an ability radar chart, progress curves for key movement parameters, and a summary of problems, and automatically recommending a training mode for the next stage. Specifically, after a set of training exercises is completed, the system generates a comprehensive report containing an ability radar chart, progress curves, and a summary of core problems based on accumulated data, and automatically recommends a training mode for the next stage. This invention provides long-term growth tracking and intelligent training planning, making training no longer fragmented, but a continuously improving scientific system.

[0056] The intelligent tennis serve training method based on computer vision provided in this invention has the following significant advantages: First, it achieves truly real-time, closed-loop training feedback, which greatly improves learning and error correction efficiency.

[0057] Existing methods of manual guidance and video playback analysis suffer from significant feedback delays. This invention, through a complete closed loop of "guidance-execution-capture-analysis-feedback," reduces the time from completing a movement to receiving professional diagnostic feedback to within 1-2 seconds. Trainees can immediately understand the strengths and weaknesses of their movements through visual playback and data during this "golden time" while muscle memory is still intact, and make immediate adjustments in the next practice session. This instantaneous and continuous feedback loop breaks the inefficient traditional model of separating "training" and "analysis," significantly accelerating the formation of correct movements.

[0058] Second, transform subjective experience-based guidance into objective, quantitative scientific analysis, so that training is based on evidence.

[0059] Traditional coaching guidance is experience-based and qualitative, while the AI ​​biomechanical analysis engine of this invention is scientific and quantitative. It can precisely quantify vague movement sensations, such as "insufficient rotation" or "disconnected force," into specific data such as "hip-shoulder separation angle is only 30 degrees (below the standard value)" or "the peak shoulder angular velocity occurs 0.1 seconds before the hip." This data-driven diagnosis provides trainees with irrefutable and clear evidence for improvement, making their training no longer aimless but like a scientist conducting an experiment—with a clear goal and a well-defined path.

[0060] Third, through innovative methods of "decoupling training" and "variable control," the core pain point of inefficient repetitive practice is fundamentally solved.

[0061] The biggest challenge with existing technologies lies in the inability to control key variables in the serving motion—especially the stability of the ball toss. This invention perfectly solves this problem through the following design: • Action Decoupling: By setting up multiple modes such as "Ball Toss Stability Training" and "Striking Point Perception Training," the complex serving action is broken down into independent, controllable training units. Beginners can focus on overcoming difficulties one by one, avoiding the confusion caused by errors in multiple aspects simultaneously.

[0062] • Constant Variables: Utilizing a "virtual hitting window" projected by a laser projection unit, a constant and clear spatial target is provided for each swing and hit by the trainee. In the "hitting point perception training" mode, the "perfect ball toss" provided by the tennis ball release device further eliminates interference from unstable ball tosses. This ensures the stability and repeatability of training conditions, allowing trainees to truly focus on "deliberate practice" of specific movements, thus breaking the vicious cycle of needing hundreds or thousands of blind repetitions due to unstable ball tosses.

[0063] Fourth, it enables non-invasive, whole-body motion data collection, resulting in more comprehensive and accurate analysis.

[0064] Compared to invasive sensors that need to be worn on the body or integrated into the racket, this invention uses a purely visual acquisition method, causing zero interference to the trainee and ensuring the naturalness and authenticity of the technical movements. More importantly, the visual method can capture motion information from all joints in the body, thereby performing a holistic analysis of the complete "kinetic chain" from leg push-off to core transmission and arm swing—something that no local sensor can achieve. This global analysis can uncover the root cause of the trainee's power generation problems.

[0065] Fifth, it combines the analytical precision of scientific research with the affordability of the mass market, greatly increasing the accessibility of professional training.

[0066] This invention cleverly utilizes relatively inexpensive cameras and AI software algorithms to achieve core analytical functions comparable to expensive optical motion capture labs (such as the Vicon system). Its modular hardware design (especially the portable L-shaped structure) makes the manufacturing cost and price of the entire system affordable for tennis enthusiasts, families, and ordinary clubs. This truly democratizes top-tier sports science analysis technology, allowing ordinary people to enjoy scientific training guidance previously only accessible to elite professionals.

[0067] Sixth, it provides an intuitive, multimodal, immersive learning experience, making the learning process more efficient and fun.

[0068] This invention goes beyond simply providing cold, impersonal data. It creates a multi-sensory, immersive learning environment by combining visualization units (such as action playback, data overlay, and simultaneous screen comparison), laser projection units (physical spatial motion path guidance), and voice broadcast units (real-time language prompts). This "virtual-real" guidance method is more intuitive and easier for learners to understand and imitate than simple verbal instruction or video playback, enhancing the fun and engagement of training.

[0069] This invention successfully solves many problems in the field of tennis serve training in the existing technology, and provides trainees with a training solution that integrates professionalism, real-time performance, scientific rigor, economy, and fun.

[0070] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0071] The computer vision-based intelligent tennis serve training method and system provided in this invention can automatically calibrate the laser guidance parameters of the guidance device according to the trainee's key physiological data to achieve personalized guidance training for the trainee. By "decoupling" complex movements and controlling and guiding key variables such as ball toss position and hitting point, it breaks the vicious cycle of inefficient repetition and achieves truly targeted and efficient key movement practice. It can also achieve non-invasive, whole-body motion data collection of the trainee's complete technical movements, ensuring more comprehensive and accurate training analysis and achieving efficient skill improvement.

[0072] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, any of the claimed embodiments can be used in any combination.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A computer vision-based intelligent tennis serve training method, characterized in that, The method includes: Acquire key physiological data of the trainee, and automatically calibrate the laser guidance parameters of the guidance device based on the key physiological data; The tennis serve motion is decoupled into multiple sub-technical training items, and each sub-technical training item is displayed based on a preset human-computer interaction interface for trainees to choose from. The target control parameters that match the sub-technical movement training item selected by the trainee are extracted from the laser guidance parameters. The control device projects laser guidance according to the target control parameters to guide the trainee to perform the corresponding sub-technical movement in accordance with the guidance. Acquire image sequence data of the trainee performing sub-technique actions, construct a human motion dynamic model based on the image sequence data, and determine the trainee's key action parameters; After the trainee completes the sub-technical action, a motion video with a superimposed rendering of the human motion dynamic model and the key action parameters is generated for the analysis results to be displayed.

2. The method according to claim 1, characterized in that, The method further includes: The human motion dynamic model is compared with a preset expert motion model. Based on the motion comparison results and the key motion parameters, problems in the trainee's sub-technical motion training are diagnosed, and the diagnosis results and improvement suggestions are highlighted and / or broadcast by voice.

3. The method according to claim 1, characterized in that, A human motion dynamic model is constructed based on the image sequence data, and key movement parameters of the trainee are determined, including: Identify key human joints in each image frame of the image sequence data to extract three-dimensional skeletal points, and construct a human motion dynamic model based on the temporal changes of each three-dimensional skeletal point. Based on the extraction results of each three-dimensional skeletal point, the hip-shoulder separation angle formed by hip rotation and shoulder rotation during the movement cycle is calculated, and the trunk twisting amplitude and the body's power transmission efficiency are determined according to the hip-shoulder separation angle. Calculate the angle changes and angular velocity timing of each three-dimensional skeletal point during the movement process. Analyze the order of force exertion and force transmission efficiency of each joint from the lower limbs and trunk to the upper limbs based on the angle changes and angular velocity timing. Determine whether the order of force exertion of each limb is correct based on the order of force exertion and force transmission efficiency of each joint.

4. The method according to claim 3, characterized in that, Constructing a human motion dynamic model based on the image sequence data and determining the trainee's key movement parameters also includes: Identify the tennis ball in each image frame of the image sequence data, fit the three-dimensional motion trajectory of the ball based on the continuous motion coordinates of the tennis ball, compare the three-dimensional motion trajectory of the ball with the preset standard trajectory, and obtain the trajectory spatial offset and the spatial error of the highest point of the ball throw. The system identifies the three-dimensional spatial coordinates of the tennis ball and the hitting tool at the moment of impact, locates the actual hitting position, and compares the actual hitting position with the standard ideal hitting point to calculate the deviation of the hitting point position.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: After completing a set of training exercises for the selected sub-technical movements, a comprehensive report is generated, including a capability radar chart, progress curves for key movement parameters, and a summary of problems. The system also automatically recommends the training mode for the next stage.

6. A computer vision-based intelligent tennis serve training system, characterized in that, The system includes a guidance and ball supply subsystem, a motion execution and acquisition subsystem, a human-computer interaction and feedback unit, and a central processing and analysis unit. The ball guidance and feeding subsystem is used to project laser guidance marks into the training area according to the control instructions of the central processing and analysis unit, and to store and feed tennis balls. The human-computer interaction and feedback unit is used to display multiple sub-technical movement training items decoupled from the tennis serve motion for trainees to choose from. The motion execution and acquisition subsystem is used to capture image sequence data of the trainee performing the selected sub-technique in a non-contact manner, and transmit the acquired image sequence data to the central processing and analysis unit in real time. A central processing and analysis unit is used to execute the computer vision-based intelligent tennis serve training method as described in any one of claims 1-4; The human-computer interaction and feedback unit is also used to receive and display motion videos sent by the central processing and analysis unit, which are overlaid with a human motion dynamic model and the key motion parameters.

7. The intelligent tennis serve training system based on computer vision according to claim 6, characterized in that, The central processing and analysis unit is also used to compare the human motion dynamic model with the preset expert motion model, combine the motion comparison results and the key motion parameters to diagnose problems in the trainee's sub-technical motion training, and highlight and / or broadcast the diagnosis results and improvement suggestions.

8. The intelligent tennis serve training system based on computer vision according to claim 6, characterized in that, The guiding and ball supply subsystem includes: The laser projection module includes a miniature laser projector for projecting a virtual hitting window and / or a position guide cursor onto the training area according to control instructions from the central processing and analysis unit. The storage and gravity ball supply module includes a ball storage compartment and a ball delivery slide, which is used to transport tennis balls in the ball storage compartment to the release buffer through the ball delivery slide using the principle of gravity; A buffer to be released, used to load and position tennis balls for free-fall training from the ball release chute; A tennis ball release device is used to release a tennis ball from a release buffer according to control instructions from a central processing and analysis unit.

9. The intelligent tennis serve training system based on computer vision according to claim 6, characterized in that, The motion execution and acquisition subsystem includes at least one 3D depth camera and at least one high-speed camera, used to capture image sequence data when the trainee performs the selected sub-technique action, and transmit the acquired image sequence data to the central processing and analysis unit in real time.

10. The intelligent tennis serve training system based on computer vision according to claim 6, characterized in that, The human-computer interaction and feedback unit includes: The visualization module, implemented using a touch screen, is used to display multiple sub-technical movement training items decoupled from the tennis serve motion, and to monitor the trainee's touch selection commands; as well as to display motion videos sent by the central processing and analysis unit, which are overlaid with and rendered with the human motion dynamic model and the key motion parameters. The voice broadcast module is used to broadcast diagnostic results and improvement suggestions via voice.