Intelligent swimming pool surfer training guidance system and method based on visual interaction

The smart pool surfer, which uses visual interaction technology, captures users' gestures and movements in real time, providing contactless operation and professional guidance. This solves the problems of inconvenient operation and boring training of existing equipment, and improves the user's training experience and safety.

CN122441071APending Publication Date: 2026-07-24YITUO ELECTRIC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YITUO ELECTRIC CO LTD
Filing Date
2026-03-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing pool surfing machines are cumbersome and unsafe to operate, lack the ability to capture and analyze users' surfing postures, make it difficult for users to obtain scientifically quantifiable exercise data, and make the training process tedious and lacking in interactivity.

Method used

The intelligent pool surfing device adopts a vision-interaction-based system. It captures user gestures and full-body motion images in real time through a video acquisition module. The main control processing unit performs gesture recognition and skeletal point extraction to generate control commands and training guidance information. Combined with preset teaching videos and a cloud-based competition system, it enables contactless operation and real-time feedback.

Benefits of technology

It enables convenient and safe water flow control, provides professional training guidance, enhances the fun and interactivity of training, avoids misoperation and sports injuries, and provides scientific training assessment and data feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122441071A_ABST
    Figure CN122441071A_ABST
Patent Text Reader

Abstract

The application relates to a visual interaction-based intelligent swimming pool surfer training guidance system and method, which comprises a swimming pool surfer main body, a video image acquisition module and a main control processing unit. The swimming pool surfer main body is internally provided with a water flow driving module, which is used for receiving control instructions and adjusting water flow speed and jetting direction. The video image acquisition module is used for collecting gesture images and whole body action images of a user in water. The main control processing unit is used for gesture recognition based on the gesture images and generation of control instructions, and bone point extraction and action evaluation based on the whole body action images and generation of real-time training guidance information. The application realizes non-contact gesture control through visual interaction, avoids the safety hidden danger of wet hand operation, realizes private coach level action correction through bone recognition and quantitative evaluation, prevents sports injury, and improves training interest and user enthusiasm in combination with AR follow-up training and cloud competition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent water sports equipment, and in particular to an intelligent pool surfer training and guidance system and method based on visual interaction. Background Technology

[0002] With the widespread adoption of the national fitness concept and the popularization of surfing, pool surfing machines that can simulate surfing training in regular swimming pools have been widely used in various scenarios such as commercial fitness venues and private swimming pools. The market has placed higher demands on the operational safety, training professionalism, and user experience of these devices.

[0003] Current mainstream pool surfing machines typically only offer basic functions like adjusting water flow speed and angle, requiring users to operate them via physical buttons on the device or a remote control. In practice, users must locate and press buttons in water or with wet hands, or use the remote control to switch speeds. This method is not only cumbersome but also prone to misoperation due to slippery hands, and even poses safety hazards due to contact with electrical components. Furthermore, existing devices lack the ability to capture and analyze the user's surfing posture; users can only judge training effectiveness based on subjective feelings, making it difficult to obtain scientifically quantifiable exercise data.

[0004] Furthermore, existing surfing machines cannot provide guidance on surfing posture, making it difficult for users to self-assess their training effectiveness. The training process is also tedious, lacking interactivity and competitiveness, and fails to meet users' comprehensive needs for professional surfing training, safe and convenient operation, and a fun sporting experience. Summary of the Invention

[0005] In view of this, it is necessary to provide a visual interaction-based intelligent pool surfer training guidance system and method to solve the above-mentioned problems of the prior art.

[0006] To address the aforementioned problems, in a first aspect, embodiments of the present invention provide a visual interaction-based intelligent pool surfer training guidance system, comprising:

[0007] The main body of the pool surfer is installed on the side wall of the pool and has a built-in water flow drive module. The water flow drive module is used to receive control commands and adjust the water flow speed and jet direction.

[0008] The video acquisition module is used to capture images of the user's hand gestures and full-body movements in the water;

[0009] The main control processing unit is connected to the video acquisition module and the water flow driving module respectively. It is used to perform gesture recognition based on the gesture image and generate control commands based on the recognition results; and to extract skeletal points and evaluate movements based on the whole-body motion image to generate real-time training guidance information.

[0010] Secondly, embodiments of the present invention provide a training guidance method for an intelligent pool surfer based on visual interaction, comprising:

[0011] The system retrieves the training course selected by the user, calls up the corresponding preset teaching video or standard action video, and displays it to the user.

[0012] The system captures real-time images of the user's hand gestures and full-body movements in the water and transmits the captured image data to the main control processing unit.

[0013] Based on the gesture image, gesture recognition is performed, and control commands are generated according to the recognition results and sent to the water flow drive module to adjust the water flow speed and spray direction;

[0014] Based on the full-body motion images, skeletal points are extracted and motion is evaluated, training guidance information is generated in real time, and feedback is provided to the user.

[0015] After training, the system summarizes the user's training data throughout the entire period, and generates and outputs a complete training report that includes training duration, calorie consumption, and posture score.

[0016] Furthermore, gesture recognition based on the gesture image includes:

[0017] Based on the human detection algorithm, the user's head and shoulder positions are located. A region of interest for gesture detection is generated and updated in real time with the center of the shoulder as a reference. Gesture recognition is performed on the image data within the region of interest.

[0018] A dynamic background model is established, and a dynamic water ripple background is removed by a background removal algorithm based on semantic segmentation. The background difference method is combined to distinguish the foreground target from the water background, and the high-brightness reflective area is separated by the HSV color space to extract the hand foreground image.

[0019] The gesture features in consecutive frames of images are matched and verified. When the duration of the same gesture feature exceeds a preset time threshold, it is confirmed as a valid gesture. Control commands are generated based on the spatial displacement vector of the valid gesture.

[0020] Skeletal point extraction and motion assessment based on the full-body motion images include:

[0021] Collect static standing posture image data for a preset duration from the user, extract the skeletal angles of the neutral standing posture, and generate personalized motion evaluation benchmark data for the user;

[0022] High-precision skeletal key point detection is performed on the user's upper body, and Kalman filtering is applied to the coordinates of the detected skeletal points. When the confidence of the key points of the lower body is lower than the preset confidence threshold, the centroid motion trend prediction mode is switched to predict the coordinates of the current joint points based on the motion trajectory before occlusion, and continuous human skeletal model data is output.

[0023] Based on human skeletal model data and motion assessment benchmark data, quantitative assessment parameters are calculated, including hip joint stability score, posture accuracy rate, movement frequency, and calorie consumption value.

[0024] The quantitative evaluation parameters are compared with preset standard thresholds. When the quantitative evaluation parameters deviate from the preset standard thresholds, corresponding real-time training guidance information is generated.

[0025] The intelligent pool surfer training guidance system and method based on visual interaction provided by this invention has the following advantages compared with the prior art:

[0026] (1) The present invention adopts a non-contact gesture control scheme based on visual interaction to replace the traditional physical buttons and remote control operation, avoids the safety hazards such as accidental touch and electric shock caused by wet hands operation, greatly simplifies the surfer operation process, realizes convenient and safe control of the water flow parameters of the equipment, and solves the problems of inconvenient operation and insufficient safety of existing equipment.

[0027] (2) This invention constructs a human motion model through skeletal recognition technology, and combines personalized action benchmarks and quantitative evaluation systems to capture users’ surfing movements in real time and generate professional guidance information, digitize professional surfing training experience, realize private coach-level movement correction, effectively avoid sports injuries caused by incorrect movements, and help users scientifically improve training results.

[0028] (3) This invention breaks the monotonous mode of traditional surfing training by using preset teaching videos, AR standard action superimposed training mode, and cloud-based competitive interaction system, transforming professional physical training into an immersive and interactive sports experience, effectively improving users' training enthusiasm. Attached Figure Description

[0029] Figure 1 The architecture diagram of the intelligent pool surfer training guidance system based on visual interaction provided by the present invention;

[0030] Figure 2 This is a structural block diagram of the non-contact control system based on gesture recognition provided by the present invention;

[0031] Figure 3 This is a structural block diagram of the motion training and evaluation system based on skeleton recognition provided by the present invention;

[0032] Figure 4 The flowchart of the intelligent pool surfer training guidance method based on visual interaction provided by the present invention. Detailed Implementation

[0033] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0034] The naming or numbering of steps in the embodiments of the present invention does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved.

[0035] Figure 1 This is an architecture diagram of the visual interaction-based intelligent pool surfer training guidance system provided by the present invention, with reference to... Figure 1 The intelligent pool surfer training guidance system based on visual interaction provided by this invention includes:

[0036] The main body of the pool surfer is installed on the side wall of the pool and has a built-in water flow drive module. The water flow drive module is used to receive control commands and adjust the water flow speed and spray direction. The water flow drive module includes a variable speed motor, an adjustable-angle nozzle, and a stepper motor that drives the nozzle to deflect. The variable speed motor and the stepper motor are respectively connected to the main control processing unit. The control commands of the main control processing unit include adjusting the PWM duty cycle of the variable speed motor to change the water flow speed and adjusting the rotation angle of the stepper motor to change the spray direction of the nozzle.

[0037] The video acquisition module is used to capture images of the user's hand gestures and full-body movements in the water;

[0038] The main control processing unit is connected to the video acquisition module and the water flow driving module respectively. It is used to perform gesture recognition based on the gesture image and generate control commands based on the recognition results; and to extract skeletal points and evaluate movements based on the whole-body motion image to generate real-time training guidance information.

[0039] Specifically, the video acquisition module refers to a waterproof camera equipped with a wide-angle lens and infrared illumination, which can be built into the top of the surfer or float on the water. The water flow drive module is integrated inside the surfer body, including a variable-speed motor, an adjustable-angle nozzle, and a stepper motor that drives the nozzle deflection. The main control processing unit changes the water flow speed by adjusting the PWM (Pulse Width Modulation) duty cycle of the variable-speed motor, and changes the nozzle spray direction by adjusting the rotation angle of the stepper motor, thereby simulating surfing environments of different intensities.

[0040] The main control processing unit is connected to both the image acquisition module and the water flow drive module. It incorporates a Neural-network Processing Unit (NPU), also known as an AI neural network processor. An NPU is a hardware accelerator designed for the efficient execution of artificial intelligence algorithms. This invention integrates an NPU into the main control processing unit because gesture recognition and skeletal point extraction involve a large number of convolutional neural network operations. Traditional general-purpose processors struggle to achieve millisecond-level real-time processing under low power conditions. The NPU, however, can perform image recognition, skeletal point extraction, and data processing tasks in parallel with extremely high energy efficiency, ensuring a real-time interactive experience with simultaneous feedback and practice during user browsing.

[0041] In practical implementation, upon power-on, the video acquisition module automatically calibrates and enters standby mode. After the user selects a training course, the system retrieves the corresponding preset instructional video or standard motion footage and displays it through the user's interactive terminal. The video acquisition module captures real-time images of the user's gestures and full-body movements in the water and transmits them to the main control processing unit.

[0042] The main control processing unit performs gesture recognition based on gesture images, generates control commands based on the recognition results, and sends them to the water flow drive module. This adjusts the PWM duty cycle of the variable-speed motor to change the water flow speed, or adjusts the rotation angle of the stepper motor to change the nozzle spray direction, achieving non-contact water flow control. Simultaneously, the main control processing unit extracts skeletal points and evaluates movements based on full-body motion images, generating real-time training guidance information. This guidance is displayed on a screen with a semi-transparent skeletal image of the standard movements overlaid, or through a voice prompt component, providing real-time error correction and feedback during training. After training, the system summarizes all training data and generates a complete training report including training duration, calorie consumption, and posture scores.

[0043] This invention replaces traditional physical operation with visual interaction technology, avoiding the safety hazards and accidental touches associated with wet hands, and achieving convenient and safe control of the surfer. Leveraging the real-time computing power of its built-in AI neural network processor, it can capture and scientifically evaluate the user's surfing movements in real time, providing professional training guidance and effectively preventing sports injuries caused by incorrect movements. Simultaneously, through an immersive visual interaction mode, it breaks away from the monotony of traditional surfing training, enhancing training fun and user engagement, and meeting the diverse surfing training needs of users.

[0044] In some embodiments of this application, the visual interaction-based intelligent pool surfer training guidance system further includes a user interaction terminal, which includes a display screen and a voice broadcast component.

[0045] The display screen is used to display the user's image captured by the video acquisition module in real time, and to overlay a semi-transparent skeletal image of the standard movement on the user's image. At the same time, it displays real-time training guidance information, training score data and online competition ranking data. The voice broadcast component is used to broadcast the real-time training guidance information in voice form.

[0046] Specifically, the user interaction terminal refers to an external device that is connected to the main control processing unit via signals. User interaction terminals include smart terminals such as mobile phones, computers, and tablets, all of which are equipped with displays and voice broadcast components. The user interaction terminal and the main control processing unit are paired for bidirectional communication, ensuring that various types of information generated by the main control processing unit can be transmitted to the user interaction terminal in real time, and that user operation commands can be synchronously transmitted back to the main control processing unit.

[0047] During training, the video acquisition module captures real-time images of the user's full-body movements and gestures in the water, transmitting them synchronously to the main control processing unit. The main control processing unit then transmits the captured real-time user footage to the user's interactive terminal for display. Simultaneously, the main control processing unit's built-in video follow-up and cloud-based competition system utilizes semi-transparent skeletal images of standard movements from the corresponding course. Through AR (Augmented Reality) fusion technology, these semi-transparent skeletal images are overlaid onto the user's real-time footage displayed on the screen. Users can then intuitively compare their movements to the standard movements on the screen, enabling immersive follow-along training.

[0048] The voice broadcast component is connected to the main control processing unit. When the motion training and evaluation system based on skeletal recognition detects that the user's action deviates from the preset standard threshold, the real-time training guidance information generated by the system will be broadcast in voice form in an instant, ensuring that the user can obtain professional guidance without being distracted by looking at the screen in the water environment.

[0049] In some embodiments of this application, reference is made to Figure 1The main control processing unit includes:

[0050] The non-contact control system based on gesture recognition is used to receive the gesture image, generate corresponding control commands through gesture recognition processing, and send the control commands to the water flow drive module;

[0051] A motion training and evaluation system based on skeleton recognition is used to receive the full-body motion images, extract skeletal points to construct a human skeleton model, calculate motion parameters, and generate real-time training guidance information.

[0052] The video training and cloud-based competition system outputs preset instructional videos or standard movement footage, compares the user's real-time movements with the standard movements to generate training scores, and uploads the user's training data to the cloud to participate in online competition rankings.

[0053] Specifically, the main control processing unit integrates three core systems, respectively realizing contactless gesture control, real-time motion evaluation, and immersive interactive training functions. Among them, the gesture recognition-based contactless control system aims to solve the inconvenience and safety hazards caused by users operating physical buttons or remote controls with wet hands, achieving the technical goal of precise control of the surfer without contact. The gesture recognition-based contactless control system has a preset gesture command set, which stores the mapping relationship between various preset gestures and device control commands. For example, if a user makes a fist gesture and holds it for more than a preset time threshold, the system recognizes it as a start or acceleration command; an open palm with fingers pointing upwards is recognized as a pause or stop command; and making an OK gesture is recognized as a command to reduce the water flow. For adjusting the water jet direction and wave pattern, the system recognizes the dynamic trajectory of the hand: an upward wave of the hand corresponds to the nozzle lifting upwards, and a downward wave corresponds to the nozzle pressing downwards; two consecutive upward waves increase wave irregularity, and two consecutive downward waves decrease wave irregularity.

[0054] In the specific control logic, the video acquisition module captures the user's video stream in the water in real time and transmits it to the main control processing unit. The main control processing unit first extracts the coordinates of key hand points, and then uses a lightweight convolutional neural network model, such as MobileNet, to recognize and classify the gesture category. MobileNet is a lightweight convolutional neural network architecture designed specifically for mobile and embedded vision applications, enabling efficient image classification and feature extraction with limited computing resources. The recognized gesture category is mapped to a corresponding control signal, which the main control processing unit uses to adjust the PWM duty cycle of the variable-speed motor in the water flow drive module to change the water flow speed, or to adjust the rotation angle of the stepper motor to change the nozzle's spray direction.

[0055] The skeletal recognition-based motion training and evaluation system aims to address the problem that existing surfing devices cannot scientifically guide user posture. It digitizes the evaluation experience of professional instructors, enabling real-time motion correction and quantitative evaluation. In this embodiment, the skeletal recognition-based motion training and evaluation system constructs a real-time human skeletal model of the user and uses the OpenPose algorithm for skeletal point extraction. OpenPose is an open-source algorithm for real-time multi-person keypoint detection, capable of identifying major human joints such as the shoulder, elbow, hip, knee, and ankle, and outputting the two-dimensional coordinates and confidence scores of each keypoint. Based on the extracted skeletal point data, the system calculates multiple motion parameters.

[0056] Stability score is obtained by calculating the fluctuation range of the hip center point in the vertical direction, i.e., the Y-axis. The smaller the fluctuation of the hip center point, the more stable the user's core, and the higher the score. Posture accuracy is obtained by comparing the user's real-time joint angles, such as the knee flexion and extension angles, with preset standard thresholds in the standard movement library. When the deviation exceeds 15%, it is judged as a posture error. Movement frequency is obtained by counting the number of limb reciprocating cycles per unit time, such as counting the number of arm strokes or leg push-offs per unit time. During real-time evaluation, when the skeletal recognition-based motion training and evaluation system detects that the user's center of gravity is excessively shifted or the leg flexion angle is insufficient, it immediately provides real-time prompts through the voice broadcast component or the APP interface of the user interaction terminal, such as issuing voice reminders such as "Please lower your center of gravity" or "Leg flexion angle is insufficient," realizing a real-time error correction function with feedback during training.

[0057] The video-based training and cloud-based competition system aims to solve the problem of monotonous training by enhancing its fun and interactivity through AR fusion technology and a cloud-based competition mechanism. The system offers a variety of pre-set surfing instructional videos, including beginner balance lessons and advanced skill lessons, allowing users to choose courses based on their skill level. During training, the system uses AR fusion technology to overlay a semi-transparent skeletal image of the standard movement onto the user's real-time view captured by the video acquisition module. Users can then "look in a mirror" and follow the training, intuitively adjusting their movements to match the standard. After training, the system generates an instant training report based on the degree of overlap between the user's real-time movements and the standard movements, categorizing the results into four levels: S, A, B, and C, for a clear understanding of the training effect. The system also supports a cloud competition mode, allowing users to initiate or join online surfing rooms to compete against other users in real-time. In cloud competition mode, the system issues standardized challenge tasks, such as maintaining balance for 5 minutes in a level 3 current or completing 100 strokes. The system uploads data such as the participants' posture stability score and endurance time to the cloud in real time. The system updates the leaderboard in real time and supports replaying highlights to create a competitive atmosphere.

[0058] In practice, the user turns on the device, and the video acquisition module automatically completes calibration and enters standby mode. The user makes a fist gesture in the water; the non-contact control system, based on gesture recognition, recognizes the gesture and generates a start command, sending it to the water flow drive module to activate the surfer. The user selects a basic balance training course through the user interface terminal. The video practice and cloud-based competition system calls the corresponding preset instructional video, which is simultaneously displayed on the screen as standard movement footage. During the user's practice, the motion training and evaluation system based on skeletal recognition captures the user's movements in real time. When insufficient knee flexion is detected, a prompt such as "Bend your knee a little more" is issued via a voice broadcast component. After training, the video practice and cloud-based competition system summarizes the training data for the entire period and displays the training report through the user interface terminal's APP interface, including the duration of 15 minutes, calories burned (180 kcal), and posture score (85 points).

[0059] In some embodiments of this application, reference is made to Figure 2 The gesture recognition-based contactless control system includes:

[0060] The dynamic region of interest (ROI) locking unit is used to acquire image data collected by the video acquisition module, lock the user's head and shoulder positions according to the human detection algorithm, generate and update the region of interest for gesture detection in real time with the shoulder center as the reference, and output only the image data within the region of interest.

[0061] An anti-interference preprocessing unit, connected to a dynamic region of interest locking unit, is used to establish a dynamic background model for image data within the region of interest. It removes dynamic water ripple backgrounds and highlights human body contours using a background removal algorithm based on semantic segmentation. It also distinguishes foreground targets from water backgrounds using background subtraction and separates high-brightness reflective areas using the HSV color space, preserves skin tone features, and extracts foreground images of hands.

[0062] The gesture state machine interaction unit is connected to the anti-interference preprocessing unit. It has a built-in state machine with idle state, detection state, confirmation state and execution state. It is used to match and verify the gesture features in continuous frame images. When the duration of the same gesture feature exceeds the preset time threshold, it is confirmed as a valid gesture and the corresponding control trigger signal is generated.

[0063] The gesture control mapping unit, connected to the gesture state machine interaction unit, is used to convert the spatial displacement vector of an effective gesture into motor control pulses; wherein, the direction of the displacement vector determines the adjustment direction, and the magnitude of the displacement vector determines the adjustment step size.

[0064] The long-distance posture recognition unit is connected to the video acquisition module. When the distance between the user and the video acquisition module exceeds a preset distance threshold, it switches to arm posture recognition mode and generates control commands based on the vector direction of the line connecting the wrist and elbow, which are then sent to the water flow drive module.

[0065] Understandably, existing gesture control schemes for pool surfers struggle to adapt to the complex physical environment of a pool. The pool environment presents three main sources of interference: water ripples and reflections can obscure hand features or cause image distortion; splashing water can generate noise that is misidentified as key hand points; and the variable postures of users—standing, crouching, or half-submerged—lead to inconsistent hand positions. To address these issues, this invention introduces dynamic region of interest locking, multimodal fusion filtering, and a state machine-based anti-mistouch mechanism to ensure accurate and reliable non-contact gesture control in complex underwater environments.

[0066] Specifically, the implementation process of the gesture recognition-based contactless control system is as follows:

[0067] S11, Dynamic Region of Interest (ROI) Locking. The dynamic ROI refers to a rectangular area for gesture detection generated with the center of the user's shoulder as a reference. After the user enters the field of view of the image acquisition module, human detection algorithms such as the lightweight object detection algorithm YOLOv8-tiny are used to lock the user's head and shoulder positions. A rectangular frame is then expanded outward from the shoulder center as a gesture detection window. As the body sways during surfing, the coordinates of this detection window are updated in real time, sending only moving objects within the detection window to the gesture recognition network, thus filtering out distant water birds, ripples, or background interference.

[0068] S12, Anti-water ripple interference preprocessing. The anti-interference preprocessing unit refers to the module that sequentially performs semantic segmentation background removal, background subtraction foreground extraction, and HSV color space reflection suppression on the image data within the region of interest. This is used to extract a clear hand foreground image from dynamic water ripples and high-brightness reflections. First, a dynamic background model is established, and a semantic segmentation-based background removal algorithm is used to remove the dynamic water ripple background and highlight the human body contour. Then, combined with the background subtraction method, only when the object and the background produce a significant grayscale or color difference and have a specific movement speed are they determined to be the foreground target, i.e., the hand. Finally, through HSV (Hue, Saturation, Value) color space separation, high-brightness white reflection points are filtered out, skin tone features are preserved, and a clean hand foreground image is extracted.

[0069] S13, Gesture State Machine Interaction Logic. The gesture state machine interaction unit refers to a state machine with four built-in states: Idle, Detect, Confirm, and Execute. It is used to perform spatiotemporal consistency verification of gesture features. Only when the same gesture feature is maintained for a duration exceeding a preset time threshold is it confirmed as a valid operation signal, thus preventing false triggering caused by unintentional actions such as wiping sweat or scratching the head. In the Idle state, it continuously monitors but does not perform any action. When a gesture candidate is detected, the non-contact control system based on gesture recognition enters the Detect state and starts a timer. If the gesture shape disappears within a time duration T1 (e.g., only a sweat-wiping action), it resets back to the Idle state and does not respond. If the same gesture remains stable for more than the preset time threshold T1 (e.g., 0.5 seconds), it enters the Confirm state and is determined to be a valid gesture, finally entering the Execute state to send control commands to the motor. This state machine mechanism effectively prevents false triggering caused by unintentional actions.

[0070] S14, Deep mapping of gesture commands and control parameters. The gesture control mapping unit converts the spatial displacement vector of an effective gesture into motor control pulses. For water flow rate control, a pushing gesture is defined as pushing the palm forward or backward with the five fingers together, detecting the movement trajectory of the palm perpendicular to the ground: pushing upward increases the water flow rate, and the acceleration is calculated based on the speed or distance of the palm's upward movement; the faster or higher the push, the greater the increase in flow rate level; pressing downward decreases the water flow rate. The pool surfer uses changes in the pitch of a buzzer to provide feedback on the currently set flow rate level, enabling blind operation confirmation. For deflector angle control, a horizontal waving gesture is defined as moving the hand left and right in the horizontal direction. Waving the hand to the left deflects the deflector to the left, creating a right-handed wave; waving the hand to the right deflects the deflector to the right, creating a left-handed wave. The amplitude of the hand wave determines the deflection angle; for example, a 30-degree wave corresponds to a 15-degree deflection of the deflector. For surge control, the opening and closing gestures are defined as the hands moving from together to open or from open to together. When the distance between the hands increases, the irregularity of the surge is increased to simulate the roughness of natural ocean waves. When the hands are closed, the surge is reduced to make the water flow laminar, which is suitable for long-distance gliding training.

[0071] S15, Optimized for Long-Distance Interaction. When the user is far from the video capture module (e.g., 3-5 meters), the pixel ratio of the hand is too small, making it difficult to recognize fine gestures. The long-distance posture recognition unit switches to arm posture recognition mode, no longer recognizing finger details, but instead detecting the vector direction of the line connecting the user's wrist and elbow, and generating control commands based on the direction of arm movement. For example, raising the entire arm overhead corresponds to the maximum power command, while the arm hanging naturally corresponds to the stop command.

[0072] This invention presents a non-contact control system based on gesture recognition. Through dynamic region of interest locking and multi-step anti-interference preprocessing, it effectively solves the recognition interference problems caused by water ripples, reflections, and splashes in swimming pool scenarios, significantly improving gesture recognition accuracy. A four-level state machine logic enables precise verification of valid gestures, completely avoiding the risk of accidental touches from unintentional actions and ensuring the safety of equipment control. Through multi-dimensional gesture control mapping and long-distance posture adaptation, it achieves fine-grained real-time adjustment of water flow parameters, covering all scenario control needs, replacing traditional physical button operations, and improving the convenience and safety of surfing device use.

[0073] In some embodiments of this application, reference is made to Figure 3 The motion training and evaluation system based on skeleton recognition includes:

[0074] The dynamic baseline calibration unit is used to collect static standing posture image data for a preset duration, extract neutral standing posture skeletal angles, and generate personalized motion evaluation baseline data for the user.

[0075] The upper and lower body decoupled tracking unit is used to perform high-precision skeletal key point detection on the user's upper body, and perform Kalman filtering on the detected skeletal point coordinates to suppress misjudgment caused by water flow turbulence; when the confidence of the lower body key points is lower than the preset confidence threshold, it switches to the centroid motion trend prediction mode, predicts the current joint coordinates based on the motion trajectory before being occluded, and outputs continuous human skeletal model data.

[0076] The motion parameter calculation unit is used to calculate quantitative evaluation parameters based on human skeletal model data and motion evaluation benchmark data; wherein, the quantitative evaluation parameters include hip joint stability score, posture accuracy rate, movement frequency and calorie consumption value.

[0077] The real-time guidance feedback unit is used to compare the quantitative evaluation parameters with preset standard thresholds. When the quantitative evaluation parameters deviate from the preset standard thresholds, corresponding real-time training guidance information is generated.

[0078] Understandably, in pool surfing scenarios, the user's body is in an unstable, dynamic environment. Water reflections and splashes often obscure key body points such as ankles and knees, making traditional skeletal recognition algorithms prone to failure. Furthermore, individual differences in height and leg length among users render the use of uniform absolute values ​​for motion assessment unscientific. Additionally, water currents cause random fluctuations in skeletal coordinates, affecting the accuracy of the assessment results. To address these issues, this invention achieves personalized assessment through dynamic baseline calibration, ensures the continuity and robustness of the skeletal model through decoupled upper and lower body tracking and Kalman filtering, and provides scientific motion guidance through multi-dimensional quantitative evaluation metrics.

[0079] In this embodiment, the specific implementation process of the motion training and evaluation system based on skeleton recognition is as follows:

[0080] The first step is dynamic baseline calibration. Before training begins, static standing posture image data is collected for a preset duration (e.g., 3 seconds) from the user. Neutral standing posture skeletal angles are extracted, and the user's personalized baseline skeletal angles are recorded. These personalized baseline values ​​are used as the subsequent dynamic evaluation reference values, rather than rigid absolute values. This baseline data includes the initial angle values ​​of key joints such as the hip, knee, and ankle when the user is standing.

[0081] The second step is decoupled tracking of the upper and lower body. The system performs high-precision skeletal keypoint detection on the user's upper body, using a lightweight, high-resolution network HRNet (High-Resolution Net) or the OpenPose algorithm for human pose estimation. HRNet is a convolutional neural network architecture that can maintain high-resolution feature representations and is suitable for keypoint detection in complex scenes. For the detected skeletal point coordinates, the system performs Kalman filtering, using the motion trajectory and velocity of the past few frames to predict the current coordinates and suppress misjudgments caused by water flow. The system assigns a confidence score to each identified skeletal point. When water splashes cause the confidence score of lower body keypoints such as ankles and knees to fall below a preset confidence threshold, the system switches to a centroid motion trend prediction mode, predicting the current joint coordinates based on the motion trajectory before occlusion, ensuring that the skeletal model does not fall apart. At the same time, the camera automatically locks onto the lower body region based on the surfboard's position, prioritizing the extraction of hip, knee, and ankle keypoints to ensure the accuracy of core motion capture, ultimately outputting continuous human skeletal model data.

[0082] The third step is the calculation of motion parameters. Based on human skeletal model data and motion assessment benchmark data, the motion parameter calculation unit calculates multiple quantitative assessment parameters, including hip joint stability score, posture accuracy, movement frequency, and calorie consumption.

[0083] The hip joint stability score is used to characterize the stability of the user's core, and its calculation formula is as follows:

[0084]

[0085] In the formula, Score stability The score represents hip joint stability, where α is the penalty coefficient, T is the total number of sampling points within the sampling duration, t is the time interval of a single sampling point, and Y is the average value of the hip joint. hip (t) represents the real-time motion value of the hip joint collected at time t. This is the arithmetic mean of the hip joint motion values ​​within time period T. The formula calculates the mean square error of the real-time hip joint position over the time series, subtracts a fluctuation penalty term from a base score of 100, and obtains a stability score; the smaller the fluctuation, the higher the score.

[0086] The core logic of hip joint stability score calculation is as follows: first, calculate the fluctuation range of real-time hip joint motion values ​​over time; then, map this fluctuation range back to a stability score; the smaller the fluctuation, the higher the score. The calculation process for hip joint stability score includes the following steps 1 to 3:

[0087] Step 1: Calculate the fluctuation at a single moment and obtain the motion deviation at a single moment.

[0088] For each sampling time t, calculate the deviation Y between the real-time hip joint motion value at that time and the mean hip joint motion value over the entire time period. hip (t)- This difference reflects the absolute deviation of the hip joint movement from the baseline mean at time t, and characterizes the instantaneous fluctuation amplitude at that time.

[0089] Step 2: Quantify the overall fluctuation and calculate the mean square error (MSE).

[0090] The deviation at each time step is squared to eliminate the problem of positive and negative deviations canceling each other out. Then, the squared deviations at all time steps are summed, and the arithmetic mean is taken to obtain the mean square error (MSE).

[0091]

[0092] A larger MSE indicates more severe overall fluctuations in hip joint movement during time interval T, and poorer core stability in the user. When MSE=0, the real-time hip joint motion values ​​Y at all times are... hip (t) are all equal to the benchmark mean. This indicates that the hip joint is completely stable and without any fluctuations.

[0093] Step 3: The fluctuation is mapped to a standardized score to obtain the final hip joint stability score.

[0094] Using 100 points as the baseline and a maximum score, the total penalty score is obtained by multiplying the penalty coefficient by the mean squared error (MSE). The final stability score is then obtained by subtracting the total penalty score from the maximum score.

[0095] Score stability =100−α·MSE

[0096] The formula establishes an inverse linear relationship between fluctuation amplitude and score. The greater the fluctuation of hip joint movement, the higher the MSE value, the higher the total penalty score α·MSE, and the lower the final score. Conversely, the smaller the fluctuation, the lower the MSE value, and the closer the final score is to the full score of 100, which is in line with users' intuitive understanding of stability rating.

[0097] Posture accuracy is used to characterize the degree to which a user's actions match a standard action, and its calculation formula is as follows:

[0098]

[0099] In the formula, Accuracy represents the posture accuracy rate, ranging from 0 to 1. The closer to 1, the higher the match between the user's action and the standard action, and the more standardized the posture. N represents the total number of key angles involved in the evaluation, which in this embodiment includes at least two core angles: the knee angle and the elbow angle. The two key angles are defined as follows: θ knee The knee angle is the angle formed by the lines connecting the hip, knee, and ankle. Standard surfing technique requires a slight bend in the knee, with a standard angle range of 120 to 150 degrees. elbow The elbow angle is the angle formed by the lines connecting the shoulder, elbow, and wrist, used to determine if the user's stroke posture is stiff. i is the key angle for a single assessment; β is the sensitivity coefficient, used to adjust the degree of influence of angle deviation on accuracy; θ user (i) represents the real-time angle value of the i-th key angle in the user's real-time action; θ standard (i) is the standard angle threshold of the i-th key angle in the standard action library.

[0100] Movement frequency (MRF) refers to the number of complete reciprocating cycles of a standard movement performed by a user's limbs per unit of time, including arm stroke frequency and body flexion / extension frequency. MRF is obtained by counting the number of reciprocating cycles of the limbs per unit of time, such as counting the number of times the wrist crosses the zero point in the left-right direction during arm strokes or leg push-off movements.

[0101] Calorie expenditure is used to characterize the energy consumption of a user during training, and its calculation formula is as follows:

[0102]

[0103]

[0104] In the formula, Power represents the user's real-time motion power, characterizing the energy consumed by the user's movement per unit time; joints represents all skeletal joints involved in the calculation; Mass... joints The equivalent mass of the corresponding joint; a joints v represents the real-time motion acceleration of the corresponding joint. joints K represents the real-time motion velocity of the corresponding joint. drag The water flow resistance coefficient is determined by the current flow rate setting of the surfer; the higher the setting, the greater the water flow resistance coefficient. flow This represents the current water flow speed of the pool surfer.

[0105] Power(t) represents the user's real-time exercise power at time t. The integration range is the user's effective training duration. By integrating the real-time power over time during the training period, the user's total energy consumption during training is obtained and converted into calorie values.

[0106] The fourth step is real-time guidance and feedback. The real-time guidance and feedback unit receives the quantitative evaluation parameters updated in real time, compares each parameter with preset standard thresholds, and generates corresponding real-time training guidance information when the quantitative evaluation parameters deviate from the preset standard thresholds. Specific implementation scenarios include three categories, fully covering the needs of different training stages.

[0107] The first scenario is a beginner's balance training scenario. The user selects a beginner balance course, the surfer is set to low-speed water flow, and the standard posture is feet shoulder-width apart, knees slightly bent, upper body upright, and eyes looking forward. When the system detects that the user's hip center point Y-coordinate fluctuates more than 15 centimeters within 2 seconds, it determines that the center of gravity is unstable; when the system detects that the user's knee angle remains consistently at 175 degrees, it determines that the knee is too straight. The real-time guidance feedback unit generates corresponding training guidance information, broadcasting a prompt through the user's interactive terminal's voice broadcast component: "Knee flexion is insufficient, please lower your center of gravity." Simultaneously, on the user's interactive terminal display screen, the user's virtual hip position is marked in red, and the standard position is displayed as a green guide box, prompting the user to adjust their posture.

[0108] The second scenario is for advanced gliding skill training. Users select a dynamic gliding course, the surfer activates a medium-speed water flow, and the standard movement involves the body springing and stretching with the waves. The real-time guidance and feedback unit uses augmented reality fusion technology to overlay a semi-transparent golden skeleton image of the standard movement onto the user's leg position in real-time on the user's interactive terminal display screen. It simultaneously displays the real-time synchronization rate between the user's movement and the standard movement, guiding the user to adjust their movements to align with the standard skeleton image, achieving immersive practice.

[0109] The third scenario is competitive paddling training. When a user selects a HIIT (High-Intensity Interval Training) course, the system focuses on monitoring the frequency and amplitude of the user's upper limb movements. By calculating the number of times the wrist crosses the zero point in the left-right direction, the system obtains the user's real-time paddling frequency. When the system detects that the user's paddling frequency is below a standard threshold, the real-time guidance feedback unit generates corresponding prompts, providing voice guidance that the current paddling frequency is too low and urging the user to increase the pace.

[0110] After training, the real-time guidance and feedback unit summarizes the user's quantitative evaluation parameters throughout the entire training period, generating a complete training report, which is then displayed to the user through the user's interactive terminal. The training report includes core data such as the user's hip stability score, average posture accuracy, training duration, calorie consumption, and movement frequency statistics. It also identifies the periods of lowest movement stability during the user's training and provides targeted improvement suggestions. Simultaneously, the system uploads the user's training data to the cloud for online competitive ranking, enabling cloud-based competitive interaction.

[0111] This invention provides a motion training and evaluation system based on skeletal recognition. Through decoupled upper and lower body tracking and a Kalman filter algorithm, it effectively solves the problem of skeletal tracking failure caused by water splashes and current turbulence in swimming pool scenarios, significantly improving the robustness of the algorithm in complex fluid environments. Relying on personalized benchmark calibration and a multi-dimensional biomechanical quantification model, it achieves scientific and accurate evaluation of surfing movements, replacing the traditional subjective judgment method. Through real-time multimodal error correction guidance and AR immersive training, it helps users standardize movements and avoid sports injuries. Simultaneously, it combines a cloud-based competitive mode to enhance the fun of training, comprehensively improving the professionalism and user experience of surfing training.

[0112] Figure 4 The flowchart of the visual interaction-based intelligent pool surfer training guidance method provided by the present invention is shown below. Figure 4 The present invention provides a training guidance method for a visually interactive intelligent pool surfer, comprising:

[0113] S1 obtains the training course selected by the user, calls up the preset teaching video or standard action video corresponding to the training course, and displays it to the user.

[0114] S2, which collects real-time images of the user's gestures and full-body movements in the water, and transmits the collected image data to the main control processing unit;

[0115] S3, perform gesture recognition based on the gesture image, generate control commands based on the recognition results and send them to the water flow drive module to adjust the water flow speed and spray direction;

[0116] S4. Based on the full-body motion image, extract skeletal points and evaluate movements, generate training guidance information in real time, and provide feedback to the user.

[0117] S5, after training, summarizes the user's training data throughout the entire period, generates and outputs a complete training report including training duration, calorie consumption, and posture score.

[0118] The visual interaction-based intelligent pool surfer training guidance method provided by this invention is applied to the visual interaction-based intelligent pool surfer training guidance system described in the foregoing embodiments. The system-related hardware architecture and algorithm logic have been described in detail in the foregoing embodiments, and will not be repeated here.

[0119] Based on the content of the above embodiments, step S3, which involves performing gesture recognition based on the gesture image, includes:

[0120] Based on the human detection algorithm, the user's head and shoulder positions are located. A region of interest for gesture detection is generated and updated in real time with the center of the shoulder as a reference. Gesture recognition is performed on the image data within the region of interest.

[0121] A dynamic background model is established, and a dynamic water ripple background is removed by a background removal algorithm based on semantic segmentation. The background difference method is combined to distinguish the foreground target from the water background, and the high-brightness reflective area is separated by the HSV color space to extract the hand foreground image.

[0122] The gesture features in consecutive frames of images are matched and verified. When the duration of the same gesture feature exceeds a preset time threshold, it is confirmed as a valid gesture. Control commands are generated based on the spatial displacement vector of the valid gesture.

[0123] Based on the content of the above embodiments, step S4, which involves skeletal point extraction and motion evaluation based on the whole-body motion image, includes:

[0124] Collect static standing posture image data for a preset duration from the user, extract the skeletal angles of the neutral standing posture, and generate personalized motion evaluation benchmark data for the user;

[0125] High-precision skeletal key point detection is performed on the user's upper body, and Kalman filtering is applied to the coordinates of the detected skeletal points. When the confidence of the key points of the lower body is lower than the preset confidence threshold, the centroid motion trend prediction mode is switched to predict the coordinates of the current joint points based on the motion trajectory before occlusion, and continuous human skeletal model data is output.

[0126] Based on human skeletal model data and motion assessment benchmark data, quantitative assessment parameters are calculated, including hip joint stability score, posture accuracy rate, movement frequency, and calorie consumption value.

[0127] The quantitative evaluation parameters are compared with preset standard thresholds. When the quantitative evaluation parameters deviate from the preset standard thresholds, corresponding real-time training guidance information is generated.

[0128] This invention replaces traditional physical operation with visual interaction technology, avoiding the safety hazards and accidental touches associated with wet hands, and achieving convenient and safe control of the surfer. Leveraging the real-time computing power of its built-in AI neural network processor, it can capture and scientifically evaluate the user's surfing movements in real time, providing professional training guidance and effectively preventing sports injuries caused by incorrect movements. Simultaneously, through an immersive visual interaction mode, it breaks away from the monotony of traditional surfing training, enhancing training fun and user engagement, and meeting the diverse surfing training needs of users.

[0129] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual interaction-based intelligent pool surfer training and guidance system, characterized in that, include: The main body of the pool surfer is installed on the side wall of the pool and has a built-in water flow drive module. The water flow drive module is used to receive control commands and adjust the water flow speed and jet direction. The video acquisition module is used to capture images of the user's hand gestures and full-body movements in the water; The main control processing unit is connected to the video acquisition module and the water flow driving module respectively, and is used to perform gesture recognition based on the gesture image and generate control commands according to the recognition results. Based on the full-body motion images, skeletal points are extracted and motion is evaluated to generate real-time training guidance information.

2. The intelligent pool surfer training and guidance system based on visual interaction according to claim 1, characterized in that, The main control processing unit includes: The non-contact control system based on gesture recognition is used to receive the gesture image, generate corresponding control commands through gesture recognition processing, and send the control commands to the water flow drive module; A motion training and evaluation system based on skeleton recognition is used to receive the full-body motion images, extract skeletal points to construct a human skeleton model, calculate motion parameters, and generate real-time training guidance information. The video training and cloud-based competition system outputs preset instructional videos or standard movement footage, compares the user's real-time movements with the standard movements to generate training scores, and uploads the user's training data to the cloud for online competition ranking.

3. The intelligent pool surfer training and guidance system based on visual interaction according to claim 2, characterized in that, The gesture recognition-based contactless control system includes: The dynamic region of interest (ROI) locking unit is used to acquire image data collected by the video acquisition module, lock the user's head and shoulder positions according to the human detection algorithm, generate and update the region of interest for gesture detection in real time with the shoulder center as a reference, and output only the image data within the region of interest. An anti-interference preprocessing unit, connected to a dynamic region of interest locking unit, is used to establish a dynamic background model for image data within the region of interest. It removes dynamic water ripple background and highlights human body contours through a background removal algorithm based on semantic segmentation. It distinguishes foreground targets from water backgrounds using background subtraction and separates high-brightness reflective areas using the HSV color space, preserves skin tone features, and extracts hand foreground images. The gesture state machine interaction unit is connected to the anti-interference preprocessing unit. It has a built-in state machine with idle state, detection state, confirmation state and execution state. It is used to match and verify the gesture features in continuous frame images. When the duration of the same gesture feature exceeds the preset time threshold, it is confirmed as a valid gesture and the corresponding control trigger signal is generated.

4. The intelligent pool surfer training and guidance system based on visual interaction according to claim 3, characterized in that, The gesture recognition-based contactless control system also includes: A gesture control mapping unit, connected to the gesture state machine interaction unit, is used to convert the spatial displacement vector of an effective gesture into motor control pulses; wherein, the direction of the displacement vector determines the adjustment direction, and the magnitude of the displacement vector determines the adjustment step size. The long-distance posture recognition unit is connected to the video acquisition module. When the distance between the user and the video acquisition module exceeds a preset distance threshold, it switches to arm posture recognition mode and generates control commands based on the vector direction of the line connecting the wrist and elbow, which are then sent to the water flow drive module.

5. The intelligent pool surfer training and guidance system based on visual interaction according to claim 2, characterized in that, The motion training and evaluation system based on skeleton recognition includes: The dynamic baseline calibration unit is used to collect static standing posture image data for a preset duration, extract neutral standing posture skeletal angles, and generate personalized motion evaluation baseline data for the user. The upper and lower body decoupled tracking unit is used to perform high-precision skeletal key point detection on the user's upper body, and perform Kalman filtering on the detected skeletal point coordinates to suppress misjudgment caused by water flow turbulence; when the confidence of the lower body key points is lower than the preset confidence threshold, it switches to the centroid motion trend prediction mode, predicts the current joint coordinates based on the motion trajectory before being occluded, and outputs continuous human skeletal model data. The motion parameter calculation unit is used to calculate quantitative evaluation parameters based on human skeletal model data and motion evaluation benchmark data; wherein, the quantitative evaluation parameters include hip joint stability score, posture accuracy rate, movement frequency and calorie consumption value. The real-time guidance feedback unit is used to compare the quantitative evaluation parameters with preset standard thresholds. When the quantitative evaluation parameters deviate from the preset standard thresholds, corresponding real-time training guidance information is generated.

6. The intelligent pool surfer training and guidance system based on visual interaction according to claim 5, characterized in that, The motion parameter calculation unit includes a hip joint stability scoring model, the expression of which is: In the formula, Score stability The score represents hip joint stability, where α is the penalty coefficient, T is the total number of sampling points within the sampling duration, t is the time interval of a single sampling point, and Y is the average value of the hip joint. hip (t) represents the real-time motion value of the hip joint collected at time t. This is the arithmetic mean of the hip joint motion values ​​during time interval T.

7. The intelligent pool surfer training and guidance system based on visual interaction according to claim 1, characterized in that, The water flow drive module includes a variable speed motor, an adjustable-angle nozzle, and a stepper motor for driving the nozzle to deflect. The variable speed motor and the stepper motor are respectively connected to the main control processing unit. The control commands of the main control processing unit include adjusting the PWM duty cycle of the variable speed motor to change the water flow speed, and adjusting the rotation angle of the stepper motor to change the spray direction of the nozzle.

8. The intelligent pool surfer training and guidance system based on visual interaction according to claim 1, characterized in that, The system also includes a user interaction terminal, which includes a display screen and a voice broadcast component; The display screen is used to display the user's image captured by the video acquisition module in real time, and to overlay a semi-transparent skeletal image of the standard movement on the user's image, while also displaying real-time training guidance information, training score data and online competition ranking data; The voice broadcast component is used to broadcast the real-time training guidance information in voice form.

9. A visual interaction-based intelligent pool surfer training guidance method, applied to the visual interaction-based intelligent pool surfer training guidance system as described in any one of claims 1-8, characterized in that, The method includes: The system retrieves the training course selected by the user, calls up the corresponding preset teaching video or standard action video, and displays it to the user. The system captures real-time images of the user's hand gestures and full-body movements in the water and transmits the captured image data to the main control processing unit. Based on the gesture image, gesture recognition is performed, and control commands are generated according to the recognition results and sent to the water flow drive module to adjust the water flow speed and spray direction; Based on the full-body motion images, skeletal points are extracted and motion is evaluated, training guidance information is generated in real time, and feedback is provided to the user. After training, the system summarizes the user's training data throughout the entire period, and generates and outputs a complete training report that includes training duration, calorie consumption, and posture score.

10. The training guidance method for an intelligent swimming pool surfer based on visual interaction according to claim 9, characterized in that, Gesture recognition based on the gesture image includes: Based on the human detection algorithm, the user's head and shoulder positions are located. A region of interest for gesture detection is generated and updated in real time with the shoulder center as a reference. Gesture recognition is performed on the image data within the region of interest. A dynamic background model is established, and a dynamic water ripple background is removed by a background removal algorithm based on semantic segmentation. The background difference method is combined to distinguish the foreground target from the water background, and the high-brightness reflective area is separated by the HSV color space to extract the hand foreground image. The gesture features in consecutive frames of images are matched and verified. When the duration of the same gesture feature exceeds a preset time threshold, it is confirmed as a valid gesture. Control commands are generated based on the spatial displacement vector of the valid gesture. Skeletal point extraction and motion assessment based on the full-body motion images include: Collect static standing posture image data for a preset duration from the user, extract the skeletal angles of the neutral standing posture, and generate personalized motion evaluation benchmark data for the user; High-precision skeletal key point detection is performed on the user's upper body, and Kalman filtering is applied to the coordinates of the detected skeletal points. When the confidence of the key points of the lower body is lower than the preset confidence threshold, the centroid motion trend prediction mode is switched to predict the coordinates of the current joint points based on the motion trajectory before occlusion, and continuous human skeletal model data is output. Based on human skeletal model data and motion assessment benchmark data, quantitative assessment parameters are calculated, including hip joint stability score, posture accuracy rate, movement frequency, and calorie consumption value. The quantitative evaluation parameters are compared with preset standard thresholds. When the quantitative evaluation parameters deviate from the preset standard thresholds, corresponding real-time training guidance information is generated.