Intelligent AI golf teaching and accompanying system

Through the hybrid computing architecture of intelligent AI cameras and mobile terminals, combined with a dual-camera system and inertial measurement unit, the problems of high cost, poor portability and lack of interactivity of golf teaching equipment are solved, personalized swing analysis and emotional feedback are achieved, and teaching efficiency and user experience are improved.

CN120661897APending Publication Date: 2025-09-19SHANGHAI XUNLING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510861446.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing golf teaching equipment and software are expensive, have poor portability, lack personalized analysis and interactivity, and have insufficient emotional connection, making it difficult to achieve multi-dimensional data fusion and real-time comprehensive analysis.

Method used

It adopts a hybrid computing architecture of intelligent AI cameras, mobile terminals and cloud servers. Through the combination of dual-camera system, inertial measurement unit and mobile terminal, it realizes real-time collection and analysis of swing posture and trajectory data, generates multimodal feedback information, and combines cloud-based deep analysis and local lightweight models to provide personalized teaching and emotional interaction.

Benefits of technology

It realizes personalized swing analysis and teaching, improves teaching efficiency, provides instant voice guidance and emotional feedback, reduces equipment costs, and enhances user experience and interactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120661897A_ABST
    Figure CN120661897A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent AI golf teaching and accompanying system which comprises an intelligent sensing terminal, a mobile terminal and a cloud server, and the intelligent sensing terminal is an intelligent AI camera terminal and comprises a master control processing unit SoC, a dual-camera system, an inertial measurement unit IMU, a display screen, a network module and a storage module. The beneficial effects of the invention are that the system can help users with different heights, body types, strengths and habits through double-camera cooperation, teaching according to materials, observation of swing actions, tracking of the actual flight path of the ball, ballistic result analysis, reverse optimization and guidance of swing postures, and a result-oriented teaching method, and can improve the training efficiency of the users with different heights, body types, strengths and habits. The most efficient exclusive swing scheme which is really suitable for the user is found, and the cloud-local hybrid computing architecture ensures that the core analysis function can still be provided when the network is poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to intelligent systems, and specifically to an intelligent AI golf teaching and accompanying system. Background Art

[0002] Currently, golf enthusiasts rely primarily on traditional coaching, general video analysis software, or commercially available golf assistive devices to improve their skills. Traditional coaching is expensive, inflexible, and difficult to provide long-term, personalized tracking and analysis of a user's swing data. While commercially available golf analysis devices, such as trajectory trackers based on radar or dedicated high-speed cameras, can provide shot data, they are typically expensive, bulky, and difficult to port. Their functionality focuses primarily on data presentation, lacking effective instructional interaction and personalized training recommendations.

[0003] Existing technical teaching methods suffer from numerous flaws and shortcomings, including high costs and limited accessibility: Professional coaches and high-end analysis equipment are expensive, limiting their use by casual enthusiasts. Lack of personalization and continuity: Existing solutions struggle to track each user's swing data over time and conduct in-depth, personalized analysis, making it difficult to dynamically adjust teaching plans based on their unique style and progress. Feedback is unintuitive and interactive: Existing devices and software often offer cold, unintelligible data, lacking effective interactive mechanisms and targeted training guidance. Lack of emotional connection: Existing products are often purely tools, lacking emotional interaction and companionship, making them difficult to motivate continued use and generate positive emotional value. Inadequate intelligent and automated user experience: Existing devices often require complex manual setup and frequent angle adjustments, resulting in a high barrier to entry, inconvenience, and a lack of intelligent features like automatic alignment. Data is limited in dimension and comprehensive analysis capabilities are weak: It is difficult to effectively integrate and analyze multi-dimensional information, such as swing form, golf ball flight trajectory, and impact sound, in real time. Summary of the Invention

[0004] In view of the above shortcomings, the present invention adopts the following technical solutions: An intelligent AI golf teaching and companionship system includes an intelligent sensing terminal, a mobile terminal, and a cloud server. The intelligent sensing terminal is an intelligent AI camera, which mainly includes a main control processing unit, a dual-camera system, an inertial measurement unit (IMU), a display screen, a network module, and a storage module. The mobile terminal is a mobile phone and runs a supporting application for data interaction and visual feedback. The dual-camera system includes a posture camera and a trajectory camera. The posture camera is designed to operate continuously in a low-power mode to detect the user's golf preparation posture and trigger system wake-up. The trajectory camera is designed to activate after the system wakes up and collect raw data on the golf ball's motion trajectory. The synchronous posture camera collects raw data on the club's motion trajectory and swing posture data. The inertial measurement unit (IMU) is configured to sense the device's posture and output a gravity vector. The reference horizontal plane is automatically calibrated based on the gravity vector for three-dimensional reconstruction of the trajectory data. The main control processing unit is configured as follows: In response to a ready posture detected by the posture camera, waking up the system and activating the track camera; Synchronously collect posture camera video, trajectory camera raw data and IMU sensor data; Based on the reference horizontal plane calibrated by the IMU, the main control processing unit will calculate the running trajectory and optimize the calibration according to the data, and simultaneously calculate the hitting parameters; Combine the front-facing pose video recorded by the mobile phone's front camera and the side-facing pose video recorded by the pose camera to extract skeletal key points, which makes the extraction more accurate. Generate multimodal feedback information based on the analysis results; Controlling the display to dynamically display emotional visual content, in addition to displaying corresponding expressions when providing feedback suggestions, normal status information such as low battery and medium connection status will also be synchronously fed back through expressions; Feedback information and visual data are sent to the mobile terminal through the network module for interactive presentation and cloud server storage summary.

[0005] Furthermore, the main control processing unit is further designed to send instructions to the mobile terminal after the system wakes up, triggering the mobile terminal to capture supplementary perspective video through the front camera of the mobile phone and capture the audio of the hitting sound through the microphone; receive the supplementary data collected by the mobile terminal, and store it synchronously with the local data according to the timestamp; send the generated voice feedback data to the mobile terminal for broadcast.

[0006] Furthermore, the main control processing unit is designed to pre-process the multimodal data collected locally to generate structured data; when the network status meets the conditions, the structured data is uploaded to the cloud server, and the cloud visual language model VLM combines the user's historical data to generate in-depth analysis results; when the network status does not meet the conditions, the core analysis results are generated through the local lightweight model.

[0007] Furthermore, the multimodal feedback information includes an action playback video that integrates the annotation of key points of the human skeleton and highlighted problems; a virtual ballistic trajectory projected in the action playback video; VIP users are identified by the main control processing unit, and ordinary users receive coaching-style voice guidance generated by the text-to-speech TTS engine. Emotional attitudes are specified through text, and the main control selects corresponding expressions from the expression library for display on the display screen; VIP users can obtain comprehensive feedback directly generated by the multimodal large model, including voice, text and video highlights, and emotional emoticons are dynamically presented through the display screen.

[0008] Furthermore, it also includes a magnetic automatic alignment gimbal, the head of which is provided with a magnetic interface that matches the intelligent sensing terminal, an integrated wide-angle camera and a control algorithm, which visually recognizes human targets and drives the motor to rotate, so that the dual cameras of the intelligent sensing terminal are automatically aligned with the user.

[0009] Furthermore, the main control processing unit is a system-on-chip (SoC), including an artificial intelligence co-processing module, providing at least 6 TOPS computing power; it includes a hardware codec that supports concurrent encoding and decoding of 1080p@60fps video streams; the trajectory camera outputs linear Raw Bayer format raw data or 4k images with a resolution of up to 4K@60fps, and is used for both outdoor scenes and indoor simulator scenes to capture the hitting data and trajectory displayed on the simulator screen; the posture camera outputs a 1080p@60fps RGB video stream processed by the ISP.

[0010] Furthermore, the front of the OLED display screen of the intelligent sensing terminal is covered with a customized black glass panel; since the dual-camera system is not installed perpendicular to the glass panel, its imaging light cone intersects with the inclined glass panel, forming two oblique elliptical transparent areas; the cover color piece behind the glass panel and the camera opening on the front shell are also designed to be elliptical, and the hole wall is a conical surface that matches the dual-camera system, presenting an anthropomorphic "eyelash" design, forming a smart visual effect.

[0011] Furthermore, the system operation process is as follows: R1. Use the posture camera to monitor the user's preparation posture with low power consumption and trigger the system to wake up; R2. The main control processing unit activates the tracking camera and simultaneously starts the mobile terminal's supplementary data collection; R3. Automatically calibrate the reference horizontal plane based on the IMU gravity vector using the built-in IMU sensor module; R4. Collect and store multimodal data and identify the timestamp of the hitting event; R5. Local data preprocessing: extracting human skeleton key points, reconstructing 3D trajectories, and calculating hitting parameters; R6. Select cloud or local model to generate analysis results based on network status; R7. Dynamically generate multimodal feedback, including text feedback, video annotation, voice guidance, and emotional expressions; R8. Send the feedback data to the mobile terminal for visual display and transmit it to the cloud server at the same time. The cloud server will store the data and summarize the feedback at a specific time to achieve personalized teaching.

[0012] Furthermore, the timestamp of the hitting event includes the moment when the posture camera detects that the ball is hit and records 1 second before the start of the backswing as the starting timestamp; 1 second after the hitting is the ending timestamp; the posture camera will always monitor the action of starting the backswing, and only if there is indeed a hitting action after the backswing will it record the timestamps of starting the backswing, 1 second before the start of the backswing, hitting the ball, and 1 second after the hitting; if no actual hitting is detected after the backswing, these data will be discarded, and the system will save the complete multimodal data from the start to the end timestamps through the memory of the storage module.

[0013] Furthermore, the main control processing unit is a system-on-chip (SoC), and its operating system runs an operating system based on the Linux kernel, preferably Ubuntu.

[0014] The beneficial effects of this invention are: 1. Dual-camera synergy enables personalized instruction, observing swing movements and tracking the actual trajectory of the ball. By analyzing trajectory results, reverse-optimizing and guiding swing posture, this results-oriented teaching method can help users of different heights, body types, strengths, and habits find the most effective swing plan that truly suits them, rather than relying on a cookie-cutter imitation. 2. This invention offers faster response times, higher data processing efficiency, and better protection of user privacy. Its unique "cloud-on-premises" hybrid computing architecture ensures that core analysis functions are still available even in poor network conditions, while leveraging the more powerful cloud-based Visual Language Model (VLM) for in-depth analysis when the network is good, balancing reliability with top-tier performance. 3. The anthropomorphic expression design makes the AI ​​more dynamic and approachable, providing emotional encouragement and support to users, building long-term trust and a sense of companionship. 4. The system not only analyzes posture and trajectory videos but also integrates audio information from the shot through a mobile app, enabling comprehensive analysis of multimodal visual and auditory data for more accurate judgments and providing instant voice guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of the system operation when the network is good; Figure 2 This is a schematic diagram of the AI ​​camera dual-camera system structure of the present invention.

[0016] Reference numerals: 1-posture camera, 2-track camera, 3-glass panel, 4-main control processing unit, 5-display screen. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0018] Example 1 The hardware configuration of the present invention is as follows: An intelligent AI golf teaching and companionship system, mainly consisting of an AI camera hardware terminal (BirdieCoach hardware), a mobile application (App) running on the user's smartphone, and an optional automatic alignment magnetic gimbal System computing architecture: This system uses a "cloud-local" hybrid computing architecture. When the network is good, data is uploaded to the cloud server and deeply analyzed using the powerful visual language model (VLM). When the network is poor or offline, the core AI calculations are completed by the main control processing unit built into the BirdieCoach hardware to ensure the availability of basic functions. The user's smartphone only serves as a display and interaction medium and does not perform AI algorithm calculations. BirdieCoach Hardware Module: System-on-Chip (SoC): Serving as the "brain" of the system, it integrates a multi-core CPU, GPU, video codec unit (VPU), and artificial intelligence co-processing module. AI computing power: The AI ​​co-processing module (which can be implemented by a combination of NPU, CPU, GPU, or DSP) must have a computing power of at least 6 TOPS, and the ideal goal is about 13 TOPS to meet the requirements of algorithms such as YOLO and OpenPose.

[0019] Hardware codec: Integrated hardware codecs supporting formats such as H.264 and H.265, capable of concurrently encoding at least one 1080p@60fps video stream and decoding one 1080p@60fps network video stream.

[0020] Operating system: running an operating system based on the Linux kernel (preferably Ubuntu); Dual-camera system: Accurate frame synchronization acquisition is achieved through hardware trigger signals generated by the SoC's GPIO pins; Posture Camera 1: Used to capture the user's swing motion. Its ISP processing flow retains standard basic processing such as color correction, white balance, and noise reduction, but does not include electronic image stabilization (EIS) or additional filters, and outputs a high-quality 1080p@60fps RGB video stream. The ISP must have excellent auto exposure (AE) and wide dynamic range (WDR / HDR) processing capabilities to adapt to changing outdoor lighting conditions. Its horizontal field of view (FOV) is approximately 65°; Track Camera 2: Used to capture golf ball flight paths. To ensure computational accuracy, its image data undergoes basic correction, bypassing the ISP's nonlinear processing and directly outputting raw data in linear Raw Bayer format at up to 4K resolution and 60fps. Its focal length is no less than 7mm, its focusing distance is approximately 2 meters, and its horizontal field of view (FOV) is approximately 70°. General camera requirements: Both cameras must support manual / auto ISO control via the API, with manual settings taking precedence over auto. The cameras must be optimized for low light conditions to ensure clear recognition of digital content on the indoor golf simulator screen. Track Camera 2 must support real-time, frame-by-frame adjustment of the CMOS region of interest (ROI) reading via the API to reduce data volume. Emotional display and interaction module: The front of the device features an OLED display 5, and the front of the screen is covered with a custom black glass panel 3, achieving a "borderless" and black-out visual effect. The screen will display dynamic expressions driven by AI based on real-time status (such as thinking, celebrating), enabling emotional interaction with users; The camera opening on the glass panel 3 uses a unique anthropomorphic "eyelash" design; IMU sensor module: This module includes a built-in 6-axis IMU sensor with automatic calibration and low long-term drift, used to obtain device attitude. The fused attitude information it outputs relative to gravity (such as pitch and roll angles) is primarily used to correct the reference horizontal plane of the trajectory tracking algorithm to improve the accuracy of trajectory calculations. Audio interaction module: The device does not have a built-in speaker or microphone. Audio interaction is achieved through cooperation with the user's mobile app: audio input (such as the sound of hitting a ball) is received from the mobile app via Wi-Fi, and system-generated voice feedback (such as coaching instructions) is packaged as audio information and transmitted to the mobile phone, which is played through the phone's speaker or connected headphones. The system must be able to receive and transmit network audio streams; Network module and communication protocol: The Wi-Fi module must have throughput sufficient to concurrently process 1080p@60fps video stream transmission, 1080p@60fps video stream reception, and audio stream reception. High-reliability data such as control instructions and analysis results are transmitted via connection-oriented protocols such as TCP / IP; Real-time streaming data such as video and audio are transmitted via low-latency, connectionless protocols such as UDP / RTP; Storage and Memory: Configure at least 16GB of available storage space and at least 4GB (8GB is optimal) of available memory.

[0021] The system operation process is as follows: Key AI algorithm process: S1. System Initialization and Connection: The user launches the BirdieCoach hardware and mobile app, which automatically establish a connection using SoftAP or STA mode. If using a gimbal, the gimbal uses its built-in camera to visually identify the user and automatically aligns the gimbal.

[0022] S2. Equipment Readiness and Calibration: Before each tracking session, BirdieCoach hardware uses the IMU sensor to obtain the gravity vector and automatically calibrates the reference horizontal plane to prepare for subsequent accurate trajectory calculations.

[0023] S3. Intelligent perception and simultaneous multimodal data acquisition: The BirdieCoach hardware's posture camera 1 operates continuously at low power consumption, and the system wakes up when it detects the user assuming a golf stance and starting a swing.

[0024] The system immediately activates the track camera 2 to start collecting data, and sends instructions to the mobile terminal phone app to start collecting the swing video with a supplementary perspective through the front camera and the audio of the hitting sound through the microphone.

[0025] When gesture camera 1 detects the ball being hit, the system records the start timestamp of the hit event. A hit event ends approximately one second after the hit. Based on the start and end timestamps, the system saves data from all sensors (gesture video, raw trajectory data, mobile phone video, and mobile phone audio) during this time period.

[0026] S4. Data Preprocessing and Analysis: Local preprocessing: BirdieCoach hardware preprocesses the collected multimodal data, including extracting human skeleton key points from the swing video; cropping the audio of the impact sound into valid segments; reconstructing the 2D trajectory data of Track Camera 2 into a 3D trajectory based on the IMU-corrected horizontal plane; and reversely calculating various impact parameters.

[0027] Hybrid AI analysis: When the network is good, preprocessed structured data is sent to the cloud, where it is deeply analyzed by the Visual Language Model (VLM) combined with the user's historical data. When the network is poor, a local lightweight model performs core analysis.

[0028] S5. Hierarchical multi-dimensional feedback generation: Generated content: The system generates different levels of feedback based on the user's VIP status. For example, VIP users may receive comprehensive feedback directly generated by a large multimodal model, including voice, text, and video highlights; ordinary users may receive text feedback generated by another model.

[0029] Voice feedback: The system converts text feedback into coaching-style voice through the local text-to-speech (TTS) engine, and transmits the generated audio information to the mobile app, which then plays it.

[0030] Emotional feedback: During the analysis and feedback process, corresponding dynamic expressions such as thinking and encouragement will be displayed synchronously on the front OLED screen.

[0031] Visual data feedback: The system generates an action replay video with highlighted annotations (e.g., key posture issues) and sends the action replay video along with the projected ballistic trajectory to the mobile app for visual display.

[0032] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

[0033] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. An intelligent AI golf teaching and accompanying system, characterized by: The system comprises an intelligent sensing terminal, a mobile terminal, and a cloud server, wherein the intelligent sensing terminal is an intelligent AI camera, and the intelligent AI camera mainly comprises a main control processing unit (4), a dual-camera system, an inertial measurement unit (IMU), a display screen (5), a network module, and a storage module; the mobile terminal is a mobile phone and runs a supporting application for data interaction and visual feedback; the dual-camera system comprises a posture camera (1) and a track camera (2), and the posture camera (1) is designed to operate continuously in a low-power mode to detect the user's golf preparation posture and trigger system wake-up; The trajectory camera (2) is designed to be activated after the system wakes up, and collects the original data of the golf ball's motion trajectory, and the synchronous posture camera (1) collects the original data of the club's motion trajectory and the swing posture data; the inertial measurement unit (IMU) is configured to sense the device posture and output the gravity vector, and automatically calibrate the reference horizontal plane based on the gravity vector for three-dimensional reconstruction of the trajectory data; The main control processing unit (4) is configured to: In response to a ready posture detected by the posture camera (1), waking up the system and activating the track camera (2); Synchronously collect the posture camera (1) video, the trajectory camera (2) raw data and the IMU sensor data; Based on the reference horizontal plane calibrated by the IMU, the main control processing unit will calculate the running trajectory and optimize the calibration according to the data, and simultaneously calculate the hitting parameters; Combine the front posture video recorded by the front camera of the mobile terminal and the side posture video recorded by the posture camera (1) to extract the skeleton key points, so that the extraction is more accurate; Generate multimodal feedback information based on the analysis results; Controlling the display screen (5) to dynamically display emotional visual content, in addition to displaying corresponding expressions when providing feedback suggestions, status information such as low battery and medium connection status will also be synchronously fed back through expressions; Feedback information and visual data are sent to the mobile terminal through the network module for interactive presentation and cloud server storage summary.

2. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: The main control processing unit (4) is further designed to send a command to the mobile terminal after the system wakes up, triggering the mobile terminal to capture the supplementary perspective video through the front camera of the mobile phone and the audio of the hitting sound through the microphone; Receive the supplementary data collected by the mobile terminal and store it synchronously with the local data according to the timestamp; send the generated voice feedback data to the mobile terminal for broadcast.

3. The intelligent AI golf teaching and accompanying system according to claim 1 or 2, characterized in that: The main control processing unit (4) is designed to pre-process the locally collected multimodal data to generate structured data; When the network status meets the conditions, the structured data is uploaded to the cloud server, and the cloud-based visual language model (VLM) combines the user's historical data to generate in-depth analysis results; When the network status does not meet the conditions, the core analysis results are generated through the local lightweight model.

4. The intelligent AI golf teaching and accompanying system according to claim 3, characterized in that: The multimodal feedback information includes an action playback video that integrates the annotation of key points of the human skeleton and highlighted problems; a virtual ballistic trajectory projected on the action playback video; VIP users are identified by the main control processing unit (4), and ordinary users receive coaching voice guidance generated by the text-to-speech TTS engine. Emotional attitudes are specified by text, and the main control selects corresponding expressions from the expression library and displays them on the display screen; VIP users can obtain comprehensive feedback directly generated by the multimodal large model, including voice, text, video highlight areas, and expressions from the specified expression library, and emotional expression symbols are dynamically presented through the display screen (5).

5. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: It also includes a magnetic automatic alignment gimbal, the head of which is provided with a magnetic interface that matches the intelligent sensing terminal, an integrated wide-angle camera and a control algorithm, which recognizes human targets through vision and drives the motor to rotate, so that the dual cameras of the intelligent sensing terminal are automatically aimed at the user.

6. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: The main control processing unit (4) is a system-on-chip (SoC), including an artificial intelligence co-processing module, providing at least 6 TOPS computing power; The system includes a hardware codec that supports concurrent processing of encoding and decoding of 1080p@60fps video streams; the trajectory camera (2) outputs raw data in a linear RawBayer format or outputs a 4k image with a resolution of up to 4K@60fps, and is used in both outdoor scenes and indoor simulator scenes to capture the hitting data and trajectory displayed on the simulator screen; The posture camera (1) outputs a 1080p@60fps RGB video stream processed by the ISP.

7. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: The front of the OLED display screen (5) of the intelligent sensing terminal is covered with a customized black glass panel (3); since the dual-camera system is not installed perpendicular to the glass panel (2), its imaging light cone intersects with the inclined glass panel (3), forming two oblique elliptical transparent areas; the cover color piece behind the glass panel (3) and the camera opening on the front shell are also designed to be elliptical, and the hole wall is a conical surface that matches the dual-camera system, presenting an anthropomorphic "eyelash"-shaped design, forming a smart visual effect.

8. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: The system operation process is as follows: R1. Use the posture camera (1) to monitor the user's preparation posture with low power consumption and trigger the system to wake up; R2. The main control processing unit (4) activates the track camera (2) and simultaneously starts the supplementary data collection of the mobile terminal; R3. Automatically calibrate the reference horizontal plane based on the IMU gravity vector using the built-in IMU sensor module; R4. Collect and store multimodal data and identify the timestamp of the hitting event; R5. Local data preprocessing: extracting human skeleton key points, reconstructing 3D trajectories, and calculating hitting parameters; R6. Select cloud or local model to generate analysis results based on network status; R7. Dynamically generate multimodal feedback, including text feedback, video annotation, voice guidance, and emotional expressions; R8. Send the feedback data to the mobile terminal for visual display and transmit it to the cloud server at the same time. The cloud server will store the data and summarize the feedback at a specific time to achieve personalized teaching.

9. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: The timestamp of the ball hitting event includes the moment when the posture camera (1) detects that the ball is hit and records the moment when the backswing starts as a starting timestamp; and the end timestamp of 1 second after the ball is hit; The posture camera (1) will always monitor the action of starting the backswing. If there is indeed a hitting action after the backswing, the timestamps of starting the backswing, 1 second before starting the backswing, hitting the ball, and 1 second after hitting the ball will be recorded; if no actual hitting is detected after the backswing, these data will be discarded, and the system will save the complete multimodal data from the start to the end timestamps through the memory of the storage module.

10. The intelligent AI golf teaching and accompanying system according to claim 1, characterized in that: The main control processing unit (4) is a system-on-chip (SoC), and its operating system runs an operating system based on the Linux kernel.