Multi-view fusion augmented reality fitness action guidance method and system
By using a multi-view fusion augmented reality system, multiple cameras are used to collect fitness movements and equipment information, enabling high-precision analysis of complex fitness movements and comprehensive training records. This solves the problems of blind spots in posture recognition and inconvenient interaction in existing technologies, and improves the accuracy and consistency of fitness guidance.
Patent Information
- Application Number
- CN202511836347.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-02-24
AI Technical Summary
Existing fitness apps cannot effectively identify blind spots in complex fitness movements, lack contextualized interaction, have incomplete data recording, and have inconvenient interaction methods, resulting in unreliable posture analysis and incomplete training records.
The augmented reality system employs multi-view fusion, which acquires multi-angle video streams through first and second cameras, combines skeletal key point data and instrument recognition, provides real-time feedback and automated data recording, and supports voice and gesture interaction.
It improves the accuracy of posture recognition and the integrity of training data, ensuring the continuity and immersive experience of the fitness process, and reducing the risk of sports injuries and the error rate of data collection.
Smart Images

Figure CN121550671A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of augmented reality and computer vision technology, and more specifically, to a multi-view fusion augmented reality fitness movement guidance method and system. Background Technology
[0002] With the increasing health awareness of the public and the digital transformation of the fitness industry, providing scientific and safe personalized fitness guidance through technological means has become an important development trend. Existing technological solutions have significant shortcomings in deeply integrating virtual information with real training scenarios and accurately assessing and intervening in the quality of user movements in real time. Therefore, developing a system that can understand the training environment, perceive multi-dimensional movement information, and provide immersive guidance has significant practical implications.
[0003] Current mainstream technologies have inherent limitations. Smartphone-based fitness apps rely on a single front-facing camera, whose fixed frontal view physically fails to capture crucial posture information from the user's side or back. This results in blind spots in recognizing key evaluation criteria such as back curvature and knee joint trajectory during multi-joint compound movements like squats and deadlifts, making posture analysis and feedback unreliable. Furthermore, as closed "outsider observers," these systems cannot identify the user's actual fitness environment and equipment type, thus failing to provide contextualized guidance linked to specific equipment, limiting their application scope. In addition, their touchscreen-based interaction is cumbersome when the user's hands are occupied by equipment or they are fatigued, severely disrupting the continuity of training. At the data recording level, existing solutions generally lack the ability to automatically collect objective parameters of equipment (such as weight and speed), resulting in incomplete training logs and difficulty in forming a comprehensive user training profile.
[0004] Therefore, this paper proposes a multi-view fusion augmented reality fitness movement guidance method and system to address the above problems. The system aims to solve the following issues: how to overcome the blind spots in posture recognition caused by a single fixed viewpoint to achieve high-precision analysis of complex fitness movements; how to deeply integrate virtual guidance with the user's real physical training environment to achieve scenario-based interaction; how to provide a smooth, uninterrupted, and convenient human-computer interaction method in typical fitness scenarios where both hands are occupied; and how to achieve automated and seamless collection of multi-dimensional data from the training process to construct a complete training record. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a multi-view fusion augmented reality fitness movement guidance method and system to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a multi-view fusion augmented reality fitness movement guidance method, comprising the following steps: S1: The first-view image acquisition unit acquires video streams of the user's fitness environment in real time and identifies fitness equipment in the environment based on the video streams. S2: Receive training instructions from the user for the identified fitness equipment; S3: Activate at least one second-view image acquisition unit to acquire a real-time video stream of the user's fitness movements from an orientation different from that of the first-view image acquisition unit; S4: Based on the video stream acquired by the second perspective image acquisition unit, obtain the skeletal key point data of the user's body through pose estimation; S5: Compare the skeletal key point data with the preset standard movements to generate real-time feedback information on the standardization of the movements; S6: The feedback information is presented to the user through an augmented reality display terminal in a form that is superimposed on the user's real field of vision.
[0007] Preferably, after step S1, there is also step S1a: after identifying the fitness equipment, the augmented reality display terminal automatically retrieves and overlays standard movement instructions that match the equipment.
[0008] Preferably, before step S3, step S2a is included: generating and prompting the user with recommended observation orientation information of the at least one second-view image acquisition unit based on the identified fitness equipment and / or training movements.
[0009] Preferably, in step S2, receiving user training instructions is achieved through a non-contact interaction method; the non-contact interaction method includes at least one of voice interaction, gesture recognition, and eye tracking.
[0010] Preferably, the pose estimation process in step S4 includes: identifying multiple predefined skeletal key points of the human body from the video stream and calculating the angle or distance relationship between the key points; step S5 includes comparing the calculated angle or distance with a preset threshold range to determine the standardization of the action.
[0011] Preferably, the method further includes step S4a: during training, automatically acquiring the real-time operating parameters or configuration parameters of the fitness equipment, and fusing and recording the parameter information with posture data of the same time period; wherein, the parameters are acquired through image recognition technology or wireless communication connection with the fitness equipment.
[0012] Preferably, the physical form and deployment method of the at least one second-view image acquisition unit is any one or more combinations of an independent mobile device, a module integrated inside the fitness equipment, or a camera device fixedly installed in the fitness environment.
[0013] Preferably, when there are two or more second-view image acquisition units, step S4 further includes: fusing the skeletal key point data from different view image acquisition units to obtain an optimized pose estimation result.
[0014] A multi-view fusion augmented reality fitness movement guidance system, used to implement the above method, includes: Augmented reality display terminal, which integrates a first-view image acquisition unit for environmental perception and a display unit for information display; The computing unit is communicatively connected to the augmented reality display terminal and is equipped with an instrument recognition module and a posture estimation module. At least one second-view image acquisition unit, whose observation orientation is configured to complement the viewpoint of the first-view image acquisition unit, is used to acquire a video stream of the user's fitness movements and transmit it to the computing unit.
[0015] Preferably, the at least one second-view image acquisition unit is selected from any of the following forms or combinations: an independent portable camera device connected to the computing unit via a wireless network, a camera module integrated inside the fitness equipment and connected to the system via wired or wireless means, or a network camera device fixedly installed in the fitness environment and connected to the system network.
[0016] The technical effects and advantages of this invention are as follows: Compared to existing technologies, this invention introduces a second camera module physically separate from the augmented reality display terminal, which captures a video stream of user movements from an optimal viewing angle, such as the side or oblique rear, different from the first camera. This video stream is analyzed by a pose estimation module in the computing unit based on skeletal keypoint models such as MediaPipe or OpenPose to obtain complete body keypoint data. This approach fundamentally eliminates the blind spots present in a single frontal view, providing a reliable data foundation for judging back curvature and joint positions during movements such as squats and deadlifts. This significantly improves the accuracy and reliability of pose recognition, providing users with truly effective motion correction feedback and reducing the risk of sports injuries caused by posture errors.
[0017] Compared to existing technologies, this invention processes the video stream from the first camera using a device recognition module based on the YOLO series target detection algorithm. This automatically identifies the user's fitness environment and the type of equipment. Optionally, a device parameter recognition module based on the Tesseract OCR engine can automatically read performance parameters such as weight and speed from the equipment display panel. This design enables the system to perceive and understand the user's training scenario, allowing it to retrieve and display standard movement instructions matching specific equipment, and automatically integrate objective performance parameters with subjective movement quality data in the training log. This achieves intelligent linkage between virtual guidance content and the real physical environment, expanding the application scope from general bodyweight training to complex and diverse gym equipment training scenarios, improving the relevance of training and the completeness of data.
[0018] Compared to existing technologies, this invention integrates a voice interaction module, employing a workflow that first listens for a preset wake word to activate the system, then converts subsequent voice commands into control commands to receive user training control instructions. Simultaneously, the system integrates a gesture recognition module as a supplementary interaction method, using a first camera to capture specific gestures lasting longer than one second to execute commands. This multimodal interaction design, centered on non-contact interaction, allows users to complete system operations without interrupting the training process, even when holding the equipment with both hands or sweating profusely. This ensures the continuity of the training rhythm and the user's immersive experience, solving the inconvenience and interruption problems caused by touch interaction. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall framework of the present invention.
[0020] Figure 2 This is a schematic diagram of the workflow of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1 As attached Figures 1 to 2 The present invention relates to a multi-view fusion augmented reality fitness movement guidance method and system, which consists of three main hardware components: an augmented reality display terminal, a computing unit, and a detachable second camera module.
[0023] Augmented reality display terminals, typically all-in-one AR glasses, serve as the human-computer interaction interface. Their built-in first camera continuously captures video streams of the surrounding fitness environment from the user's first-person perspective, providing raw visual data for subsequent environmental understanding. The display unit is responsible for accurately overlaying the calculated virtual guidance information onto the user's real field of vision, thus achieving virtual-real fusion.
[0024] The computing unit, as the central processing unit of the system, is a device that operates via high-speed data lines (such as...). Standalone embedded devices connected to AR glasses (e.g., based on...) Platform construction can also be a high-performance processor integrated inside the glasses themselves.
[0025] It carries and runs a series of complex computer vision and artificial intelligence algorithms, including but not limited to instrument recognition, pose estimation, data fusion and feedback generation.
[0026] The detachable second camera module is physically independent of the AR glasses, and integrates an image sensor, battery, and wireless communication unit (such as supporting...). (Mode). Users can freely place it at the optimal observation point according to the needs of the training movement, for example, placing it approximately to the side of the body when performing squat training. This allows for the acquisition of a third-person video stream that complements the first-person perspective and is used for precise attitude analysis.
[0027] Based on the aforementioned hardware architecture, the system's workflow first involves sensing the environment and equipment, then receiving user commands, followed by the collection and fusion analysis of multi-view data, and finally transforming the analysis results into intuitive guidance information for the user.
[0028] Furthermore, after recognizing the fitness equipment through the first camera, the system automatically triggers the display function of standard movement instructions. This solves the problem of users, especially beginners, not being clear about the key points of the movements on unfamiliar equipment.
[0029] The underlying principle is that the system maintains a structured motion database, which associates each identifiable piece of fitness equipment (such as a squat rack, cable machine, or treadmill) with one or more standardized digital motion models. These digital models can be pre-rendered 3D animations or videos containing key pose sequences.
[0030] When the device recognition module uses a high confidence level (e.g., greater than), When the system detects the presence of a squat rack in the field of view, it immediately retrieves the standard animation of a barbell squat from its database. Then, using augmented reality spatial registration and rendering technology, it stably anchors this semi-transparent 3D animation model next to the real squat rack. Users can see a virtual "coach" continuously demonstrating the standard movements in their field of view without needing to consult their phones or manuals, lowering the learning curve and ensuring the accuracy of the movement imitation.
[0031] Furthermore, before activating the detachable second camera, the system proactively provides the user with a recommended placement location for the module. This ensures the quality of the input data required for subsequent posture analysis. The principle behind this is that different fitness movements have their own key observation dimensions, and a single, fixed perspective (especially a first-person perspective) cannot cover all of these dimensions.
[0032] The system has a pre-built, experience-based mapping knowledge base. It takes "equipment type" and "training exercise name" as input and outputs one or more recommended "observation angles" and "approximate distances." For example, when the input is (squat rack, barbell squat), the knowledge base outputs "frontal and lateral," to The recommended angle is "directly forward, slightly to the side" because this angle is best for observing squat depth, back curvature, and knee position; while when the input is (gantry, cable face pull), "directly forward, slightly to the side" might be recommended. The system projects virtual markers (such as circles and arrows) on the ground or equipment using an AR display unit, and provides voice prompts to guide the user in completing the placement.
[0033] This guiding step directly addresses and resolves the "fatal blind spot" issue of a single-camera viewpoint mentioned in the background technology. By ensuring the second camera is placed at the optimal side observation point, the system obtains the back curvature and squat trajectory information necessary for evaluating movements such as squats and deadlifts. This improves the measurement accuracy of key side posture angles from a state where it was impossible to determine them from a single frontal viewpoint to an absolute error of less than [value missing]. This level of accuracy provides a data foundation for subsequent high-precision attitude correction.
[0034] Furthermore, the system integrates automatic device parameter recognition during training, enabling fully automated recording of training data. This function primarily relies on optical character recognition (OCR) technology. Its workflow begins with image capture. When the system detects that the user's gaze is fixed on the equipment display panel (such as the speed screen on a treadmill or the weight display on a barbell plate), the first camera or another camera specified by the user focuses and captures one or more frames. Raw images often suffer from uneven lighting, reflections, and cluttered backgrounds, resulting in low accuracy when directly performing OCR recognition.
[0035] Therefore, the preprocessing stage begins: first, the color image is converted to a grayscale image. Subsequently, the Otsu adaptive thresholding algorithm is used for binarization. This algorithm can automatically calculate an optimal threshold T based on the image's grayscale histogram, dividing the image into foreground (characters) and background. After binarization, a mid-range filter is applied to eliminate salt-and-pepper noise. The preprocessed, clear binary image is then fed into an OCR engine (such as Tesseract) for character recognition. Finally, the recognized text (such as "10.0", "50") is parsed into structured parameter data (speed). ,weight ).
[0036] The OCR recognition process is triggered by a specific event. One triggering condition is a periodic timer (e.g., triggered every 30 seconds); a more preferred triggering condition is when the system detects that the user's gaze (estimated by the first camera) is continuously fixed on the device display panel for more than 1.5 seconds.
[0037] Once triggered, the system will capture a high-resolution image frame and first perform an image quality assessment, calculating its sharpness (such as the Laplacian gradient value).
[0038] If the clarity is below the threshold, the recognition attempt is abandoned and the system waits for the next trigger. Clear images are then fed into the aforementioned preprocessing and OCR recognition pipeline. If the confidence level returned by the OCR engine is below 80%, the system will not record the result, but will prompt the user via voice with "Parameter recognition failed, please look at the display screen."
[0039] Successfully identified parameters (such as weight) are immediately tagged with a timestamp and the current training context, and stored in a temporary database for the training session, binding them with posture data from the same time period. Through this automated and error-tolerant process, users don't need to painstakingly recall or manually input training loads after exhaustion sets; the system efficiently, seamlessly, and accurately records the data. This solves the core pain points of incomplete and error-prone data in traditional fitness recording, addressing the problem of "single data collection dimension, unable to form a complete training profile" in the background technology.
[0040] This frees users from the 5-minute manual recording task that typically occurs after each training session, and eliminates errors caused by fatigue or forgetfulness. It enables automated and seamless simultaneous collection of subjective motion quality data and objective performance parameters (such as weight and speed), providing a stronger data foundation for building comprehensive and objective user training profiles.
[0041] Furthermore, the posture estimation module's analysis process includes a quantitative assessment of the degree of motion completion, such as determining the squat depth. This determination relies on geometric calculations of the spatial relationships between key human body points. Posture estimation models (such as...) It will output the two-dimensional coordinates of 33 key skeletal points of the human body in the image coordinate system. For squat depth, the system focuses on the key points of the hip joint. Key points of the knee joint vertical coordinates and During the lowering phase of a standard squat, the hip joint gradually lowers until it is below the knee joint. Therefore, the system establishes a simple geometric criterion: When this condition is met, a single squat is considered to have reached the effective depth range, and the system counts it as a valid repetition. This quantitative, physically based criterion replaces the traditional vague "feeling" or the coach's subjective visual assessment, making the evaluation of movement standards consistent and objective.
[0042] Furthermore, posture analysis also assesses the quality of body posture, such as the back posture, which is crucial in squats and deadlifts. The system achieves this by calculating the angle between the torso and a baseline (such as a plumb line). Specifically, it uses the two-dimensional coordinates output by the posture estimation model to select the shoulder joint. Hip joint and ankle joint Three points. Vector. It can approximate the direction of the torso, while the vector It can be used to aid in calculations. This is to obtain the angle between the back and the vertical direction. It can be achieved by calculating vectors Vertical downward unit vector The solution can be found using the dot product, or through vectors. and The angle between the two sides indirectly reflects the tilt of the torso: The system presets an ideal safety angle threshold range (e.g.) to When the angle is calculated in real time For several consecutive frames below the lower limit of the safety threshold (e.g.) When this occurs, it is determined that the user has exhibited incorrect posture such as "lumbar collapse" or "pelvic tilt." This vector operation based on key point coordinates transforms the abstract posture problem into a measurable angular deviation, providing data for generating accurate corrective feedback.
[0043] Furthermore, to improve the robustness and accuracy of pose estimation, especially in cases of complex movements or partial occlusion, the system supports configuring multiple detachable second cameras. For example, one camera can be placed to the side, and another to the rear. The data fusion module of the computing unit receives both video streams simultaneously. The pose estimation module first performs independent skeletal keypoint detection on each video stream to obtain two-dimensional coordinate estimates of the same keypoint (such as the right shoulder joint) from two different viewpoints. and Due to the limitations of a single camera's field of view and estimation errors, these two coordinates may be inconsistent. The fusion module uses a weighted average strategy to calculate the final two-dimensional coordinates of the key point. : Among them, weight and The key point can be dynamically allocated based on factors such as visibility confidence and image clarity from each camera's perspective, and it must meet the following requirements: By fusing information from multiple perspectives, the system can effectively compensate for blind spots caused by occlusion in a single perspective and use geometric constraints between views to smooth out estimation errors of individual models, thereby obtaining a more stable and accurate sequence of key point positions on a two-dimensional plane.
[0044] Furthermore, at the end of the training session, the system automatically invokes the data recording and report generation module to provide users with a structured training summary. This module integrates all the fragmented data generated during training, forming a global view. It extracts and aggregates various information from the data stream: the number of movements, depth achievement rate, and trigger frequency of various posture warnings (such as "lower back collapse" and "knee valgus") from the posture analysis module; the average speed, maximum weight, and total number of sets from the device parameter recognition module; and the training duration and rest time between sets recorded by the system.
[0045] All this data is fed into a pre-defined visualization template, generating a comprehensive report that includes statistical tables and various charts (such as speed-time curves and motion quality radar charts). This report not only objectively records the "quantity" and "quality" of this training session, but also reveals the user's progress trends and potential bottlenecks by comparing them with historical data.
[0046] Furthermore, to enhance the system's adaptability to different environments, the augmented reality display terminal also integrates gesture recognition functionality. This function complements voice interaction, addressing the issue of decreased voice command recognition rates in noisy gym environments or when users are unable to speak. Its principle involves continuously capturing images of the user's hand area using a first camera and running a lightweight hand keypoint detection model (such as...). ).
[0047] The system predefines several simple, easily distinguishable, and clearly defined gestures and maps them to core control commands. For example, "Continuously push the palm forward..." "Mapped to "Pause / Continue" command, "thumbs up" This is mapped to a "Confirm / Start" command. By introducing a duration check, unintentional gesture triggers can be effectively avoided.
[0048] Once the user performs the required movement and holds it stably, the gesture recognition module generates the corresponding control signal. This contactless interaction method is well-suited to the characteristics of fitness scenarios where users' hands are often occupied by equipment or covered in sweat, ensuring that users can smoothly interact with the system in any environment.
[0049] The following example illustrates the workflow of a user performing barbell squats at a gym.
[0050] 1. Environmental Awareness and Initialization: The user puts on the AR glasses and walks to the squat rack. The AR glasses' first camera continuously captures environmental video streams and transmits them to the computing unit.
[0051] The real-time video stream captured by the first camera is transmitted to the computing unit in H.264 encoded format. The video decoding service within the computing unit first decodes it into consecutive RGB image frames. Subsequently, the instrument recognition module (based on the YOLOv8 model) performs inference on each frame and outputs a structured list containing bounding box coordinates, instrument category labels, and confidence scores.
[0052] When a "squat rack" is identified and the confidence level is higher than 90%, the module will publish an EquipmentIdentified event message to the data fusion module and the AR rendering module through an internal event bus (such as a microservice call based on gRPC). The message body includes the equipment type, the bounding box position in the image, and the confidence level.
[0053] Upon receiving this message, the AR rendering module immediately renders a green 3D border at the corresponding location in the user's AR field of view as a marker.
[0054] 2. Command Reception and Preparation: The user speaks the voice command "Start Squat Training". The voice interaction module is first activated by the wake word, then recognizes the command content, and the system officially enters the squat training mode.
[0055] Following this, based on the mapping knowledge base, the system locates the object approximately to the side of the user's body within their AR field of vision. A glowing virtual footprint icon is projected onto the ground, accompanied by a voice prompt: "Please place the split camera at the side footprint location."
[0056] The user places the second camera as instructed. At the same time, the system retrieves a 3D animation of a standard barbell squat from the database and overlays it next to the squat rack for the user's reference.
[0057] 3. Multi-view data acquisition and real-time analysis: The user begins to perform a squatting motion.
[0058] The second camera captures full-body motion video from a side view, and the encoded video stream is transmitted to the computing unit in real time via a Wi-Fi link operating in AP mode using the RTPoverUDP protocol.
[0059] The video stream receiving service within the computing unit receives the data stream through a Socket interface, performs decoding and buffering, and forms an image frame queue.
[0060] The pose estimation module (based on MediaPipePose) acts as a consumer, retrieving image frames sequentially from the queue and performing inference. For each frame, the module outputs a structured JSON object containing a timestamp, frame number, and an array of 33 skeletal keypoints, each containing its two-dimensional coordinates. and its visibility confidence This JSON object is pushed to the data cache pool of the data fusion module in real time.
[0061] 4. Data Fusion and Feedback Generation: The data fusion module maintains a timestamp-based sliding window cache pool for temporarily storing the most recent data. The data originates from asynchronous data from the instrument identification, attitude estimation, and equipment parameter identification modules.
[0062] When the pose estimation module sends a new frame of keypoint data, the fusion module uses its timestamp as a reference to search the cache pool for the closest timestamp (with a difference of less than 1). The module associates the device identification results with the data in this frame. Then, for this frame of data, the module performs the following core calculations and judgments: Squat depth determination: Read the ordinate of key points of the hip joint and the ordinate of key points of the knee joint In a certain frame, it was measured , ,satisfy The squat depth was determined to be valid.
[0063] Back posture assessment: Select the shoulder joint Hip joint Ankle Coordinates (unit: ). Calculate vectors , Substituting into the included angle formula: Keep your back straight at this time.
[0064] In another frame, due to the body leaning forward, the coordinates become... , , .
[0065] Calculated , The calculated included angle Reduce to The system immediately identified it as "waist collapse".
[0066] 5. Immersive Feedback Presentation: When the system detects a "back collapse" error, the feedback generation module is triggered. In the AR view, the user's back skeletal lines are highlighted in a striking red in real time. Simultaneously, the voice module broadcasts a prompt: "Please keep your back straight!" The system counts and displays the number of valid squats in the corner of the field of view.
[0067] 6. Automated Data Recording and Reporting: After completing a set of exercises, the user rests. During the training process, the system automatically recorded the weight of the barbell plates using OCR technology. At the end of the entire training session, the user says "End Training." The report generation module then starts, integrating all data from this training session: Total completed... Groups, 10 times per group, average depth attainment rate Back posture accuracy Using weight This generates a training report with both text and images and stores it in the user's personal log.
[0068] Example 2: Multiple Implementation Methods of Computing Units Based on the same inventive concept, the computing unit can adopt different physical forms and deployment methods according to the different requirements of the system for computing power, power consumption and integration.
[0069] In the first implementation, the computing unit takes the form of a standalone hardware box. This standalone box connects to the augmented reality display terminal via a USB-C data cable, and houses a high-performance embedded processor responsible for centrally running all algorithms for device recognition, attitude estimation, and data fusion. This approach facilitates heat dissipation and computing power expansion.
[0070] In the second implementation, the computing unit is highly integrated. Its hardware is miniaturized and integrated into the temples or main frame of the augmented reality display terminal, forming an all-in-one device. In this case, the system requires no external wiring, and all computing tasks are completed locally on the glasses, achieving maximum portability and ease of use.
[0071] In the third implementation, the computing unit adopts a distributed form with edge-cloud collaboration. The augmented reality display terminal and the detachable camera serve as data acquisition terminals, establishing a communication connection with a remote cloud server via a 5G or Wi-Fi network. The device recognition module and posture estimation module are deployed on the cloud server; the data acquisition terminal is responsible for encoding and uploading the video stream data, and receiving skeletal keypoint data and feedback command results from the cloud server. This approach leverages the powerful computing capabilities of the cloud, enabling the execution of more complex models while reducing the hardware cost and power consumption of the terminal device.
[0072] Example 3: Alternative Solution for Environmental Perception and Parameter Acquisition Based on the same inventive concept, the system can use a variety of technical approaches to perceive environmental and instrument parameters.
[0073] In terms of equipment recognition, besides using deep learning models such as YOLO for label-free recognition, the system can also employ a visually labeled recognition scheme. Specifically, visually identifiable labels, such as QR codes or ArUco tags, are pre-set at specific locations on the fitness equipment. The equipment recognition module captures images containing these labels using a first camera and quickly obtains the equipment's identity information and preset parameters by calculating the visual features of the labels. This scheme offers fast recognition speed and high stability, making it particularly suitable for equipment with inconspicuous features or scenarios requiring extremely fast initialization.
[0074] Regarding equipment parameter acquisition, in addition to image analysis via optical character recognition (OCR) technology, the system can also employ a direct sensor connection solution. Specifically, the equipment parameter recognition module establishes a data connection with smart fitness equipment possessing wireless communication capabilities via Bluetooth or Wi-Fi communication interfaces, directly reading structured parameter data from the equipment's built-in sensors. Furthermore, for traditional weight plates lacking electronic display functionality, a physical tag identification solution can be used: passive RFID or NFC tags storing weight information are attached to the weight plates; simultaneously, a corresponding RFID reader or NFC sensing module is integrated into the augmented reality display terminal or computing unit. When the user picks up the weight plate, the weight parameters stored in the tag are instantly read via near-field communication (NFC) technology.
[0075] Example 4: Alternatives to the Interaction Method Based on the same inventive concept, the system's human-computer interaction methods can integrate other non-contact interaction modalities in addition to voice.
[0076] The augmented reality display terminal can be further integrated with an eye-tracking module. This module captures the user's eye movements using a built-in infrared camera and calculates the coordinates of the user's gaze point within the display unit's field of view in real time using an eye-tracking algorithm. When the system detects that the user's gaze point remains on a virtual interactive control (such as the "Start" button) for more than a preset threshold of 1 second, it triggers the corresponding operation command for that control. This eye-tracking interaction scheme provides an efficient and silent control method for users when their hands are occupied by equipment or in noisy environments.
[0077] Finally, the following points should be noted: First, in the description of this application, it should be noted that, unless otherwise specified and limited, the terms "installation", "connection", and "linkage" should be interpreted broadly, and can be mechanical or electrical connections, or internal connections between two components, or direct connections. "Up", "down", "left", "right", etc. are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may change. Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other. In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-view fusion augmented reality fitness movement guidance method, characterized in that, Includes the following steps: S1: The first-view image acquisition unit acquires video streams of the user's fitness environment in real time and identifies fitness equipment in the environment based on the video streams. S2: Receive training instructions from the user for the identified fitness equipment; S3: Activate at least one second-view image acquisition unit to acquire a real-time video stream of the user's fitness movements from an orientation different from that of the first-view image acquisition unit; S4: Based on the video stream acquired by the second perspective image acquisition unit, obtain the skeletal key point data of the user's body through pose estimation; S5: Compare the skeletal key point data with the preset standard movements to generate real-time feedback information on the standardization of the movements; S6: The feedback information is presented to the user through an augmented reality display terminal in a form that is superimposed on the user's real field of vision.
2. The augmented reality fitness movement guidance method with multi-view fusion according to claim 1, characterized in that, The process includes step S1a after step S1: after identifying the fitness equipment, the augmented reality display terminal automatically retrieves and overlays standard movement instructions that match the equipment.
3. The augmented reality fitness movement guidance method with multi-view fusion according to claim 1, characterized in that, Before step S3, there is also step S2a: based on the identified fitness equipment and / or training movements, generate and prompt the user with recommended observation orientation information of the at least one second-view image acquisition unit.
4. The augmented reality fitness movement guidance method with multi-view fusion according to claim 1, characterized in that, In step S2, receiving user training instructions is achieved through a non-contact interaction method; the non-contact interaction method includes at least one of voice interaction, gesture recognition, and eye tracking.
5. The augmented reality fitness movement guidance method with multi-view fusion according to claim 1, characterized in that, The pose estimation process in step S4 includes: identifying multiple predefined skeletal key points of the human body from the video stream and calculating the angle or distance relationship between the key points; step S5 includes comparing the calculated angle or distance with a preset threshold range to determine the standardization of the action.
6. The augmented reality fitness movement guidance method with multi-view fusion according to claim 1, characterized in that, It also includes step S4a: during the training process, automatically acquire the real-time operating parameters or configuration parameters of the fitness equipment, and fuse the parameter information with the posture data of the same time period and record it; wherein, the parameters are acquired through image recognition technology or wireless communication connection with the fitness equipment.
7. The augmented reality fitness movement guidance method with multi-view fusion according to claim 1, characterized in that, The physical form and deployment method of the at least one second-view image acquisition unit can be any one or more combinations of an independent mobile device, a module integrated inside the fitness equipment, or a camera device fixed in the fitness environment.
8. The augmented reality fitness movement guidance method with multi-view fusion according to claim 7, characterized in that, When there are two or more second-view image acquisition units, step S4 further includes: fusing the skeletal key point data from different view image acquisition units to obtain an optimized pose estimation result.
9. A multi-view fusion augmented reality fitness movement guidance system, used to implement the method of any one of claims 1 to 8, characterized in that, include: Augmented reality display terminal, which integrates a first-view image acquisition unit for environmental perception and a display unit for information display; The computing unit is communicatively connected to the augmented reality display terminal and is equipped with an instrument recognition module and a posture estimation module. At least one second-view image acquisition unit, whose observation orientation is configured to complement the viewpoint of the first-view image acquisition unit, is used to acquire a video stream of the user's fitness movements and transmit it to the computing unit.
10. The multi-view fusion augmented reality fitness movement guidance system according to claim 9, characterized in that, The at least one second-view image acquisition unit is selected from any of the following forms or combinations: an independent portable camera device connected to the computing unit via a wireless network, a camera module integrated inside the fitness equipment and connected to the system via wired or wireless means, or a network camera device fixedly installed in the fitness environment and connected to the system network.