Fitness guidance system and method based on AI posture analysis and MR human-computer interaction

By integrating AI posture analysis and MR human-computer interaction technology in the fitness guidance system, the problem of insufficient personalization and interaction in the existing fitness guidance technology is solved, and accurate analysis of user exercise status and recommendation of personalized fitness plans are achieved, improving user experience and motivation.

CN119649997BActive Publication Date: 2025-05-02SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510175712.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-02
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing fitness guidance technology lacks personalization and interactivity, the accuracy and real-timeness of the sports posture analysis algorithm are insufficient, the fitness plan recommendation algorithm is low in personalization, and the AI ​​interaction intelligence is limited, resulting in poor user experience.

Method used

The fitness guidance system based on AI posture analysis and MR human-computer interaction is adopted, including a 3D depth global camera module, an AI-based fitness plan recommendation module, an AI-based fitness movement analysis module and an AI-based virtual coach interaction module. Through multimodal data fusion and deep learning technology, personalized fitness guidance and feedback are provided.

Benefits of technology

It realizes accurate capture and analysis of user's exercise status, provides personalized fitness plans and instant feedback, improves users' fitness experience and motivation, and enhances the intelligence of AI interaction and user participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649997B_ABST
    Figure CN119649997B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent fitness guidance system and method based on AI posture analysis and MR human-computer interaction. The system includes: a 3D depth global camera module for capturing RGBD image data of a user; an AI-based fitness plan recommendation module for planning a fitness plan for the user; a dynamic adjustment module for analyzing the user's physical fitness changes and dynamically adjusting the training plan according to the user's actual physical fitness; an instant feedback module for collecting user feedback on the training effect; an AI-based fitness action analysis module for identifying and analyzing the user's fitness actions based on the user's multimodal fitness data; an AI-based virtual coach interaction module for real-time communication with the user based on a large fitness knowledge model; a 3DMR imaging-based fitness action demonstration module for projecting a virtual 3D fitness coach and a virtual character similar to the user's body shape. The invention can provide users with intelligent and personalized fitness guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The system of the present invention relates to the fields of computer vision technology, deep learning technology, 3D depth camera array and augmented reality technology, and in particular to a fitness guidance system and method based on AI posture analysis and MR human-computer interaction. Background Art

[0002] In today's society, as people's health awareness increases, fitness has become a common lifestyle. However, the existing fitness business model has great drawbacks and shortcomings; fitness guidance and interactive technology also have obvious limitations and deficiencies, making it difficult to meet market demand.

[0003] From a business model perspective, the limitations of existing fitness guidance technology are particularly prominent for fitness users. Traditional fitness methods often lack personalization and interactivity, making it difficult for users to maintain interest and motivation. At the same time, complex and diverse fitness equipment also increases the learning difficulty and time cost of users, further weakening the user's fitness experience.

[0004] On the technical level, the corresponding functions of some existing auxiliary hardware equipped with intelligent modules, as well as software devices such as fitness apps, cannot meet the actual needs of fitness people. First, the current motion posture analysis algorithm is inaccurate, has a limited range of freedom, and poor real-time performance: the existing posture analysis algorithm often cannot reach the ideal level of accuracy when identifying complex movements and subtle posture changes, making it difficult for users to obtain accurate feedback and guidance; many algorithms have recognition blind spots or delays when processing large-scale or high-speed movements, and cannot fully capture the user's motion state; and due to the limitations of computational complexity and hardware performance, some algorithms perform poorly in real time and cannot give immediate feedback, affecting user experience. Secondly, the existing fitness plan recommendation algorithm has shortcomings such as low personalization, poor dynamic adaptability, and lack of comprehensive coordination ability: the current fitness plan recommendation algorithm is often based on simple user preferences or historical data, which makes it difficult to achieve true personalization and cannot meet the diverse needs of different users; when the user's physical condition, fitness goals or environmental conditions change, the algorithm cannot adjust the recommended plan in time, resulting in poor fitness results; and lacks comprehensive consideration of the user's living habits such as diet and sleep, and cannot provide a comprehensive health management solution. The interactive intelligence of existing fitness system AI also has great limitations. The existing virtual interactive AI still needs to be improved in natural language understanding and generation, and it is difficult to communicate with users smoothly and naturally. Due to the lack of emotional computing and interaction capabilities, it cannot adjust the interaction strategy according to the user's emotional changes, which will directly affect the user experience and customer stickiness; and when providing fitness guidance and suggestions, there is a lack of immediate and intuitive feedback mechanism, which makes it difficult for users to quickly understand and master the correct movements. At present, the mainstream camera array technology does not capture depth information enough, resulting in inaccurate spatial information acquisition of key points of the human body in posture analysis. Due to unreasonable array layout and insufficient viewing angles, the field of view is not fully covered and there are blind spots, which cannot fully capture the user's movement status. In terms of performance, due to inaccurate calibration or insufficient data processing capabilities, the camera array often cannot meet the requirements in terms of real-time and stability. Most of the augmented reality devices on the market currently rely on physical media such as display screens as interactive interfaces, which limits the immersion and convenience of user experience. In complex or dynamic environments, the environmental perception capabilities of current devices are often insufficient to provide users with accurate guidance and feedback.

[0005] For example, the "Fitness Coaching System Using Personalized Augmented Reality Technology" disclosed in the existing patent document KR101970687B1 can provide fitness guidance, but it cannot provide customized fitness services based on the user's personal exercise level and preferences. It also lacks comprehensive consideration of the user's historical fitness data, current status, etc., and it is difficult to provide a comprehensive fitness plan; although the camera module is used for motion recognition, it lacks depth information capture and the posture analysis capability is limited; in addition, in the virtual reality feedback mode, the solution is limited to using non-embodied media such as display screens as the interactive interface, which greatly restricts the immersion and convenience of use. Summary of the invention

[0006] In order to solve at least one of the problems existing in the prior art, the present invention provides a fitness guidance system and method based on AI posture analysis and MR (mixed reality) human-computer interaction to provide users with intelligent and personalized fitness guidance.

[0007] In order to achieve the purpose of the present invention, the fitness guidance system and method based on AI posture analysis and MR human-computer interaction provided by the present invention include the following modules:

[0008] 3D depth global camera module, used to capture RGBD image data of users during exercise and transmit it to the server;

[0009] The AI-based fitness plan recommendation module is used to plan personalized fitness plans for users. The AI-based fitness plan recommendation module includes the LightGCN-att algorithm module, the dynamic adjustment module and the instant feedback module. The LightGCN-att algorithm module generates a training plan based on the user's historical fitness data, current status and fitness level score, combined with the preset fitness goals; the dynamic adjustment module is used to analyze the user's physical fitness changes in real time and dynamically adjust the training plan according to the user's actual physical fitness; the instant feedback module is used to collect user feedback on the training effect;

[0010] An AI-based fitness action analysis module is used to identify and analyze the user's fitness actions based on the user's multimodal fitness data to obtain the user's fitness action analysis results;

[0011] The AI-based virtual coach interaction module is used to communicate with users in real time based on a large model of fitness knowledge.

[0012] Furthermore, in the 3D depth global camera module, the common feature points in the multi-perspective RGBD image data collected by multiple 3D depth cameras are matched through the multi-view stereo matching module, and high-precision three-dimensional point cloud data is generated; the matching errors are corrected through the deep learning-based matching compensation module, and data gaps are filled or erroneous feature points are repaired during the matching process; and the real-time calibration module is used to perform dynamic calibration between multiple 3D depth cameras.

[0013] Furthermore, the multi-view stereo matching module uses an improved multi-view stereo matching algorithm to accurately match common feature points in multi-view RGBD image data collected by multiple 3D depth cameras to obtain preliminary matching results. The improvements made by the improved multi-view stereo matching algorithm include: introducing a multi-scale feature fusion mechanism to extract the approximate outline and shape features of the image at low resolution to obtain low-resolution features, and to capture more subtle texture and edge features at high resolution to obtain high-resolution features, and then fuse the low-resolution features with the high-resolution features; optimizing the cost aggregation strategy to increase the weight of clear and reliable feature points and reduce the weight of unreliable feature points.

[0014] Furthermore, the real-time calibration module adopts an adaptive calibration algorithm to accurately match feature points in the overlapping fields of view between 3D depth cameras, eliminates optical distortion and angle errors through a dynamic weight mechanism, and performs global optimization of three-dimensional point cloud data through a depth alignment mechanism; the dynamic weight mechanism refers to real-time adjustment of the weight distribution of feature points in the matching process between different 3D depth cameras according to the reliability of the feature points, and the depth alignment mechanism uses the common scene space reference points in the multi-view RGBD image data obtained by multiple 3D depth cameras to align the three-dimensional point cloud data.

[0015] Furthermore, the LightGCN-att algorithm module includes the LightGCN-att fitness plan recommendation algorithm. The LightGCN-att fitness plan recommendation algorithm includes a graph convolutional network LightGCN embedded with an attention mechanism. In the graph convolutional network LightGCN, by constructing a user-fitness activity bipartite graph and using the node's neighbor information to update the node's embedded representation, the similarity between users and the correlation between fitness activities are captured, and the attention mechanism is used to dynamically adjust the attention to different features in the user's fitness data.

[0016] Furthermore, the AI-based fitness action analysis module adopts a multimodal posture analysis algorithm to identify fitness actions; the multimodal posture analysis algorithm includes a convolutional neural network CNN and a Transformer model, and the Transformer model includes an encoder, a decoder and a classifier; or the multimodal posture analysis algorithm includes a convolutional neural network CNN, a Transformer model and a classifier, and the Transformer model includes an encoder and a decoder;

[0017] The convolutional neural network (CNN) is used to extract features from RGBD image data to obtain visual features and three-dimensional spatial features, and to combine the visual features, three-dimensional spatial features and physical features into multimodal features, wherein the physical features include kinematic features, which are extracted from physiological index data and motion data;

[0018] The encoder is used to fuse and encode multimodal features to obtain a comprehensive feature representation, which contains the key features of the user action;

[0019] The decoder is used to decode the encoded comprehensive feature representation into a natural language description or animation demonstration;

[0020] The classifier is used to analyze the comprehensive feature representation output by the encoder, and obtains the fitness action analysis result by comparing the comprehensive feature representation of the current posture with the preset standard posture template.

[0021] Furthermore, a generative adversarial network is applied to the training of the Transformer model. The generative adversarial network includes a generator and a discriminator. The generator generates multimodal pseudo-features by learning the distribution of multimodal fitness data, and the discriminator is used to compare multimodal features with multimodal pseudo-features to improve the recognition ability of pseudo data.

[0022] Furthermore, it also includes a fitness movement demonstration module based on 3DMR imaging, which is used to project a virtual 3D fitness coach and a virtual character with a body shape similar to the user so that the virtual 3D fitness coach can demonstrate the correct movements and the virtual character can reflect the user's movements in real time.

[0023] Furthermore, the fitness action demonstration module based on 3DMR imaging includes an intelligent hardware device based on augmented reality, a posture estimation network and a skeleton mapping model, wherein the intelligent hardware device based on augmented reality is used to project the virtual 3D fitness coach and the virtual character; the posture estimation network is used to generate the skeleton data of the user, and the skeleton data of the user includes the three-dimensional coordinates of the joint points and the connection relationship between the joint points; the skeleton mapping model generates a standardized skeleton model based on the skeleton data of the user; and the virtual character is constructed based on the standardized skeleton model;

[0024] The process of generating a standardized skeleton model by mapping the skeleton mapping model includes:

[0025] Parse and preprocess the user's skeleton data;

[0026] Determine the root node in the user's skeleton data;

[0027] Predict the connection relationship between the joint points in the user's skeleton data to obtain a set of the user's skeleton connection lines;

[0028] Optimize the user's bone connection set to obtain the preliminary structure of the skeleton;

[0029] Align the preliminary structure of the skeleton with the standard skeleton template to obtain a standardized skeleton model;

[0030] The process of constructing a virtual character based on the standardized skeleton model comprises the following steps:

[0031] Use a 3D modeling method based on mesh generation to add appearance information to the standardized skeleton model;

[0032] Based on texture mapping technology, add skin, clothing and other visual elements to the surface of standardized skeleton models;

[0033] Bind the appearance mesh of the avatar to the standardized skeleton model;

[0034] Rendering of virtual characters.

[0035] The fitness guidance method based on AI posture analysis and MR human-computer interaction provided by the present invention is implemented by the aforementioned system and includes the following steps:

[0036] The AI-based fitness plan recommendation module personalizes fitness plans based on the user's multimodal fitness data;

[0037] The user exercises according to the fitness plan, and the user's physiological index data and motion data are collected during the fitness process. The user's RGBD image data is collected in real time through the 3D depth global camera module, and the multi-modal fitness data is transmitted to the server;

[0038] The AI-based fitness action analysis module recognizes and analyzes the user's fitness actions in real time based on multimodal fitness data, performs posture recognition and posture evaluation, and obtains fitness action analysis results;

[0039] The AI-based virtual coach interaction module communicates and guides users in real time based on a large fitness knowledge model;

[0040] When the system also includes a fitness action demonstration module based on 3DMR imaging, the fitness action demonstration module based on 3DMR imaging performs 3D projection in an intelligent hardware device based on augmented reality, projects a virtual 3D fitness coach and a virtual character, and the user performs corresponding fitness actions according to the action demonstration of the virtual 3D fitness coach, and the virtual character reflects the user's action status in real time.

[0041] Compared with the prior art, the present invention can at least achieve the following beneficial effects:

[0042] (1) The present invention collects multi-view RGBD image data of a user by setting a 3D depth global camera module, which can capture depth information and comprehensively capture the user's motion state.

[0043] (2) In the AI-based fitness action analysis module, a multimodal posture analysis algorithm is used, and a Transformer-based adversarial neural network is adopted for recognition and analysis, which can improve the accuracy and robustness of posture analysis.

[0044] (3) In the AI-based fitness plan recommendation module, the LightGCN-att algorithm module is used to generate personalized training plans. The LightGCN-att algorithm module combines the advantages of the lightweight graph convolutional network (LightGCN) and the attention mechanism through the combination of graph convolutional networks and attention mechanisms to conduct in-depth analysis of user fitness data to customize personalized training plans.

[0045] (4) In terms of user interaction, artificial intelligence trained based on a large language model is used to communicate and guide users in real time. The fitness action demonstration module based on 3DMR imaging can use augmented reality technology to provide a scientific, safe and realistic interactive experience.

[0046] (5) The present invention can accurately identify the user's posture and movements, provide personalized guidance and correction, and bring the user an intelligent and efficient fitness experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a framework diagram of a fitness guidance system based on AI posture analysis and MR human-computer interaction provided by an embodiment of the present invention.

[0048] Figure 2 It is a schematic diagram of the fitness guidance process of the fitness guidance system based on AI posture analysis and MR human-computer interaction provided by an embodiment of the present invention.

[0049] Figure 3 It is a schematic diagram of the working process of the fitness guidance system based on AI posture analysis and MR human-computer interaction provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0051] See also Figure 3 The fitness guidance system based on AI posture analysis and MR human-computer interaction provided by the present invention includes a 3D depth global camera module, an AI-based fitness plan recommendation module, an AI-based fitness action analysis module, an AI-based virtual coach interaction module and a 3DMR imaging-based fitness action demonstration module.

[0052] The 3D depth global camera module is used to capture multi-view RGBD image data of users in the fitness scene in real time through multiple 3D depth cameras arranged in the fitness scene, and generate high-precision three-dimensional point cloud data.

[0053] In the 3D depth global camera module, the multi-view stereo matching module is used to accurately match the common feature points in the multi-view RGBD image data collected by multiple 3D depth cameras, and generate high-precision 3D point cloud data. The matching objects are the same object or human joints from different perspectives; the matching compensation module based on deep learning is used to correct the matching errors caused by noise or data loss; the real-time calibration module is used to perform dynamic calibration between multiple 3D depth cameras to ensure the geometric consistency and data fusion accuracy between 3D depth cameras. The multi-view stereo matching module, the matching compensation module based on deep learning, and the real-time calibration module are all installed in the server. The data collected by the 3D depth camera is uploaded to the server for processing, thereby ensuring the calculation efficiency of the algorithm and reducing the dependence on the camera-side hardware.

[0054] In order to improve the accuracy and real-time performance of fitness posture analysis, the present invention sets a 3D depth global camera module to collect the user's fitness data in the fitness scene. The 3D depth global camera module uses multiple 3D depth cameras arranged in the gym, combined with stereo vision technology and structured light technology, to capture the user's posture when using fitness equipment in real time. The high-resolution sensor and wide-angle lens of each 3D depth camera can ensure accurate capture in complex lighting environments and dynamic scenes.

[0055] In some embodiments of the present invention, the 3D depth camera of the present invention selects a 3D depth camera device Astra+ that integrates an RGB color camera and a depth sensor, which can simultaneously capture a color image and depth information of a scene.

[0056] In some embodiments of the present invention, in order to fully cover the fitness scene, the 3D depth camera device Astra+ is arranged in the fitness scene in the form of a linear array or a matrix array. This arrangement can ensure a wide coverage of the field of view while reducing blind spots. When arranging, factors such as field of view coverage, angle selection, distance adjustment, and synchronization mechanism need to be comprehensively considered to ensure that all 3D depth cameras can capture as comprehensive and detailed scene information as possible.

[0057] Among them, the multi-view stereo matching module uses an improved multi-view stereo matching algorithm to accurately match the common feature points in the multi-view RGBD image data collected by multiple 3D depth cameras to obtain preliminary matching results.

[0058] In order to improve the reconstruction accuracy of three-dimensional point cloud data, the present invention adopts an improved multi-eye stereo matching algorithm. Compared with the traditional multi-eye stereo matching algorithm, the improvements made to the improved multi-eye stereo matching algorithm include:

[0059] ①Introduced a multi-scale feature fusion mechanism.

[0060] The multi-scale feature fusion mechanism includes: extracting the general outline and shape features of the image at low resolution to obtain low-resolution features; capturing more subtle texture and edge features at high resolution to obtain high-resolution features, and fusing the low-resolution features with the high-resolution features.

[0061] The present invention can analyze images at multiple scales and extract feature points at different scales, thereby capturing texture information in the image more comprehensively and improving the matching robustness of complex texture areas. Through this multi-scale analysis, the improved multi-view stereo matching algorithm can better handle complex texture areas and improve the accuracy and robustness of feature matching.

[0062] ②Optimize the cost aggregation strategy.

[0063] The cost aggregation strategy uses a dynamic weight allocation mechanism to assign different weights to feature points. The dynamic weight allocation mechanism is: increase the weight of clear and reliable feature points, and reduce the weight of unreliable feature points.

[0064] Reduce error accumulation in occluded areas and high-contrast scenes through a dynamic weight allocation mechanism: In occluded areas or high-contrast scenes, some feature points may become unreliable due to environmental factors. At this time, by reducing the weights of these feature points, their impact on the final matching results can be reduced. For clear and reliable feature points, their weights are increased to ensure that they play a greater role in the matching process. Reliability refers to whether the feature point can maintain consistency and stability in RGBD image data from different perspectives, and its contribution to the matching results. The improved multi-view stereo matching algorithm analyzes the depth features and texture features of the pixel area around the input feature point, combines the multi-dimensional characteristics of the feature point such as contextual information, pixel continuity and gradient change, and outputs a confidence score to indicate whether the feature point is a clear and reliable matching point. This dynamic weight allocation mechanism helps to reduce error accumulation and improve matching accuracy.

[0065] Among them, the preliminary matching results are finely adjusted through the matching compensation module based on deep learning to enhance the correction ability of noise points and missing points in the image. The matching compensation module based on deep learning uses deep learning technology to identify and compensate errors in the matching process through a neural network model. Identification errors refer to the detection of possible noise points or missing points by analyzing the differences between feature points and surrounding background features. Compensation errors refer to the generation of matching results for missing areas or the correction of noise points through the prediction ability of the neural network model to make them consistent with the features in the real scene.

[0066] The implementation of matching compensation is based on the deep feature learning ability of the neural network model, that is, through large-scale training data, learning the regularity and common features in multi-view RGBD image data, so as to fill data gaps or repair erroneous feature points in the matching process, making the final matching results more accurate and robust.

[0067] The deep learning model can learn complex feature representations, thereby predicting the correct matching points in the presence of noise or missing data. The deep learning-based matching compensation module can be used as a post-processing step to fine-tune the initial matching results and enhance the algorithm's ability to correct noise points and missing points, thereby achieving higher-precision 3D reconstruction.

[0068] Among them, the real-time calibration module is used to perform dynamic calibration between multiple 3D depth cameras, and the following strategies are used to ensure the geometric consistency and data fusion accuracy between 3D depth cameras: ① The real-time calibration module uses the embedded gyroscope to detect the tiny displacement of the 3D depth camera, and dynamically adjusts the internal and external parameters of each 3D depth camera to ensure that the spatial position and viewing angle between the 3D depth cameras are accurately aligned; ② An adaptive calibration algorithm is used to accurately match the feature points in the overlapping fields of view between 3D depth cameras, and a dynamic weight mechanism is used to eliminate optical distortion and angular errors; ③ Combined with the depth alignment mechanism of multi-view images, that is, the spatial reference points of the common scene are used to perform global optimization of the three-dimensional point cloud data, thereby significantly improving the calibration accuracy and stability of the system under multi-camera configuration.

[0069] Specifically, the adaptive calibration algorithm can accurately match feature points within the overlapping fields of view between 3D depth cameras, and can dynamically adjust the calibration weights according to the actual environment and 3D depth camera characteristics. By analyzing the RGBD image data captured by the 3D depth camera in real time, the adaptive calibration algorithm can recognize and adapt to different lighting conditions, camera characteristics, and scene changes. This recognition and adjustment is achieved through a dynamic weight mechanism. After analyzing the lighting conditions, camera characteristics, and scene changes, the adaptive calibration algorithm automatically calculates the reliability score of each feature point, and dynamically adjusts the weight of the feature point based on the reliability score, thereby giving higher weights to high-reliability feature points during the calibration process, thereby improving the accuracy of feature point matching.

[0070] The adaptive calibration algorithm eliminates optical distortion and angular errors through a dynamic weight mechanism. The dynamic weight mechanism automatically adjusts the weights between different 3D depth cameras to compensate for the distortion caused by the different positions and angles of the 3D depth cameras. The dynamic weight mechanism refers to the real-time adjustment of the weight distribution of feature points in the matching process between different 3D depth cameras based on their reliability. Feature points refer to the corresponding spatial points in the multi-view RGBD images collected by the 3D depth camera. These points usually include edges, corners, and texture-significant areas in the scene. The reliability of feature points refers to the consistency, clarity, and texture integrity of the feature points under different viewing angles. The weight is the quantitative representation of the reliability in the matching calculation. The dynamic weight mechanism calculates the reliability of the feature points captured by each camera in real time and dynamically adjusts the weights based on these reliabilities.

[0071] Different from the traditional static weight allocation, the dynamic weight mechanism can automatically adjust according to the changes in the scene and the performance of the 3D depth camera, thereby more effectively reducing errors and improving the accuracy of 3D point cloud data.

[0072] The depth alignment mechanism uses the common scene space reference points in the multi-view RGBD image data obtained by multiple 3D depth cameras to align the three-dimensional point cloud data. This method is different from the existing depth estimation technology based on a single view. It improves the accuracy of depth estimation by integrating information from multiple viewpoints. In a multi-camera configuration, the depth alignment mechanism can ensure that the data captured by different 3D depth cameras maintains consistency in depth information, thereby achieving global optimization and significantly improving the calibration accuracy and stability of the system.

[0073] Through the 3D depth global camera module, the present invention can efficiently process large-scale three-dimensional data to subsequently achieve rapid human skeleton extraction and posture recognition, and maintain high stability and accuracy even when the user moves quickly or performs complex movements.

[0074] The 3D depth global camera module of the present invention ensures that the 3D depth global camera module can accurately capture the user's posture through high-precision three-dimensional reconstruction technology, as well as the layout and real-time calibration strategy of the 3D depth camera array.

[0075] The AI-based fitness plan recommendation module is used to plan personalized fitness routes for users, including the LightGCN-att algorithm module, dynamic adjustment module and instant feedback module.

[0076] The LightGCN-att algorithm module includes the LightGCN-att fitness plan recommendation algorithm, which includes a graph convolutional network LightGCN embedded with an attention mechanism. In the graph convolutional network LightGCN, by constructing a user-fitness activity bipartite graph, the node's neighbor information is used to update the node's embedded representation, thereby capturing the similarity between users and the correlation between fitness activities. The attention mechanism is used to assign different weights to each neighbor node, so that the graph convolutional network LightGCN can focus on the factors that have the greatest impact on the user's fitness plan.

[0077] The user-fitness activity bipartite graph is a graph structure representation in which nodes are divided into two categories: one category represents users and the other category represents fitness activities. The edges in the bipartite graph represent the historical records or behavioral relationships of users participating in a certain fitness activity. Through this graph structure, the graph convolutional network LightGCN is able to capture the interaction information between users and fitness activities as well as the similarities between users. For example, if two users participate in similar fitness activities, their nodes in the bipartite graph will be closer, and recommendations can be made based on this relationship. The graph convolutional network LightGCN is particularly suitable for processing large-scale user data and can effectively learn from users' historical activities and preferences to recommend the most suitable fitness activities and progress.

[0078] The present invention improves the traditional LightGCN algorithm. The improved LightGCN-att algorithm introduces an attention mechanism into the LightGCN algorithm, so that the attention to different features in the user's fitness data (various aspects of the user's historical fitness data, including exercise type, exercise intensity, time period, duration, physical condition, and the user's fitness goals and preferences, etc.) can be dynamically adjusted to enhance the personalized recommendation effect. The attention mechanism is a mechanism that simulates human attention in a deep learning model, which allows the model to give different degrees of attention to different parts of the input data when processing information. In the LightGCN-att algorithm module of the present invention, the attention mechanism assigns different weights to different features of the user's fitness data, thereby identifying and emphasizing the features that have the greatest impact on the fitness effect. For example, in some embodiments of the present invention, if a user shows better health indicators after aerobic exercise, the attention mechanism will increase the weight of aerobic exercise-related features. In fitness plan recommendations, adaptive adjustments can be made according to the user's physical condition and action performance. For example, when the user is tired, the system will reduce the training intensity and recommend restorative actions; when the user is progressing well, more challenging training is recommended. The attention mechanism provides each user with a scientific and dynamic fitness plan by weighting the data of users and similar nodes, and develops a fitness guidance plan that is both scientific and personalized. The LightGCN-att algorithm module cleverly combines the advantages of the lightweight graph convolutional network (LightGCN) and the attention mechanism (Attention Mechanism). By constructing a user-activity bipartite graph, it deeply explores the similarities between users and the correlation between fitness activities. This process not only takes into account the user's personal physical condition and sports preferences, but also takes into account the actual effect of fitness activities and the actual needs of users.

[0079] Based on the traditional LightGCN, the graph convolutional network LightGCN adds an attention mechanism, which enables the graph convolutional network LightGCN to pay more attention to the key information in the user's historical fitness activities, such as the user's preferred exercise type, intensity, time period, etc. This improvement not only improves the personalization of the graph convolutional network LightGCN, but also makes the generated fitness plan more in line with the actual needs of users.

[0080] As the user's training progresses, the system continuously monitors the user's training results and uses the dynamic adjustment module to optimize the training plan in real time to adapt to the user's changes. This dynamic adaptability ensures that the user can achieve the best fitness results in the shortest time while avoiding the negative effects caused by overtraining or undertraining.

[0081] The dynamic adjustment module can continuously monitor the training effect and physical reaction of the user. Through machine learning technology, the dynamic adjustment module analyzes the user's physical fitness changes in real time and dynamically adjusts the training plan according to the user's actual physical fitness. The dynamic adjustment module uses a multivariate regression model or a reinforcement learning method to dynamically adjust the training intensity, training type and training duration according to the user's current state (including actual physical fitness). For example, in some embodiments of the present invention, when it is detected that the user's training effect is improved, the intensity will be gradually increased; when the user's heart rate and fatigue are beyond the safe range, the intensity will be reduced or rest training will be recommended. In this way, the training effect and safety can be balanced. The dynamic adjustment module also obtains the user's current fitness level by comprehensively analyzing the user's training data and real-time feedback. This analysis is based on a weighted evaluation of multiple indicators, such as the user's action accuracy, cardiopulmonary endurance changes, training intensity adaptability, etc., to generate a fitness level score (such as a score or grade). Based on this fitness level score, the subsequent training plan is further optimized to ensure that the recommended content meets the user's current level and promotes the user to gradually improve.

[0082] In order to further enhance the user experience, the system also provides personalized instant feedback through an instant feedback module, which collects user feedback on the training effect to further optimize and personalize the training plan. The instant feedback module visualizes the fitness action analysis results (such as action type, whether the action is wrong, and how to correct it) obtained by the AI-based fitness action analysis module to convert the fitness action analysis results into easy-to-understand feedback information (including action accuracy, improvement suggestions and progress reports), which can be displayed to users instantly through APP or smart hardware devices based on augmented reality. Users can view the feedback information in real time through APP or smart hardware devices based on augmented reality, and can also view their own training data, and make self-adjustments or seek guidance from AI virtual coaches based on the feedback information. This closed-loop feedback not only enhances the user's sense of participation and motivation, but also helps the system to continuously optimize and adjust the training plan to achieve the best personalized guidance effect.

[0083] In the process of customizing personalized training plans, the LightGCN-att algorithm module generates a series of targeted training plans through the Adam optimizer in the LightGCN-att fitness plan recommendation algorithm based on the user's historical fitness data and current status, as well as the fitness level score obtained by the dynamic adjustment module, combined with preset fitness goals and constraints (such as time, venue, and equipment). These training plans not only include detailed information such as specific action arrangements, number of sets, repetitions, and interval time, but also take into account the physical challenges and movement difficulty that users may face, ensuring that the training plans are both challenging and feasible; continuously monitor the user's training effects, and dynamically adjust the training plans through the dynamic adjustment module to ensure that users can achieve the best fitness results in the shortest time. The dynamic adjustment module can continuously monitor the user's training effects and physical reactions.

[0084] The AI-based fitness plan recommendation module adopts the LightGCN-att fitness plan recommendation algorithm to deeply analyze user data and provide personalized training plans and suggestions. The algorithm combines lightweight graph convolutional networks and attention mechanisms to capture the similarities between users and the correlations between activities by constructing a user-activity bipartite graph. It further dynamically adjusts the training plan to adapt to user changes through a dynamic adjustment module to ensure that users achieve the best fitness results in the shortest time.

[0085] In some embodiments of the present invention, the present invention can record user fitness data through a fitness APP, including action accuracy, number of times, duration, etc.

[0086] The AI-based fitness motion analysis module adopts the Transformer-based adversarial neural network, i.e., the multimodal posture analysis algorithm for training, and uses the attention mechanism to give adaptive weights to the information of the visual modality from the RGBD image data and the physical modality from the sensor data to give full play to their advantages.

[0087] The multimodal posture analysis algorithm includes a convolutional neural network (CNN) and a Transformer model. The Transformer model includes an encoder, a decoder and a classifier. The multimodal posture analysis algorithm uses a convolutional neural network (CNN) and a Transformer model to extract and fuse features from the integrated multimodal fitness data to obtain multimodal features for posture recognition.

[0088] The RGBD image data is input into the convolutional neural network CNN to obtain the visual features extracted from the RGB color image and the three-dimensional spatial features extracted from the depth image, and then the visual features, three-dimensional spatial features and physical features are merged into multimodal features. Among them, the physical features include kinematic features, the visual features include color and texture, and the three-dimensional spatial features include human body contours and bone points. The kinematic features are extracted from the sensor data through signal processing technology. The extraction of kinematic features adopts existing mature technologies, such as data processing based on Kalman filter, frequency domain analysis method based on Fourier transform and time series decomposition method based on wavelet transform. It will not be elaborated here. The sensor data includes physiological indicator data and motion data, and the kinematic features include the changing trends of acceleration and angular velocity. The present invention uses a variety of sensors for data collection, including an accelerometer for capturing the user's linear acceleration data to reflect the speed changes of the user's movements; a gyroscope for measuring the user's angular velocity and tracking the rotation characteristics during the movement; a magnetometer for working in conjunction with the gyroscope and accelerometer to provide the user's direction information and improve the accuracy of movement recognition; and a heart rate sensor for monitoring the user's heart rate changes and evaluating the exercise intensity and the user's physiological load status.

[0089] In the Transformer model, the encoder includes multiple cascaded Transformer encoder layers, which are used to fuse and encode multimodal features to obtain a comprehensive feature representation; each encoder layer is provided with a self-attention mechanism, which is used to capture the long-distance dependencies between multimodal features. The self-attention mechanism can improve the computational efficiency and feature capture accuracy, and improve the Transformer model's ability to recognize complex postures. In order to further improve the fusion ability between multimodal features, the encoder also introduces a cross-attention mechanism, which is used to help the Transformer model capture the relevance of another modality when processing data of a certain modality, and can enhance the Transformer model's ability to understand multimodal features and the accuracy of posture recognition.

[0090] The encoder part of the Transformer model is responsible for posture recognition. The encoder processes the input multimodal features through a multi-layer self-attention mechanism and a cross-attention mechanism. After the multimodal features are processed by the encoder, a comprehensive feature representation is formed, and the user's fitness action type can be subsequently identified based on the comprehensive feature representation. The relationship between the comprehensive feature representation and the fitness action type is that the comprehensive feature representation contains the key features of the user's action, and after further processing by the classifier, it can be mapped to a specific action type label. For example, in some embodiments of the present invention, a comprehensive feature representation may correspond to fitness action types such as squats, lunges, push-ups, etc., and may also contain specific details of the action, such as the depth of the squat, whether the angle of the lunge is correct, etc. Based on these mapping relationships, the specific action type that the user is performing can be accurately identified.

[0091] The decoder is used to decode the encoded comprehensive feature representation into a natural language description or animation demonstration, which is convenient for

[0092] Generate animated demonstrations or natural language descriptions in scenarios where posture descriptions or guidance suggestions are needed.

[0093] The classifier is used to further analyze the comprehensive feature representation output by the encoder to obtain the fitness action analysis results. The classifier can be part of the Transformer model or an independent module that works with the Transformer model.

[0094] Specifically, the classifier compares the comprehensive feature representation of the current posture with the preset standard posture template to obtain the fitness action analysis result, which includes: identifying the user's fitness action type and judging whether there is an error in the fitness action. Specifically, judging whether there is an error in the fitness action is by judging the user's posture, the speed of movement, etc.

[0095] Specific error judgment and type recognition are achieved through classifiers. During the training process, the classifier learns the feature representations of different error types, such as non-standard posture, uneven force distribution, too fast or too slow speed, etc. In practical applications, the classifier outputs the fitness action type and the corresponding error probability distribution based on the comprehensive feature representation of the input, and points out the possibility of various error types. For example, if the model determines that the probability of non-standard posture is the highest, then the system will point out that the user's action has a non-standard posture error. In some embodiments of the present invention, specific error types may include: arms not straightened when pulling the butterfly machine, shrugging shoulders, etc.

[0096] In this way, the system of the present invention can provide accurate posture recognition and error analysis, helping users to improve their fitness movements and enhance fitness effects.

[0097] In some embodiments of the present invention, the present invention also applies GAN (Generative Adversarial Network) to the training process of the Transformer model to enhance the robustness of the model. Specifically, the Generative Adversarial Network includes a generator and a discriminator. The generator generates multimodal pseudo features by learning the distribution of multimodal fitness data, and the discriminator improves the recognition ability of the Transformer model for pseudo data by comparing multimodal features with multimodal pseudo features. The generator optimizes the convergence speed and generation quality, and the discriminator ensures the stability and generalization ability of the Generative Adversarial Network GAN in complex scenarios, thereby significantly improving the robustness of the posture recognition model in real environments. By combining the Transformer model with the Generative Adversarial Network (GAN), the advantages of both are effectively utilized to improve the accuracy and robustness of posture recognition.

[0098] In order to improve the accuracy and robustness of posture analysis, the present invention adopts a Transformer-based adversarial neural network, combined with a nine-axis sensor (accelerometer, gyroscope, magnetometer, in some embodiments of the present invention, the nine-axis sensor uses an MPU9250 sensor) and multimodal fitness data captured by a 3D depth camera, and a Kalman filter is used for time synchronization and noise filtering of the data to ensure the consistency and accuracy of the data. The application of self-attention mechanism and cross-attention mechanism enables the Transformer model to focus on the dependencies within the input sequence and fuse visual and motion data across modalities, thereby improving the ability to recognize incorrect postures. The adversarial training strategy further enhances the robustness of the model, enabling it to accurately recognize postures even under occlusion and complex backgrounds.

[0099] The AI-based fitness action analysis module transmits the output fitness action analysis results to the augmented reality-based smart hardware device or the equipped device, and provides feedback to the user on posture problems through screen display, voice prompts, etc. At the same time, the fitness action analysis results (such as whether the user's fitness action is wrong, and if so, how to improve it) can provide posture data for the AI-based virtual coach to generate personalized guidance suggestions.

[0100] The multimodal posture analysis algorithm processes the input multimodal fitness data (including RGBD image data, physiological indicators, and motion data) through the Transformer model, and shows a strong ability to capture spatiotemporal dependencies in sequence modeling. With the help of the multi-layer self-attention mechanism, the Transformer model can achieve a detailed analysis of the dynamic changes of human posture and can more accurately identify different posture features, especially in fast-changing and nonlinear movements. At the same time, the introduction of the adversarial learning mechanism in the generative adversarial network (GAN) further enhances the performance of the model in complex environments. The generative adversarial network improves the sample generation quality of the model through the competition between the generator and the discriminator: the generator is responsible for generating realistic posture samples with occlusions, simulating the occlusions and complex backgrounds that users may encounter in actual scenes; the discriminator is responsible for distinguishing the generated samples from the real samples and providing feedback on the quality of the samples generated by the generator. The generator and the discriminator compete with each other during the training process. Under the supervision of the discriminator, the generator continuously generates more realistic occluded samples, gradually improving the performance of the model in processing occluded data. In this way, the generator continuously generates more realistic and indistinguishable occluded posture samples, making the Transformer model robust in complex backgrounds and occluded conditions. Through this deep fusion, the multimodal posture analysis algorithm can not only significantly improve the accuracy of posture recognition, but also enable the Transformer model to more flexibly adapt to the diverse needs in fitness scenarios, expanding the application of the Transformer model in the fitness field. For example, when the user's movements are partially obscured by equipment or other objects, the Transformer model can still accurately identify the user's movement posture, thereby providing users with precise feedback and guidance. This multimodal posture analysis method based on Transformer's spatiotemporal modeling and GAN-generated sample enhancement can be more practical and adaptable in the fitness environment.

[0101] In addition to traditional video input, the multimodal posture analysis algorithm also integrates multimodal information such as physical feature data. Through multimodal data fusion technology, the present invention can efficiently integrate data from different sources and comprehensively utilize the complementary advantages of various sensor data. This multimodal fusion not only improves the accuracy of posture recognition, but also enables the algorithm to more comprehensively capture the user's motion state, providing richer data support for subsequent posture evaluation and feedback.

[0102] The multimodal posture analysis algorithm not only has powerful posture recognition capabilities, but also has advanced posture evaluation functions. Through quantitative evaluation of user posture, including the standardization of posture, speed of movement, force distribution, etc., it can provide users with instant feedback and suggestions, which helps users to adjust their movements in time and improve fitness effects. It also provides an important basis for the formulation of personalized fitness plans.

[0103] The fitness action demonstration module based on 3DMR imaging is used to project a virtual 3D fitness coach and a virtual character with a body shape similar to that of the user so as to demonstrate the correct actions through the virtual 3D fitness coach and reflect the user's actions in real time through the virtual character.

[0104] The fitness action demonstration module based on 3DMR imaging includes an intelligent hardware device based on augmented reality, a posture estimation network and a skeleton mapping model, and the posture estimation network and the skeleton mapping model are both installed in a server. In some embodiments of the present invention, the intelligent hardware device based on augmented reality can be any one of AR glasses, HoloLens One MR glasses, and a helmet-mounted display.

[0105] The fitness action demonstration module based on 3DMR imaging provides users with a three-dimensional dynamic virtual 3D fitness coach through advanced image processing and augmented reality technology, and projects the virtual 3D fitness coach into an intelligent hardware device based on augmented reality; a complete fitness action demonstration animation is generated through a preset character model combined with a deep learning algorithm, and then the fitness action demonstration animation is processed using the mixed reality light field rendering and ray tracing technology based on Unity, so that the performance of the virtual 3D fitness coach under different lighting and viewing angles is more natural and realistic, so as to enhance the realism of the user experience.

[0106] The fitness action demonstration module based on 3DMR imaging can perform 3D projection in smart hardware devices based on augmented reality, such as MR glasses, and project a virtual 3D fitness coach in front of the user's field of vision. The virtual 3D fitness coach will show the correct demonstration of the current action. Users can watch from multiple angles, so as to learn the correct fitness posture more intuitively.

[0107] The animation of the virtual 3D fitness coach is rendered in real time into the user's field of view, and can interact with the user through voice, gestures, etc. At the same time, the guidance actions of the virtual 3D fitness coach can be dynamically adjusted according to the user's feedback and action performance to achieve a personalized teaching experience.

[0108] The present invention generates the appearance and action sequence of a virtual 3D fitness coach to make it both professional and attractive. Then, the animation of the virtual 3D fitness coach is presented in real time and vividly in the user's field of vision using graphics rendering. This process can be achieved through intelligent hardware devices based on augmented reality such as MR head displays, bringing an immersive interactive experience to users.

[0109] In some embodiments of the present invention, during the action demonstration, users can watch the dynamic demonstration of the virtual 3D fitness coach by wearing MR glasses or a helmet-type display. The multi-angle surround viewing function is supported, and users can change the viewing angle and distance by moving their heads or operating the handles, so as to more comprehensively understand and learn the correct fitness posture. Interaction is also possible, allowing users to pause, replay, slow down, etc. during the viewing process, so as to better grasp the details and key points of the action.

[0110] In the user's field of vision, in addition to the real scene and the virtual 3D fitness coach, the present invention uses a 3D depth camera to obtain depth information that a 2D camera cannot obtain, and by applying a posture estimation network (Orbbec SDK algorithm), generates skeleton data of the user consisting of 33 three-dimensional skeleton coordinate points from a relatively complex background, and then obtains a standardized skeleton model through a skeleton mapping model, and constructs a virtual character based on the standardized skeleton model.

[0111] The user's skeleton data is input into the skeleton mapping model, and the skeleton is migrated through the skeleton mapping model to generate a standardized skeleton model. The process of mapping the skeleton mapping model to generate a standardized skeleton model includes:

[0112] ① Analyze and preprocess the user’s skeleton data.

[0113] The user's skeleton data includes the three-dimensional coordinates of the joints and the connection relationships between these joints. Joints are abstract representations of key parts of the human body, such as shoulders, elbows, or knees, while connection relationships describe the anatomical relationships between these joints. In the parsing stage, the skeleton mapping model extracts the coordinate positions of all joints and the connection information between the joints to form the initial representation of the user's skeleton. In the preprocessing stage, in order to eliminate possible jitter or noise in the joint coordinates, filtering methods (such as Kalman filtering) are used to smooth the user's skeleton data to ensure the accuracy and stability of subsequent calculations.

[0114] ② Determine the root node in the user's skeleton data through the RootNet module.

[0115] The RootNet module calculates the joint point location that is most likely to be the root node by analyzing the spatial distribution of all relevant nodes in the user's skeleton data.

[0116] After the initial analysis and preprocessing of the data, the root node of the user skeleton is determined with the help of the RootNet module. The root node is the reference point of the entire skeleton and can be selected as the position of the human pelvis or hip joint, because these parts play the role of a stable center in human posture. The determination of the root node provides a reliable starting point and a stable reference system for subsequent skeleton migration.

[0117] ③ Use the GMEdgeNet module to predict the connection relationship between the joint points in the user's skeleton data and obtain the set of the user's skeleton connections.

[0118] The GMEdgeNet module is a deep learning module designed for skeleton line inference. It is used to automatically generate a judgment result on whether there is a line between every two joint points based on the coordinate information of the joint points, and output a set of skeleton lines, which describes the structural topology of the user's skeleton. These lines represent the actual connection relationship of human bones, such as the skeleton line between the upper arm and the forearm.

[0119] ④ Optimize the set of user's bone connections through the minimum spanning tree (MST) algorithm to obtain the preliminary structure of the skeleton.

[0120] After obtaining the connection information between the root node and the joint points, the minimum spanning tree (MST) algorithm is used to optimize the set of user's skeleton connections.

[0121] The minimum spanning tree (MST) is a graph algorithm used to find a subgraph with the minimum total edge length and guaranteed connectivity in a graph composed of nodes and edges. In skeleton migration, the minimum spanning tree (MST) algorithm is used to ensure that the joints in the user's skeleton connection set are connected, while reducing redundant connections and optimizing the efficiency and accuracy of the model. In the specific implementation process, the minimum spanning tree algorithm is optimized through the Prim algorithm or the Kruskal algorithm. The Prim algorithm is a node-based incremental algorithm used to process dense graphs. It starts from a starting joint point (usually the root node) and gradually expands to adjacent nodes. Each time, the edge with the smallest weight is selected, and the unconnected nodes are added to the tree to ensure that each step maintains the optimality of local connectivity. The Kruskal algorithm is an edge-based greedy algorithm used to process sparse graphs. By sorting all edges by weight from small to large, the edges that do not form a ring are gradually added to the tree until all nodes are connected. The two algorithms focus on local optimization and global optimization respectively in the process of constructing the minimum spanning tree. These two implementation methods support the core goal of the minimum spanning tree algorithm, which generates a connected and efficient skeleton structure by reducing redundant connections and optimizing paths, providing an optimized preliminary skeleton structure for subsequent standard skeleton mapping.

[0122] RootNet is a network module used to determine the root node of the human skeleton (usually the pelvis or hip joint). The root node is the basis for building a skeleton migration model. It provides a stable reference point for the entire skeleton. Other outputs are used as the starting point for building a minimum spanning tree to ensure the accuracy and stability of skeleton migration; GMEdgeNet is a network module used to predict the connection relationship between the joints of the human skeleton. It can infer the bone connection between them based on the position information of the joints. It provides edge information between the joints. These edges represent the connection relationship between the bones and are the key to building a minimum spanning tree. The minimum spanning tree algorithm ensures the connectivity of the skeleton model and minimizes the total edge length, thereby providing an accurate and efficient skeleton representation.

[0123] ⑤Align the preliminary structure of the skeleton with the standard skeleton template to obtain a standardized skeleton model.

[0124] The standard skeleton template is a predefined skeleton model, which includes the joint positions and connection relationships of a standard human skeleton.

[0125] The steps to align the preliminary structure of the skeleton with the standard skeleton template include:

[0126] a: In order to align the preliminary skeleton structure and the standard skeleton template, the scale factor of the user skeleton is first calculated.

[0127] The scaling factor is determined by comparing the length ratios between the preliminary structure of the skeleton and key bones (eg, the length from the head to the hip joint) in a standard skeleton template.

[0128] b: Based on the scale factor, the joint point coordinates of the preliminary structure of the skeleton will be scaled proportionally so that its overall size is consistent with the standard skeleton.

[0129] c: Subsequently, the position and orientation of the preliminary structure of the skeleton are further adjusted through linear transformation methods, including translation, rotation and scaling, so that it is fully aligned to the coordinate system of the standard skeleton.

[0130] d: After alignment, the connection relationship of the joints needs to be optimized. This step recalculates the length of the connecting edge between each two joints and compares it with the connection structure of the standard skeleton template to ensure that the output skeleton conforms to the topological rules of the standard skeleton template and obtain a standardized skeleton model. For example, if the length of some edges is too long or there are non-connected nodes, these connections are adjusted to ensure the integrity and rationality of the skeleton.

[0131] After the above process, the skeleton mapping model generates a standardized skeleton model, which contains the adjusted joint point positions and optimized connection relationships. This result can successfully map the user's random skeleton data to the standard skeleton structure, providing a reliable basis for subsequent posture analysis and action recognition. Throughout the process, by gradually processing the concepts of joint points, root nodes, connections and proportions, an efficient conversion from random skeleton to standard skeleton is achieved.

[0132] By using the skeleton migration technology based on RootNet and GMEdgeNet modules combined with the minimum spanning tree (MST) algorithm, the present invention constructs a virtual character with a similar body shape to the user, through which the user's movements can be reflected in real time. When a user's posture does not meet the standard, the body part that does not meet the standard will be displayed in the form of a red dot (or other forms) on the corresponding body part of the virtual character as a warning, thereby reminding the user to correct the movement.

[0133] The process of building a virtual character through the output of the skeleton mapping model, that is, the standardized skeleton model, includes the following steps:

[0134] 1. Use a 3D modeling method based on mesh generation to add appearance information to the standardized skeleton model. First, use triangular mesh technology to wrap the surface of the standardized skeleton model to generate a polygonal surface that fits the human body shape. Then, use the subdivision surface algorithm to smooth the mesh surface to ensure the delicacy and continuity of the virtual character's appearance.

[0135] 2. Based on texture mapping technology, the skin, clothing and other visual elements selected by the user are mapped to the surface of the standardized skeleton model, thereby realizing the personalized presentation of the virtual character.

[0136] 3. In order to enhance the dynamic performance of the virtual character, the present invention adopts skeletal animation technology, which enables the appearance of the virtual character to change in real time with the movement of the skeleton by binding the appearance grid of the virtual character to the standardized skeleton model. This technology can accurately capture the dynamic posture of the user and reproduce it in the virtual character with high fidelity.

[0137] 4. Combined with Unity-based real-time rendering technology, virtual characters can present realistic and natural visual effects under various lighting conditions and viewing angles.

[0138] The skeleton and key points of the human body are obtained with the help of RGBD image data, and then based on the data results obtained by the posture analysis algorithm, the present invention uses the skeleton mapping model as a visualization presentation for subsequent interactive display, so as to reflect the user's movements in real time.

[0139] The present invention uses the depth camera in the 3D depth global camera module to accurately capture the user's skeleton node data through the posture estimation network. And the skeleton mapping model is used to perform detailed preprocessing on these raw data, including removing noise interference, applying filtering technology to smooth data, and optimizing the accuracy of joint point positioning, thereby significantly improving the reliability and accuracy of the data. After the multimodal posture analysis process, the specific problems of the user's fitness movements can be obtained.

[0140] By adopting the high-precision depth perception capability of the depth camera Astra+ and the Orbbec SDK algorithm, accurate 3D reconstruction of the human skeleton nodes can be achieved. The arrangement of multiple 3D depth cameras can ensure global field of view coverage without blind spots, improving the integrity and accuracy of the data. This high-precision 3D reconstruction technology provides a solid foundation for posture assessment and motion analysis. Through the Orbbec SDK algorithm, the 3D depth global camera module can achieve low-latency real-time processing while ensuring high-precision reconstruction. This enables users to obtain instant feedback on their posture during exercise and adjust their movements in time, which not only improves the user experience, but also ensures the timeliness and effectiveness of fitness guidance.

[0141] The present invention deeply analyzes the results of fitness action analysis and can accurately identify possible problems in user actions, such as inaccurate posture, unsmooth movements, etc. Based on these problems, the present invention can provide a variety of guidance according to the user's specific posture problems and current physical state, including voice prompts and virtual coach guidance in the MR head display. Not only does it provide targeted adjustment suggestions, but it also helps users intuitively understand and correct mistakes by showing correct action demonstrations.

[0142] The AI-based virtual coach interaction module is used to communicate and guide users in real time based on a large fitness knowledge model, and to provide users with scientific fitness guidance.

[0143] The AI-based virtual coach interaction module obtains a large model of fitness knowledge by training a large language model.

[0144] When training the large language model, in terms of the construction of the knowledge base in the field of fitness, the present invention fine-tunes the large language model, combines professional knowledge websites in the field of fitness, including Darebee and bodybuilding, and constructs a large question-and-answer model covering a wide range of fitness knowledge. The system can accurately understand the user's natural language questions and provide scientific fitness guidance and answers. This knowledge base construction method not only improves the professionalism of the virtual coach, but also enhances the user experience. In terms of situational awareness and intelligent recommendation functions, the system has situational awareness capabilities and can provide more accurate fitness suggestions based on the user's current exercise state and environmental conditions. For example, indoor exercise plans are recommended on rainy days, or suggestions for relaxation and stretching are provided when the user feels tired. This situational awareness and intelligent recommendation mechanism makes the guidance of the virtual coach more intimate and practical. Through natural language processing technology, the system has emotional interaction and personalized communication capabilities, can identify the user's emotional state, and interact with the user in a more humane and encouraging language. This emotional interaction not only enhances the user's sense of participation and motivation, but also makes the guidance of the virtual coach more vivid and interesting. At the same time, the system also supports personalized communication functions, which can adjust the communication method according to the user's preferences and habits, and provide users with a more personalized service experience.

[0145] In some embodiments of the present invention, fine-tuning of the large language model includes the following aspects: First, by refining the professional knowledge in the fitness field, a large amount of structured and unstructured data from professional knowledge websites (such as Darebee and Bodybuilding) is collected, and these data are cleaned, labeled and classified to ensure their accuracy and reliability. Secondly, in the training of the large language model, question-answer pairs and contextualized interaction cases in the field of fitness are added to enable the large language model to understand and answer fitness-related questions, such as training plan formulation, nutritional advice, and sports injury prevention. Finally, the domain knowledge is transferred to the weight of the large language model through the existing knowledge distillation technology, thereby improving the large language model's understanding and reasoning ability of fitness professional issues.

[0146] In some embodiments of the present invention, the large language model of Wenxinyiyan 4.0 Turbo may be used, and in other embodiments, other large language models may also be used.

[0147] The trained fitness knowledge model has powerful natural language processing and context understanding capabilities. It can generate personalized guidance and suggestions based on user questions and feedback, and convert text into natural and fluent speech output through speech synthesis technology. In addition, the fitness knowledge model also integrates an emotional computing module, which analyzes the user's emotional changes by monitoring the user's voice and intonation characteristics in real time, and adjusts the interaction method, such as communicating with encouraging language when the user is tired, and providing more challenging guidance when the user is excited, so as to enhance the user experience. The present invention is combined with an AI virtual coach based on a large language model, that is, the guidance of the virtual 3D fitness coach AI kernel. The AI-based virtual coach interaction module can provide users with accurate and visual guidance, thereby improving the scientificity and effectiveness of fitness.

[0148] Users can interact with virtual 3D fitness coaches through voice, gestures and other methods. The AI-based virtual coach interaction module collects information through 3D depth cameras and nine-axis sensors based on the user's real-time feedback and action performance to input into the AI-based fitness action analysis module. The AI-based fitness action analysis module evaluates the user's fitness posture adjustment. If the user has adjusted to the correct fitness posture, the user will be given positive feedback; if the user's fitness posture still has problems, continue to guide to ensure that each user can get a customized, efficient and accurate personalized teaching experience. This two-way interactive mechanism not only enhances the user's sense of participation and satisfaction, but also helps to continuously optimize the teaching effect and achieve a virtuous cycle of teaching and learning.

[0149] In the fitness guidance system, the user's historical data is stored in a database. The main function of the database is to record the user's posture data, action performance, fitness plan and feedback information during training. The content stored in the database includes but is not limited to the user's fitness posture data, action accuracy, training times, duration, fitness goals, physiological indicators (such as heart rate), and analysis results and improvement suggestions generated by the system.

[0150] In some embodiments of the present invention, a fitness guidance system based on AI posture analysis and MR human-computer interaction includes a fitness action demonstration module based on 3DMR imaging, a 3D depth global camera module, an AI-based fitness plan recommendation module, an AI-based fitness action analysis module, and an AI-based virtual coach interaction module. The fitness action demonstration module based on 3DMR imaging includes an intelligent hardware device based on augmented reality, a posture estimation network, and a skeleton mapping model. This fitness guidance system is directly mounted in the intelligent hardware device based on augmented reality, and the system has a display function through the intelligent hardware device based on augmented reality. In another embodiment of the present invention, a fitness guidance system based on AI posture analysis and MR human-computer interaction includes a 3D depth global camera module, an AI-based fitness plan recommendation module, an AI-based fitness action analysis module, and an AI-based virtual coach interaction module. This fitness guidance system is mounted on a mounted device, and the system does not have a display function. The mounted device can use any electronic product that can be mounted, such as any electronic product in a smart phone, a tablet computer, or a smart TV as a mounted device.

[0151] In some embodiments of the present invention, the fitness guidance system can be configured in any intelligent hardware device based on augmented reality, an electronic product with mounting conditions, or any AI fitness system.

[0152] In some embodiments of the present invention, a fitness guidance method based on AI posture analysis and MR human-computer interaction is provided, comprising the following steps:

[0153] Step 1: The AI-based fitness plan recommendation module personalizes the fitness plan according to the user's multimodal fitness data.

[0154] In the fitness guidance system, the formulation and recommendation of personalized fitness guidance plans are key steps to ensure that users get the best fitness experience. This process is closely based on the data collected by the 3D depth global camera module and the data collected by the sensor, including visual and depth information obtained through the RGBD image set, user physiological indicator data (such as heart rate) collected by the sensor, and motion data provided by the sensor (such as MPU9250). The multimodal posture analysis algorithm deeply processes these multi-source data to accurately identify and evaluate the user's fitness movements, providing a solid foundation for the formulation of personalized training plans.

[0155] During initial use, when the guidance system has not yet accumulated the user's historical fitness data, the fitness plan is generated based on the fitness direction manually selected by the user. When the user registers or uses the system for the first time, the system allows the user to select their fitness goals (such as fat loss, muscle gain, training direction), daily training time, and desired training intensity through an interactive questionnaire; based on the information entered by the user, the system matches the most suitable initial fitness plan from the built-in general fitness template library, and provides detailed action arrangements, sets, repetitions, and training suggestions as a starting point. At the same time, the system collects the user's basic exercise data (such as action accuracy) through the first fitness test to further optimize the subsequent training plan.

[0156] Step 2: The user performs fitness according to the fitness plan. During the fitness process, the user's physiological indicator data and motion data are collected through sensors, and the user's multimodal fitness data is collected in real time through a 3D depth global camera module, and the multimodal fitness data is transmitted to the server.

[0157] In some embodiments of the present invention, during the data collection process, each RGBD camera, i.e., the 3D depth camera device Astra+, is first started, and parameters such as exposure time, frame rate, resolution, etc. are set first to ensure that the RGBD camera is in the best working state; then, the RGBD camera is used to capture the RGB color image and depth image of the fitness scene in real time to form RGBD image data, which is output in the form of a video stream and read and processed through a software interface; then the captured RGBD image data is stored in a local or cloud server. The multi-view RGBD image data captured by multiple RGBD cameras will obtain high-precision three-dimensional point cloud data through a multi-eye stereo matching module. These three-dimensional point cloud data can support subsequent posture recognition algorithms, and will be subsequently integrated with the user's physiological index data and motion data collected by the sensor for the characterization of the user's fitness mode. Among them, in the RGBD image data, the RGB color image will be used to extract visual features such as color and texture, and the depth image will be used to extract three-dimensional spatial features such as human body contours and bone points.

[0158] After collecting multimodal fitness data, the multi-source data is integrated. Multimodal fitness data includes RGBD image data and sensor data. Among them, RGBD image data is a real-time video stream captured by RGBD cameras during the user's fitness process. These RGBD cameras not only record color images, but also obtain depth information through depth sensors, providing a rich data foundation for subsequent posture analysis; sensor data includes physiological indicator data (such as heart rate) and motion data (such as acceleration, angular velocity) of users collected through wearable sensor devices. Physiological indicator data, motion data and RGBD image data complement each other to jointly build a comprehensive portrait of user fitness behavior.

[0159] In some embodiments of the present invention, the multimodal fitness data is synchronized to ensure that the data collected by all data sources are aligned on the time axis for subsequent multimodal fusion analysis; and the obtained multimodal fitness data is cleaned: the collected original multimodal fitness data is preprocessed by denoising, filtering, missing value filling and other operations to improve the quality of the multimodal fitness data. Under the premise of ensuring data integrity, a large amount of data is compressed and stored to have high-density and large-capacity storage capabilities, and user privacy can be more effectively protected through encryption technology.

[0160] The processed multimodal fitness data is used as the input of the AI-based fitness action analysis module to improve the accuracy and real-time performance of posture recognition. The user's fitness data (such as action accuracy, number of times, duration, etc.) will also serve as the analysis basis of the LightGCN-att algorithm module to generate personalized training plans.

[0161] Step 3: The AI-based fitness action analysis module identifies and analyzes the user's fitness actions in real time based on multimodal fitness data, performs posture recognition and posture evaluation, and obtains fitness action analysis results, wherein the fitness action analysis results include posture recognition results and posture evaluation results. The posture recognition results include the identified fitness action type, and the posture evaluation results include whether there are any errors in the user's fitness posture.

[0162] Step 4: The fitness action demonstration module based on 3DMR imaging performs 3D projection in an augmented reality-based smart hardware device such as an MR headset, and projects a virtual 3D fitness coach and a virtual character in front of the user's field of vision. The user makes correct fitness movements according to the action demonstration of the virtual 3D fitness coach, and the virtual character reflects the user's movements in real time.

[0163] Step 5: The AI-based virtual coach interaction module communicates and guides users in real time based on the fitness knowledge model, providing scientific fitness guidance.

[0164] In some embodiments of the present invention, the fitness guidance system on which the method is based does not have a fitness action demonstration module based on 3DMR imaging, the steps of the fitness guidance method do not have a display demonstration function, and the AI-based virtual coach interaction module does not perform display interaction through a virtual 3D fitness coach, but only voice interaction.

[0165] In a specific application example, when using the fitness guidance system provided by the present invention to perform fitness, the following steps are included:

[0166] Step 1: The user enters the gym and rents an MR headset and sensors.

[0167] Step 2: The user enters the fitness app and selects the corresponding fitness program based on the fitness content recommended by the fitness guidance system.

[0168] Step 3: The user starts exercising on the corresponding fitness equipment according to the fitness instructions in the headset.

[0169] Step 4: The 3D depth global camera module and sensor begin to acquire the user's fitness posture, obtain multimodal fitness data, and upload it to the server for multimodal posture analysis. The obtained fitness movement analysis results will be continuously transmitted to the MR headset.

[0170] Step 5: The AI-based virtual coach in the MR headset corrects and guides the user's posture according to the results of fitness movement analysis, including voice prompts and demonstrations of correct movements.

[0171] Step 6: Users can ask questions to the AI-based virtual coach at any time while exercising, and the AI-based virtual coach will give corresponding replies based on the fitness knowledge model.

[0172] Step 7: After completing the fitness, users can check their fitness status through the fitness app on their mobile phones.

[0173] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A fitness guidance system based on AI posture analysis and MR human-computer interaction, characterized in that: The system specifically comprises: 3D depth global camera module, used to capture RGBD image data of users during exercise and transmit it to the server; The AI-based fitness plan recommendation module is used to plan personalized fitness routes for users, including the LightGCN-att algorithm module, the dynamic adjustment module and the instant feedback module. The LightGCN-att algorithm module generates a training plan based on the user's historical fitness data, current status and fitness level score, combined with the preset fitness goals; the dynamic adjustment module is used to analyze the user's physical fitness changes in real time and dynamically adjust the training plan according to the user's actual physical fitness; the instant feedback module is used to collect user feedback on the training effect; An AI-based fitness action analysis module is used to identify and analyze the user's fitness actions based on the user's multimodal fitness data to obtain the user's fitness action analysis results; An AI-based virtual coach interaction module is used to communicate and guide users in real time based on a large fitness knowledge model; The LightGCN-att algorithm module includes a LightGCN-att fitness plan recommendation algorithm, which includes a graph convolutional network LightGCN embedded with an attention mechanism. In the graph convolutional network LightGCN, a user-fitness activity bipartite graph is constructed and the node's neighbor information is used to update the node's embedded representation, thereby capturing the similarity between users and the correlation between fitness activities, and dynamically adjusting the attention to different features in the user's fitness data through the attention mechanism; The AI-based fitness action analysis module adopts a multimodal posture analysis algorithm to identify fitness actions; the multimodal posture analysis algorithm includes a convolutional neural network CNN and a Transformer model, and the Transformer model includes an encoder, a decoder and a classifier, or the multimodal posture analysis algorithm includes a convolutional neural network CNN, a Transformer model and a classifier, and the Transformer model includes an encoder and a decoder; The convolutional neural network (CNN) is used to extract features from RGBD image data to obtain visual features and three-dimensional spatial features, and to combine the visual features, three-dimensional spatial features and physical features into multimodal features, wherein the physical features include kinematic features, which are extracted from physiological index data and motion data; The encoder is used to fuse and encode multimodal features to obtain a comprehensive feature representation, which contains the key features of the user action; The decoder is used to decode the encoded comprehensive feature representation into a natural language description or animation demonstration; The classifier is used to analyze the comprehensive feature representation output by the encoder, and obtain the fitness action analysis result by comparing the comprehensive feature representation of the current posture with the preset standard posture template; The system also includes a fitness action demonstration module based on 3DMR imaging, which is used to project a virtual 3D fitness coach and a virtual character with a body shape similar to that of the user so that the virtual 3D fitness coach can demonstrate the correct actions and the virtual character can reflect the user's actions in real time.

2. The fitness guidance system based on AI posture analysis and MR human-computer interaction according to claim 1, characterized in that: In the 3D depth global camera module, the multi-view stereo matching module is used to match the common feature points in the multi-view RGBD image data collected by multiple 3D depth cameras, and generate high-precision 3D point cloud data; the matching compensation module based on deep learning is used to correct the matching errors, fill in the data gaps or repair the wrong feature points during the matching process; Dynamic calibration between multiple 3D depth cameras is performed through a real-time calibration module.

3. The fitness guidance system based on AI posture analysis and MR human-computer interaction according to claim 2, characterized in that: The multi-view stereo matching module uses an improved multi-view stereo matching algorithm to accurately match common feature points in multi-view RGBD image data collected by multiple 3D depth cameras to obtain preliminary matching results. The improvements made by the improved multi-view stereo matching algorithm include: introducing a multi-scale feature fusion mechanism, extracting the approximate outline and shape features of the image at low resolution to obtain low-resolution features, capturing more subtle texture and edge features at high resolution to obtain high-resolution features, and fusing the low-resolution features with the high-resolution features; optimizing the cost aggregation strategy to increase the weight of clear and reliable feature points.

4. The fitness guidance system based on AI posture analysis and MR human-computer interaction according to claim 2, characterized in that: The real-time calibration module adopts an adaptive calibration algorithm to accurately match the feature points in the overlapping fields of view between 3D depth cameras, eliminates optical distortion and angle errors through a dynamic weight mechanism, and performs global optimization of three-dimensional point cloud data through a depth alignment mechanism; the dynamic weight mechanism refers to real-time adjustment of the weight distribution of feature points in the matching process between different 3D depth cameras according to the reliability of the feature points, and the depth alignment mechanism uses the common scene space reference points in the multi-view RGBD image data obtained by multiple 3D depth cameras to align the three-dimensional point cloud data.

5. The fitness guidance system based on AI posture analysis and MR human-computer interaction according to claim 1, characterized in that: The generative adversarial network is used in the training of the Transformer model. The generative adversarial network includes a generator and a discriminator. The generator generates multimodal pseudo features by learning the distribution of multimodal fitness data. The discriminator is used to compare multimodal features with multimodal pseudo features to improve the recognition ability of pseudo data.

6. The fitness guidance system based on AI posture analysis and MR human-computer interaction according to claim 1, characterized in that: The fitness action demonstration module based on 3DMR imaging includes an intelligent hardware device based on augmented reality, a posture estimation network and a skeleton mapping model. The intelligent hardware device based on augmented reality is used to project the virtual 3D fitness coach and the virtual character; the posture estimation network is used to generate the user's skeleton data, and the user's skeleton data includes the three-dimensional coordinates of the joint points and the connection relationship between the joint points; the skeleton mapping model generates a standardized skeleton model based on the user's skeleton data; and the virtual character is constructed based on the standardized skeleton model; The process of generating a standardized skeleton model by mapping the skeleton mapping model includes: Parse and preprocess the user's skeleton data; Determine the root node in the user's skeleton data; Predict the connection relationship between the joint points in the user's skeleton data to obtain a set of the user's skeleton connection lines; Optimize the user's bone connection set to obtain the preliminary structure of the skeleton; Align the preliminary structure of the skeleton with the standard skeleton template to obtain a standardized skeleton model; The process of constructing a virtual character based on the standardized skeleton model comprises the following steps: Use a 3D modeling method based on mesh generation to add appearance information to the standardized skeleton model; Based on texture mapping technology, add skin, clothing and other visual elements to the surface of standardized skeleton models; Bind the appearance mesh of the avatar to the standardized skeleton model; Rendering of virtual characters.

7. A fitness guidance method based on AI posture analysis and MR human-computer interaction, characterized in that: The method is implemented by using the system according to any one of claims 1 to 6, comprising the following steps: The AI-based fitness plan recommendation module personalizes fitness plans based on the user's multimodal fitness data; The user exercises according to the fitness plan, and the user's physiological index data and motion data are collected during the fitness process. The user's RGBD image data is collected in real time through the 3D depth global camera module, and the multi-modal fitness data is transmitted to the server; The AI-based fitness action analysis module recognizes and analyzes the user's fitness actions in real time based on multimodal fitness data, performs posture recognition and posture evaluation, and obtains fitness action analysis results; The AI-based virtual coach interaction module communicates and guides users in real time based on a large fitness knowledge model; When the system also includes a fitness action demonstration module based on 3DMR imaging, the fitness action demonstration module based on 3DMR imaging performs 3D projection in an intelligent hardware device based on augmented reality, projects a virtual 3D fitness coach and a virtual character, and the user performs corresponding fitness actions according to the action demonstration of the virtual 3D fitness coach, and the virtual character reflects the user's action status in real time.

Citation Information

Patent Citations

  • Fitness coaching system using personalized augmented reality technology

    KR101970687B1

  • AI-based body posture analysis assistant system and transmission method

    CN110600125A

  • Construction method of fitness training evaluation decision model based on VR virtual digital human

    CN119027558A

Cited By

  • Multi-gun-type adaptive AI interactive light weapon training method

    CN122388596A