Eight-section brocade training system and method based on digital human technology

By employing high-precision optical motion capture, Kalman filter tracking, and BERT intelligent question answering, combined with environmental perception and adaptive adjustment, the bottlenecks in motion generation fidelity, robustness of open environment interaction, and module synergy in existing technologies have been overcome, achieving high-precision motion reproduction, stable interaction, and personalized rehabilitation guidance.

CN121237308APending Publication Date: 2025-12-30ZHONGKE XINHE (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511407026.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies have bottlenecks in terms of motion generation fidelity, robustness of open environment interaction, and functional module synergy, and cannot meet the needs of rehabilitation training, especially for addicts, for high-precision motion reproduction, stable human-computer interaction, and personalized feedback.

Method used

Employing high-precision optical motion capture and human anatomical data-driven 3D modeling, Kalman filter target tracking, multimodal signal anti-interference processing, and a BERT-based intelligent question-answering module, combined with open environment perception and adaptive adjustment, it achieves digital human motion reproduction, stable interaction, and personalized guidance.

Benefits of technology

It achieves high-precision digital human motion reproduction, stable human-computer interaction in an open environment, provides personalized rehabilitation guidance, and improves user experience and training effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237308A_ABST
    Figure CN121237308A_ABST
Patent Text Reader

Abstract

The invention discloses an eight-segment brocade training system and method based on a digital human technology, and mainly relates to the technical field of artificial intelligence interaction. Comprising an eight-section brocade digital human striping module which is used for generating a high-simulation digital human model through motion capture and a 3D modeling technology; the intelligent question and answer module is used for performing semantic understanding and answer generation based on a natural language processing technology and a knowledge base; the digital human all-in-one machine integrates face capturing and character locking functions and is used for identifying and tracking a user in an open environment; the open environment interaction module is used for sensing environment interference and carrying out anti-interference processing to guarantee stable operation of the system in a complex environment. The system has the beneficial effects that high-precision digital human action reproduction, stable human-computer interaction in an open environment and intelligent question answering and training closed loop are realized, and customized guidance is specially provided for rehabilitation requirements of addicts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer intelligent interaction technology, specifically to an Eight-Section Brocade training system and method based on digital human technology. Background Technology

[0002] With the application of virtual digital human technology in the field of sports and fitness, existing technology can already demonstrate basic Baduanjin movements. However, when applied to rehabilitation training (especially for special groups such as addicts) where high accuracy and real-time interaction are required, three major insurmountable technical bottlenecks have been exposed: 1. Bottlenecks in motion generation fidelity and standardization: Existing technologies mostly employ keyframe animation or low-precision inertial motion capture, resulting in digital human movements that lack data support based on real human kinematics. This leads to significant deviations in joint rotation angles, motion trajectories, and rhythms from the standard Baduanjin exercise, failing to meet the core requirement of "accurate motion reproduction" in rehabilitation training. The root cause lies in the failure of the technological approach to deeply couple high-precision optical capture, human anatomical models, and 3D data-driven rendering. 2. Bottlenecks in robustness of interaction in open environments: Existing interactive systems rely on fixed background, lighting, and sound field environments. In real open scenarios (such as rehabilitation activity rooms or outdoors), dynamic interference factors such as people moving around, sudden changes in lighting, and environmental noise can cause facial recognition (FR) and object tracking (MOT) algorithms to fail, as well as a sharp drop in speech recognition (ASR) accuracy, making continuous and stable human-computer interaction impossible. The root cause lies in the lack of a closed-loop control mechanism that integrates environmental perception, signal interference resistance, and online adaptive adjustment of algorithm parameters.

[0003] 3. Bottlenecks in Functional Module Collaboration and Knowledge Targeting: Existing systems are mostly simple stacking of functional modules, with training demonstrations and Q&A guidance being disconnected, resulting in a fragmented user experience. More importantly, their built-in knowledge bases contain general content and lack domain knowledge graphs and intent recognition models that are highly matched to the physiological characteristics, rehabilitation stages, and common questions of specific groups (such as addicts). This makes it impossible to provide personalized feedback with medical rehabilitation guidance value, significantly reducing the practicality of the technology.

[0004] Therefore, there is an urgent need for an Baduanjin training system and method based on digital human technology to solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to provide an Eight-Section Brocade training system and method based on digital human technology. It achieves high-precision digital human movement reproduction, stable human-computer interaction in an open environment, intelligent Q&A and training closed loop, and provides customized guidance for the rehabilitation needs of addicts.

[0006] To achieve the above objectives, the present invention employs the following technical solution: On one hand, the present invention provides an Eight-Section Brocade training system based on digital human technology, comprising: The Eight-Section Brocade Digital Human with Module, Intelligent Question and Answer Module, Digital Human All-in-One Machine and Open Environment Interaction Module; The Baduanjin digital human module is used to generate a highly realistic digital human model through motion capture and 3D modeling technology, and reproduce the standard Baduanjin movements. The intelligent question-answering module is used for semantic understanding and answer generation based on natural language processing technology and a knowledge base; The digital human all-in-one machine integrates facial capture and human locking functions to identify and track users in open environments; The open environment interaction module is used to sense environmental interference and perform anti-interference processing to ensure the stable operation of the system in complex environments.

[0007] Preferably, the Eight-Section Brocade digital human-guided practice module includes: The digital human creation unit is used to build highly realistic digital human models based on 3D scanning and human anatomical data; The motion capture unit is used to collect standard Baduanjin movement data from professional coaches using optical or inertial motion capture equipment. The 3D modeling integration unit is used to match and fuse motion data with digital human models to generate interactive Eight-Section Brocade dance animations.

[0008] Preferably, the motion data collected by the motion capture unit includes three-dimensional coordinates, joint angles, and motion speed, and its data processing satisfies the following relationship: ; in, Indicates time The three-dimensional coordinate vector of the joint at any given time; Indicates the angle of joint flexion; , These are vector representations of adjacent joint segments.

[0009] Preferably, the intelligent question-answering module includes: The knowledge base building unit is used to integrate the key points, effects, precautions, and common questions and answers for people trying to quit addiction. The question-answering algorithm unit uses a BERT-based deep learning model for semantic understanding and intent recognition, and its output response probability is expressed as follows: ; in, Indicates the input text sequence; Indicates the type of the output answer; and These are the model parameters.

[0010] Preferably, the digital human-machine interface includes: The face capture unit is used to extract facial feature points of the user through a high-precision camera and facial recognition algorithm; The person locking unit is used to continuously track the user's location using a target tracking algorithm; The interactive implementation unit is used to conduct question-and-answer interactions with users using either voice or text.

[0011] Preferably, the target tracking unit uses a Kalman filter algorithm for target tracking, and its state update formula is as follows: ; in, Indicates the first Frame target state estimate; Represents the state transition matrix; Represents the control matrix; Indicates control input; Indicates Kalman gain; Represents the observed value; This represents the observation matrix.

[0012] Preferably, the open environment interaction module includes: An environmental sensing unit is used to acquire information about obstacles, lighting, and noise in the environment through cameras or sensors. The anti-interference processing unit is used to process speech and image signals using filtering and noise reduction algorithms. Its noise reduction process is described as follows: ; in, The original signal; For transfer functions; For noise; This is the processed signal.

[0013] On the other hand, the present invention also provides a training method for Baduanjin based on digital human technology, applied to the Baduanjin training system based on digital human technology as described above, including the following steps: Step S1: Start the system, and the digital human-guided striking module demonstrates the standard movements of Baduanjin. Step S2: The user asks a question via voice or gesture; Step S3: The digital human all-in-one machine identifies the user through facial capture and person locking; Step S4: The intelligent question-answering module processes the question and generates the answer; Step S5: The digital human replies to the user in voice or text format; Step S6: The open environment interaction module adjusts system parameters in real time to cope with environmental changes.

[0014] Preferably, the process of the intelligent question-answering module processing the question in step S4 includes: Semantic parsing, intent recognition, knowledge base retrieval, and answer generation are performed. Intent recognition employs an attention mechanism to enhance contextual understanding, and its attention weights are calculated as follows: ; in, Represents the query vector. Represents a keyword vector. This represents the attention weight.

[0015] Preferably, the adaptive adjustment of environmental parameters in step S6 includes: The camera exposure time is automatically adjusted based on light intensity, and the microphone gain and filter parameters are adjusted based on ambient noise levels. The adjustment strategy is as follows: ; in, Indicates microphone gain; Indicates the level of ambient noise; Indicates the reference noise value; Indicates the exposure time; Indicates ambient light intensity; This represents the reference illumination value.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention overcomes the challenge of motion fidelity through a combination of high-precision optical motion capture and anatomically constrained 3D modeling. Specifically, the digital human's motion data originates from millimeter-precision motion trajectory capture by professional coaches and is calculated using inverse kinematics algorithms. The resulting joint angle data deviates from the standard Baduanjin (Eight Pieces of Brocade) by less than 3 degrees, far exceeding the simulation effect of keyframe animation. This allows addicts to obtain standard motion references indistinguishable from those guided by professional coaches offline, ensuring the effectiveness and safety of training from the outset.

[0017] 2. This invention achieves robust interaction in open environments through the use of "Kalman filter target tracking + multimodal signal anti-interference processing" technology. Specifically, the target tracking algorithm integrated into the digital human-machine interface significantly reduces tracking loss rate under interference such as personnel movement and partial occlusion; and through microphone array beamforming and spectral noise reduction algorithms, speech recognition accuracy remains high even in 70dB of ambient noise. This results in the first-ever realization of a continuous, stable, and highly available natural human-computer interaction experience in complex open environments.

[0018] 3. This invention achieves a deep closed loop of training and Q&A through the integration of BERT-based semantic understanding and a knowledge graph in the field of addiction rehabilitation. Specifically, the intelligent Q&A module significantly improves the accuracy of user intent classification and provides faster response times. More importantly, its answer generation is not based on general network information but originates from a knowledge base specifically built for addiction patients, providing precise and safe medical guidance on specific issues such as "the range of motion for patients with liver damage" and "breathing techniques for relieving anxiety." This forms an integrated training loop where "what you see is what you ask, and what you ask is what you get," upgrading the digital human from a demonstration tool into a professional personal rehabilitation mentor. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart of the data preprocessing process of the present invention. Detailed Implementation

[0020] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined in this application.

[0021] In this invention, terms such as "upper," "lower," "left," "right," "front," "back," "vertical," "horizontal," "side," and "bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used only to facilitate the description of the structural relationships of the various components or elements of this invention and do not specifically refer to any component or element in this invention. They should not be construed as limiting the invention.

[0022] Example: like Figure 1 As shown, this embodiment provides an Baduanjin training system based on digital human technology, including: a Baduanjin digital human practice module, an intelligent question-and-answer module, a digital human all-in-one machine, and an open environment interaction module; these four modules work together logically and physically to form a complete intelligent training solution.

[0023] Specifically: 1. Eight-Section Brocade Digital Human-Assisted Practice Module: This module is the core demonstration unit of the system, responsible for generating a highly realistic digital human image and driving it to complete standard and regulated Baduanjin movements; its specific technical implementation is divided into three sub-units: Digital Human Creation Unit: This unit uses a high-precision 3D laser scanning system to perform static body scans on professional Baduanjin instructors, acquiring high-density point cloud data of their body surfaces, with a point cloud density of no less than 100 points / square centimeter. Subsequently, based on the Non-Uniform Rational B-Spline (NURBS) surface reconstruction technology in computer graphics, the point cloud data is converted into a 3D mesh model with a continuous smooth surface. To further improve the biomechanical accuracy of the model, this unit introduces a human anatomy database. This database contains parameters such as the length and density of major bones, joint types (e.g., ball-and-socket joints, hinge joints), as well as the origin and insertion points of muscles, and mechanical properties (e.g., contractile force and elastic modulus). By rigidly and elastically binding the 3D mesh model with the anatomical data, each vertex in the model is assigned corresponding bone and muscle weights, thereby constructing a highly realistic digital human model that not only looks lifelike but also has an internal structure that conforms to human kinematics. This model can accurately simulate muscle deformation and joint rotation limits during movement. Motion Capture Unit: This unit employs a motion capture system based on infrared optical markers. At least 52 reflective markers are affixed to key anatomical locations (such as the acromion, lateral epicondyle of the humerus, radial styloid process, anterior superior iliac spine, lateral femoral condyle, and lateral malleolus) by professional Baduanjin instructors. Eight high-speed infrared cameras (sampling frequency no less than 200Hz) positioned around the capture space synchronously capture the motion trajectory of these markers. The infrared light emitted by the cameras is reflected by the markers and received by image sensors within the camera lenses. Using stereoscopic vision principles, the coordinate sequence of each marker in three-dimensional space is calculated. Subsequently, using inverse kinematics (IK) algorithms, the rotational angles of the internal joint chains of the human body were calculated from the trajectories of these surface markers. For each frame of data, the joint angle The calculation is achieved by solving for the angle between the vectors, and its mathematical expression is:

[0024] in, and These represent the two limb segments connected by the joint in... The spatial vector at time points, and the final output data is a six-degree-of-freedom (position and orientation) motion data stream for all joints, including timestamps; 3D Modeling Integration Unit: This unit receives standard motion data streams and highly realistic digital human models from the motion capture unit. The integration process is completed in a professional 3D animation engine (such as Unity3D or Unreal Engine). First, the skeletal structure of the digital human model is mapped one-to-one with the joint nodes in the motion data. Then, linear interpolation (Lerp) or spherical linear interpolation (Slerp) algorithms are used to smooth the motion between keyframes, eliminating minor jitter that may occur during data acquisition. Next, skinning rendering technology is used to calculate the deformation of the model's surface mesh in real time based on the movement of the skeleton, ensuring that the movement of muscles and skin is natural and smooth. Finally, an eight-section brocade belt dance animation sequence that can be rendered in real time and controlled by external commands is generated. This sequence can be played and demonstrated at any viewpoint and speed.

[0025] 2. Intelligent Question Answering Module: This module is responsible for understanding the natural language questions posed by users and providing accurate answers; it is a core of AI-based interaction. Knowledge Base Construction Unit: This unit constructs a structured domain knowledge graph. The knowledge sources include authoritative Baduanjin classics, traditional Chinese medicine theoretical works, sports medicine literature, and clinical research data on the rehabilitation of addicts. The knowledge is stored in the form of triplets (entity-relationship-entity), such as ("Hands supporting the sky to regulate the three jiaos", "efficacy", "regulate qi", "shake head and wag tail to clear heart fire", "precautions", "patients with hypertension should have smaller range of motion"), ("addicts", "prohibited movements", "extreme forward bend"). In addition, a large number of QA pairs are collected as the corpus basis for generative responses. Question Answering Algorithm Unit: This unit uses the pre-trained language model BERT (Bidirectional Encoder Representations from Transformers) based on the Transformer architecture as the core of semantic understanding. When a user inputs a question `X`, it first undergoes text preprocessing (word segmentation, stop word removal). The processed word sequence is then input into the BERT model. The model outputs a contextual semantic embedding vector for each word. For classification questions (such as determining whether the question belongs to "action instructions" or "efficacy consultation"), intent classification is performed by connecting a fully connected layer and a Softmax function to the BERT output. The probability calculation is as follows: in, It is the output of BERT that represents the semantics of the entire sentence. The vector of the marker, and The system consists of a trainable parameter matrix and bias terms. For open-ended questions that require retrieving answers from a knowledge base, the system uses the BERT vector of the question itself to perform vector similarity retrieval in the knowledge base (e.g., using cosine similarity) to find the most relevant knowledge fragments. Finally, the answer generator integrates the retrieved information into fluent natural language sentences, which are then broadcast or displayed in text form using text-to-speech (TTS) technology.

[0026] 3. Digital Human All-in-One Machine: This module serves as the system's hardware platform and real-time interactive interface, integrating various sensors and processing units. Face capture unit: This unit integrates a high-definition RGB camera with a resolution of 1080p or higher and an infrared depth camera. It adopts a facial feature point detection algorithm based on convolutional neural network (CNN) (such as Google's MediaPipe Face Mesh). This algorithm can locate 468 3D feature points of the face in real time in each frame of the image, including the outline of the lips, eyes and eyebrows, so as to accurately describe the user's facial expressions and mouth opening and closing state. Person Tracking Unit: This unit employs a target tracking algorithm based on correlation filtering (such as the KCF algorithm) or deep learning. Once the face capture unit identifies the user, this unit continuously tracks the user in subsequent video frames. To handle short-term occlusion or rapid movement of the user, this unit introduces a Kalman filter algorithm for motion prediction. Its state equation and observation equation are as follows: State prediction: ; ; Status Update: ; ; ; in, yes The state estimation vector at any given time (such as the position and velocity of the target center point). It is the error covariance matrix. It is the state transition matrix. It is the observation matrix. yes The observed value at a given time (such as the detected target location). and These are the covariance matrices of process noise and observation noise, respectively. It is the Kalman gain. Through recursive calculation, this algorithm can effectively smooth the tracking trajectory and make predictions when observations are lost, thus maintaining the stability of the lock. Interactive Implementation Unit: This unit is a software middleware responsible for coordinating the facial capture, person locking, and intelligent question answering modules. It receives user identity confirmation signals from the facial capture unit and continuous coordinate information from the person locking unit to ensure that the digital human's gaze point and body orientation are always facing the locked user. When it detects that the user has initiated a voice question, it calls the interface of the intelligent question answering module and forwards the voice data. After receiving the text answer, it drives the digital human model to perform lip-sync (based on a phoneme-lip-sync table) and reads out the answer.

[0027] 4. Open Environment Interaction Module: This module ensures the stable operation of the system in uncontrolled environments: Environmental perception unit: In addition to using the camera on the digital human all-in-one machine, this unit can also integrate LiDAR and microphone array; LiDAR is used to scan the surrounding environment and generate point cloud maps to perceive obstacles and spatial layout, and the microphone array is used to preliminarily determine the direction of the sound source by calculating the time difference of arrival (TDOA) of sound reaching different microphones. Anti-interference processing unit: This unit contains a series of signal processing algorithms: For image signals: an algorithm based on histogram equalization and adaptive gamma correction is used to cope with changes in illumination, and a median filter is used to eliminate image noise; For speech signals: First, beamforming technology of the microphone array is used to form a pickup beam in the estimated direction of the sound source to enhance the target speech and suppress environmental noise from other directions. Then, spectral subtraction or a deep learning-based speech enhancement model (such as SEGAN) is used to further reduce noise in the acquired speech, improve the signal-to-noise ratio (SNR), and finally output a clean speech stream to the subsequent speech recognition (ASR) engine.

[0028] like Figure 2 As shown, this embodiment also provides a training method for Baduanjin (Eight Pieces of Brocade) based on digital human technology. This method is executed by the aforementioned system and includes the following steps: Step S1: System initialization and action demonstration startup, specifically: When the system is powered on, each module completes self-test and initialization. The Baduanjin digital human module loads the pre-generated high-fidelity digital human model and standard Baduanjin movement data stream. The digital human all-in-one machine begins to render and display the digital human image. The digital human begins to perform high-fidelity movement demonstrations according to the preset training courses (such as "Preparatory Posture" and "Holding the Sky with Both Hands to Regulate the Three Jiaos"). The rendering engine calculates the lighting and shadow effects in real time, so that the digital human can be integrated into the user's real-time shooting environment background, forming an augmented reality (AR) visual effect. Step S2: Capturing user interaction intent, specifically: During training, users can initiate interactions through specific wake words (such as "digital human assistant") or preset gestures (such as raising hands). The facial capture unit of the digital human all-in-one machine runs continuously. Once a valid user is identified and their interaction intention is detected, the human locking unit is immediately activated. The user is continuously tracked and locked using a Kalman filter algorithm. The open environment interaction module works synchronously. The environment perception unit monitors the ambient light and noise levels in real time, and the anti-interference processing unit dynamically adjusts the camera parameters and audio front-end processing parameters to provide the highest quality raw signal for subsequent interactions. Step S3: Problem identification and transmission, specifically: When a user asks a question in natural language, the noise-reduced and enhanced speech stream is transmitted to the speech recognition (ASR) engine and converted into text data. This text data The interface is bound to the user ID that asked the question and sent to the intelligent question-answering module. Step S4: Intelligent semantic understanding and answer generation, specifically: The question-answering algorithm unit of the intelligent question-answering module receives the question text. First, the BERT model is used to encode and semantically understand the question, enabling intent recognition and key information extraction. Then, retrieval and reasoning are performed within a constructed knowledge base dedicated to Baduanjin exercises and addiction rehabilitation. The retrieval process combines keyword-based matching and vector-based semantic similarity matching. Finally, an answer generator integrates the retrieved information into a complete, accurate, and natural answer text. ; Step S5: Multimodal presentation of the answer, specifically: The generated answer text The interactive implementation unit is returned to the digital human all-in-one machine; this unit drives two parallel processes: Firstly, The text is fed into a text-to-speech (TTS) engine to generate the corresponding audio stream. Second Real-time phoneme analysis is performed to drive the digital human model's lip movements and synthesized speech to achieve audio-visual synchronization. Ultimately, the digital human presents the answer to the user with realistic expressions and voice, while maintaining eye contact with the targeted questioner; Step S6: Environmental Adaptation and Process Cycle, specifically: Throughout the entire interaction process, the open environment interaction module runs in the background, forming a closed-loop feedback system. Its environment sensing unit continuously collects environmental data. (Light intensity and noise level), compared to the preset ideal reference values. The comparison generates an error signal; this error signal drives the control algorithm (such as a PID controller) to dynamically adjust the system parameters. (such as camera gain, exposure time, microphone gain, filter coefficients), making This is done to minimize environmental interference and ensure high quality and uninterrupted operation of the entire training and interaction process. Once completed, the system automatically returns to step S1, and the digital human continues to demonstrate actions, waiting for the next interaction, thus forming a complete "demonstration-interaction-feedback" closed loop.

[0029] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A digital human technology-based Baduanjin exercise system, characterized in that, Comprise: Baduanjin digital human belt beating module, intelligent question and answer module, digital human all-in-one machine and open environment interaction module; The baduanjin digital human belt beating module is used to generate a high-simulation digital human model through motion capture and 3D modeling technology, and reproduce standard baduanjin movements; The intelligent question and answer module is used for semantic understanding and answer generation based on natural language processing technology and knowledge base; The digital human all-in-one machine integrates face capture and character locking functions for identifying and tracking users in an open environment; The open environment interaction module is used to perceive environmental interference and perform anti-interference processing to ensure stable operation of the system in complex environments.

2. The digital human technology-based eight-section exercise training system according to claim 1, characterized in that, The baduanjin digital human belt beating module comprises: A digital human production unit for constructing a high-simulation digital human model based on three-dimensional scanning and human anatomy data; A motion capture unit for collecting professional trainer's standard baduanjin movement data through optical or inertial motion capture devices; A 3D modeling integration unit for matching and fusing movement data with the digital human model to generate interactive baduanjin belt beating animations.

3. The digital human technology-based eight-section exercise system according to claim 2, characterized in that, The movement data collected by the motion capture unit includes three-dimensional coordinates, joint angles, and movement speed, and the data processing process satisfies the following relationship: ; wherein, represents time the joint three-dimensional coordinate vector at the time point; represents the joint flexion angle; , are vector representations of adjacent joint segments, respectively.

4. The digital human technology-based eight-section exercise training system according to claim 1, characterized in that, The intelligent question and answer module comprises: A knowledge base construction unit for integrating baduanjin movement essentials, effects, precautions, and common questions and answers of drug addicts; A question and answer algorithm unit using a BERT-based deep learning model for semantic understanding and intent recognition, with the output response probability represented as: ; wherein, represents an input text sequence; represents an output answer category; and are model parameters.

5. The digital human technology-based eight-section exercise training system according to claim 1, characterized in that, The digital human all-in-one machine comprises: A face capture unit for extracting user facial feature points through high-precision cameras and facial recognition algorithms; A character locking unit for continuously tracking user locations through target tracking algorithms; An interaction implementation unit for conducting question and answer interactions with users through voice or text methods.

6. The digital human technology-based eight-section exercise system according to claim 5, characterized in that, The character locking unit uses Kalman filtering algorithm for target tracking, and its state update formula is: ; wherein, represents the target state estimate value of the frame; represents the target state estimate value of the frame; represents the state transition matrix; represents the control matrix; represents the control input; represents the Kalman gain; represents the observation value; represents the observation matrix.

7. The digital human technology-based eight-section exercise training system according to claim 1, characterized in that, The open environment interaction module comprises: An environmental perception unit for obtaining obstacle, light, and noise information in the environment through cameras or sensors; An anti-interference processing unit for processing voice and image signals using filtering and noise reduction algorithms, with the noise reduction process represented as: ; wherein is the original signal; is the transfer function; is the noise; is the processed signal.

8. A digital human technology-based Ba Duan Jin training method applied to the digital human technology-based Ba Duan Jin training system of any one of claims 1-7, characterized in that, Comprise the following steps: Step S1: Start the system, and the digital human belt beating module demonstrates standard baduanjin movements; Step S2: The user initiates a question through voice or gesture; Step S3: The digital human all-in-one machine identifies the user through face capture and character locking; Step S4: The intelligent question and answer module processes the question and generates an answer; Step S5: The digital human replies to the user in the form of voice or text; Step S6: The open environment interaction module adjusts system parameters in real time to respond to environmental changes.

9. A method according to claim 8, wherein, The process of the intelligent question and answer module processing the question in step S4 comprises: Semantic analysis, intent recognition, knowledge base retrieval, and answer generation, wherein the intent recognition uses an attention mechanism to enhance context understanding, and the attention weight calculation is: ; wherein, denotes a query vector, denotes a keyword vector, denotes an attention weight.

10. A method according to claim 8, wherein, The environmental parameter adaptive adjustment in step S6 comprises: Adjusting the camera exposure time automatically according to the light intensity, and adjusting the microphone gain and filter parameters according to the environmental noise level, with the adjustment strategy being: ; wherein, represents a microphone gain; represents an ambient noise level; represents a reference noise value; represents an exposure time; represents an ambient light intensity; represents a reference light value.