A digital assistant employee system for traditional chinese medicine health preservation based on intelligent multi-modal interaction

By constructing an intelligent, multimodal, interactive TCM health-preserving auxiliary digital employee system, the problem of dynamic motion reconstruction and real-time analysis in existing technologies has been solved. This system enables quantitative assessment and multimodal feedback of TCM guided movements, thereby improving the user learning experience and training effectiveness.

CN122117340APending Publication Date: 2026-05-29JIANGSU XINWANG VIDEO SOFTWARE TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU XINWANG VIDEO SOFTWARE TECH CO LTD
Filing Date
2026-04-30
Publication Date
2026-05-29

Smart Images

  • Figure CN122117340A_ABST
    Figure CN122117340A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on intelligent multi-modal interaction traditional Chinese medicine health preservation auxiliary digital staff system, including three-dimensional posture reconstruction module, standard action knowledge base, interactive rectification analysis module and multi-modal feedback module;System is reconstructed in real time with three-dimensional skeleton model in line with human biomechanics by monocular video stream, and is matched with the knowledge base of embedded three-dimensional meridian path and acupoint biomechanics response data;Using Fréchet distance algorithm and proxy model quantization evaluation meridian compliance and acupoint activation state, then accurately diagnose action deviation root by multi-head attention network;Finally, virtual digital staff renders meridian blood state by particle flow visualization, and outputs by language model driven situational voice to guide.The application converts abstract traditional Chinese medicine theory into computable index, realizes accurate and objective action rectification and immersive, individualized traditional Chinese medicine health preservation teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction, and in particular to a digital employee system based on intelligent multimodal interaction and traditional Chinese medicine health preservation assistance. Background Technology

[0002] In recent years, with the rapid evolution of computer vision and artificial intelligence technologies, digital health assistance systems based on 3D human posture estimation have been widely applied. These systems typically utilize visual acquisition devices to capture human movement trajectories, reconstruct 3D skeletal models, and on this basis, develop various virtual coach or digital employee systems to assist users in physical training and exercise guidance. Simultaneously, traditional Chinese medicine (TCM) exercises (such as Baduanjin and Wuqinxi) are gradually transforming towards digitalization and intelligence as a distinctive TCM health intervention method. Some existing TCM health assistance systems can already incorporate virtual humans for movement demonstrations and use conventional visual algorithms to compare the spatial coordinates of users' limb joints, thereby achieving basic movement following and preliminary posture correction feedback. This, to some extent, lowers the barrier for users to learn TCM exercises and promotes the digital dissemination of traditional health knowledge.

[0003] Currently, domestic and international patents have been researching the digital visualization of TCM theory and the intelligent positioning of TCM acupoints. CN116541007A discloses a method and system for visualizing and mapping physical signs based on TCM syndrome differentiation. This solution achieves static visualization and component-based encapsulation of TCM syndrome differentiation physical signs by constructing a TCM human body model, decomposing disease location and disease nature elements, and performing multi-layer superimposed mapping, thus lowering the understanding threshold of TCM syndrome differentiation theory. However, this solution only focuses on the design of static physical sign visualization and does not involve the three-dimensional reconstruction and real-time analysis of human dynamic movements. It cannot achieve interactive correction of guided movements, nor does it establish the correlation between dynamic movements and meridian circulation and acupoint biomechanical responses. It cannot quantitatively evaluate the TCM conditioning effect of guided movements and lacks a multimodal interactive virtual digital human guidance system, making it difficult to adapt to interactive teaching scenarios of TCM health preservation guidance. CN118860245A discloses a standardized acupoint intelligent positioning and assisted acupuncture system based on AR technology. This solution combines CT scan 3D modeling with AR technology to achieve high-precision static positioning and visualization of acupoints in acupuncture scenarios, providing an auxiliary tool for acupuncture clinical practice and teaching. However, this solution is only designed for static acupuncture acupoint positioning scenarios and does not involve the three-dimensional capture and temporal analysis of dynamic guiding movements. It lacks the ability to evaluate and correct movements, and does not consider the changes in meridian path deformation and acupoint biomechanical response during dynamic movements. It cannot establish a quantitative correlation between movements and TCM conditioning effects, and it also lacks a multimodal real-time feedback mechanism that integrates TCM theory, failing to provide users with immersive guidance and correction for guiding movements. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides an intelligent multimodal interactive TCM health-preservation auxiliary digital employee system to solve the problems mentioned in the background art.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides an intelligent multimodal interactive TCM health preservation-assisted digital employee system, comprising: A 3D pose reconstruction module is configured to receive a monocular video stream of a user's body and reconstruct a 3D human skeleton model representing the user's pose in real time based on the monocular video stream. A standard movement knowledge base, wherein the knowledge base pre-stores standard three-dimensional movement sequences corresponding to traditional Chinese medicine guiding techniques, and each frame of the standard three-dimensional movement sequence is associated with a digitized three-dimensional meridian circulation path model and preset key acupoint biomechanical response data; An interactive correction analysis module is configured to match the user's real-time three-dimensional human skeleton model with the data in the standard action knowledge base, and based on a preset evaluation algorithm, to quantitatively analyze the deviation between the user's posture and the standard posture from the aspects of kinematic morphology, meridian path conformity and acupoint activation status, and generate correction diagnosis information. A multimodal feedback module includes a virtual digital employee configured to receive the corrective diagnostic information and provide users with real-time corrective guidance that integrates with traditional Chinese medicine theory by graphically rendering the information on a display interface and outputting voice commands.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This system breaks through the technical limitations of traditional motion-assisted software, which only focuses on the comparison of three-dimensional limb coordinates. By constructing a standard movement knowledge base that links three-dimensional meridian pathways with acupoint biomechanical response data, and combining the discrete Fraser distance algorithm and lightweight proxy model, it transforms the abstract concepts of obtaining qi and guiding qi through form in traditional Chinese medicine into real-time calculable physical and geometric indicators, providing objective and scientific data support for judging the effectiveness of traditional Chinese medicine guiding movements.

[0008] 2. This invention introduces an auxiliary digital employee. Visually, it utilizes a particle system inspired by computational fluid dynamics to intuitively present the operational status of meridians and blood circulation on the user's virtual image with dynamic light effects. Auditorily, it uses a language model combined with the user's historical profile and fuzzy logic reasoning results to generate contextualized TCM colloquial voice commands. Through these multimodal interactions, the cognitive threshold for users to learn TCM health preservation techniques is reduced, and training focus is improved.

[0009] 3. Furthermore, this invention utilizes a diagnostic network based on a multi-head attention mechanism to directly trace the fundamental kinematic characteristics that lead to action deviations, breaking the "black box" of evaluation. At the same time, combined with a personalized curriculum planning module built using a reinforcement learning model, digital employees can, like traditional Chinese medicine therapists, autonomously try and dynamically adjust long-term health intervention strategies based on users' daily performance and physical fluctuations, thereby achieving a personalized approach. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is an architecture diagram of an intelligent multimodal interactive TCM health preservation auxiliary digital employee system according to an embodiment of the present invention. Detailed Implementation

[0011] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0012] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0013] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0014] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Example 1

[0015] Reference Figure 1 This is the first embodiment of the present invention, which provides an intelligent multimodal interactive TCM health preservation-assisted digital employee system, comprising: It should be noted that, in order to combine the characteristics of traditional Chinese medicine guiding techniques (such as Baduanjin and Tai Chi) that emphasize the continuity of form and the precision of force exertion, this system first needs to obtain the user's three-dimensional spatial posture.

[0016] Furthermore, the 3D pose reconstruction module receives a monocular video stream of the user's body and reconstructs a 3D human skeleton model representing the user's posture in real time based on this video stream. Considering that ordinary users in home or office settings typically only have a standard single-camera device (e.g., a mobile phone or computer webcam), and that in conventional 3D reconstruction, single-camera devices are prone to depth blur and limb scaling artifacts, leading to fatal errors in the evaluation of movements in traditional Chinese medicine (TCM) exercises (e.g., the arm extension amplitude in the "Drawing the Bow to Shoot the Eagle" movement); and that the movements in TCM exercises are not isolated static postures but continuous dynamic processes of Qi and blood flowing along meridians, with joints of the limbs interacting spatially and exhibiting strong temporal dependence, this 3D pose reconstruction module also integrates a spatiotemporal graph convolutional network to extract and generate a preliminary 3D human skeleton coordinate sequence from the monocular video stream.

[0017] Specifically, firstly, the system extracts the coordinates of two-dimensional human joints frame by frame from the monocular video stream using a front-end feature extraction backbone network (such as HRNet or ResNet). Then, these continuous sequences of two-dimensional joints are constructed into a standard spatiotemporal graph structure. , where the set of nodes Representing the human body Key nodes, edge set It includes two types of edges: the first type is skeletal connections within the same frame that conform to the natural anatomical structure of the human body (spatial edges), and the second type is trajectory connections of the same joint points between adjacent frames (temporal edges). Finally, the spatiotemporal graph convolutional network receives this spatiotemporal graph as input, and its core spatial graph convolutional layer is used to capture the coordination relationship between joints at the same time. Its update formula is as follows: in, Indicates the first The hidden layer feature matrix of the layer, its initial input This refers to the spatiotemporal two-dimensional key point coordinates and confidence data extracted above; This indicates the number of preset spatial division strategies (e.g., dividing adjacent nodes into centripetal nodes, centrifugal nodes, and static nodes to better distinguish the extension and contraction states of limbs in traditional Chinese medicine guiding exercises). It is the first The spatial adjacency matrix under this partition represents the physical connection topology between joints; It corresponds to The degree matrix is ​​used to normalize the adjacency matrix to prevent eigenvalue explosion or vanishing after multiple convolutions. It is the first The weight parameter matrix that needs to be learned through the backpropagation algorithm in a layered network; It is a non-linear activation function (e.g., the ReLU function).

[0018] Simultaneously, a temporal graph convolutional network is used to extract temporal coherence features of actions by sliding a window along the time dimension. This spatiotemporal graph convolutional network, after pre-training on a large-scale human action dataset with real 3D labels (and incorporating some action datasets specific to traditional Chinese medicine guiding exercises), can output a preliminary three-dimensional human skeleton coordinate sequence containing depth information, denoted as... (Where, 3 represents three-dimensional spatial coordinates) , Represents the number of video frames. (Represents the number of key points).

[0019] Furthermore, although the aforementioned spatiotemporal graph convolutional network can reconstruct three-dimensional coordinates, the accuracy of movement patterns is crucial in traditional Chinese medicine (TCM) health preservation scenarios. For example, in stances like the horse stance or the half-stance, the angles of the joints directly determine whether the corresponding meridians can be effectively activated. Simply relying on the initial three-dimensional human skeletal coordinate sequence output by the convolutional network often results in physical distortions such as bone length stretching over time or joint inversion. Therefore, this three-dimensional pose reconstruction module also includes a physical constraint optimization unit to apply biomechanical constraints to the initial three-dimensional human skeletal coordinate sequence, outputting a three-dimensional human skeletal model that conforms to physical laws.

[0020] Specifically, this optimization unit transforms the pose reconstruction problem into an energy minimization optimization problem. It performs secondary optimization on the initial 3D human skeleton coordinate sequence by constructing an energy function (loss function) that incorporates various biomechanical priors. The defined total energy optimization function is as follows: in, The current 3D skeleton coordinate variables to be optimized; This is a data item used to ensure that the optimized posture does not deviate excessively from the initial prediction of the neural network, thus ensuring the correct overall movement. As a constraint term for bone length consistency, since the length of the human limbs remains physically constant during continuous guided movements, this formula can be preferably defined as: , and Indicates in Constantly forming bones The three-dimensional coordinates of the two joints, The actual bone segment length of the user is calculated from the user's initial calibration frame. This constraint is used to eliminate the common phenomenon of inconsistent arm length in monocular reconstruction. This is a joint limit angle constraint term. Since the flexion and extension of human joints have fixed limits (e.g., the knee joint cannot bend forward backward), this term can be calculated by measuring the angle between adjacent bone vectors. And set up a penalty mechanism, preferably defined as: , and They are the first The maximum and minimum permissible angles of each joint are determined to ensure the rationality of the reconstructed TCM standing postures or stretching movements. As a temporal smoothness constraint, since traditional Chinese medicine guiding exercises emphasize smooth, gentle, and restrained movements, an acceleration penalty is introduced to eliminate high-frequency jitter caused by video prediction. This acceleration penalty is preferably defined as follows: This is used to make the reconstructed motion trajectory natural and smooth. The weighting coefficients for the corresponding constraint terms can be dynamically adjusted according to the different movement characteristics of traditional Chinese medicine guiding exercises.

[0021] It should be noted that the total energy optimization function is solved rapidly in the background using the gradient descent method through this physical constraint optimization unit, so that the output of the total energy optimization function reaches a minimum value. The final 3D human skeleton model output after optimization not only eliminates visual jitter and distortion, but more importantly, it also provides a 3D space that conforms to the real biomechanical structure of the human body for the calculation of the topological attachment of the meridian pathways and the force state of acupoints.

[0022] It should be noted that, unlike conventional fitness software (such as yoga or calisthenics apps) which only store a set of standard skeletal animations, the purpose of Traditional Chinese Medicine (TCM) guiding exercises is to guide Qi and blood through form. The correctness of a movement lies not only in the placement of the limbs, but also in whether the movement effectively stretches specific meridians and provides sufficient physical stimulation to the fascia tissues of specific acupoints. Therefore, this invention constructs a standard movement knowledge base, which pre-stores standard three-dimensional movement sequences corresponding to TCM guiding exercises. Each frame of the standard three-dimensional movement sequence is associated with a digitized three-dimensional meridian pathway model and preset biomechanical response data for key acupoints.

[0023] Furthermore, in the field of computer vision, standard 3D human models typically only contain surface meshes and internal skeletons. To visualize the abstract twelve regular meridians and eight extraordinary meridians of Traditional Chinese Medicine (TCM), this embodiment introduces a meridian topology binding algorithm based on a parametric human model (e.g., the SMPL model) on top of a pre-made standard 3D motion sequence. Specifically, TCM experts and anatomical experts collaborate to perform high-precision spatial modeling of the meridian pathways in the 3D space of a benchmark human model using non-uniform rational B-spline curves. Since the skin and subcutaneous tissue deform during human movement, the meridians are not rigidly attached to the bones. To ensure that the 3D direction of the meridian path conforms to the actual physiological sliding patterns of the human body in each frame of standard motion, the system uses a flexible topology binding formula based on Linear Blend Skinning (LBS) extension to define the time frame. Arbitrary control point on the meridian curve From the spatial coordinates, we obtain: in, Indicates at time In the standard action frame, the first The three-dimensional absolute coordinates of each meridian control point; This represents the total number of joints in the human skeleton. This is the meridian-skeletal-skin weighting coefficient, which reflects the first... The meridian point is affected by the first The degree of influence of skeletal joint movement is set by this parameter to distinguish between superficial and deep meridians. For example, meridian nodes attached to the ends of the limbs are greatly affected by local joints and have a concentrated weight distribution, while meridian nodes running through the trunk are affected by multiple spinal nodes and have a divergent weight distribution. Indicates time No. Rigid body transformation matrix (including rotation and translation) of each skeletal joint; The initial three-dimensional coordinates of this meridian control point in a standard upright resting state; For the nonlinear deformation compensation term related to posture, since there are many extreme twisting movements in traditional Chinese medicine exercises (such as looking backwards in the Five Labors and Seven Injuries of Baduanjin), simple linear skinning would cause the meridian curves to collapse in volume or clipping at the joint folds. Therefore, this parameter needs to be fitted to the joint angle set using a multilayer perceptron (MLP). The resulting muscle bulges and fascial sliding offsets ensure that the three-dimensional dynamic representation of the meridian model conforms to real anatomical dynamics.

[0024] It should be noted that by establishing a meridian model, each frame of standard action stored in the knowledge base carries three-dimensional meridian grid data that flows naturally with the limbs, similar to a vascular network.

[0025] Furthermore, Traditional Chinese Medicine (TCM) theory holds that the arrival of Qi at the site of illness is often accompanied by sensations of soreness, numbness, distension, and pain at the acupoint (i.e., the sensation of obtaining Qi). Modern medical and biomechanical research indicates that this sensation of obtaining Qi is essentially caused by the localized stress and strain experienced by the fascia, muscles, and connective tissues during specific movements reaching a certain threshold, thereby activating the mechanoreceptors at the base of the acupoint. Because the calculation of this physical process is extremely complex and cannot be performed in real-time on the user's mobile terminal, this system employs a pre-processing offline finite element analysis (FEA) method during the knowledge base construction phase.

[0026] Specifically, firstly, the system establishes a volumetric mesh model including bones, muscles, subcutaneous fat, and deep fascia. Key acupoints of interest in Traditional Chinese Medicine (such as Zusanli, Hegu, and Neiguan) are mapped to specific sets of physical units within this continuum mechanics model. Secondly, kinematic parameters (i.e., displacement boundary conditions) of a standard action sequence at different time points are applied to this volumetric model. Due to the strong nonlinear hyperelasticity of human soft tissues, the system employs the classical Mooney-Rivlin constitutive equation to describe the strain energy density function W of the tissue surrounding the acupoints: in, These are the first and second principal invariants of the right Cauchy-Green deformation tensor, which characterize the degree of physical deformation of acupoint tissues under macroscopic stretching movements (such as the "Hands Supporting the Sky to Regulate the Three Jiaos" exercise in Baduanjin). Empirical constants for characterizing the properties of soft tissue materials are assigned different preset values ​​according to the tissue type where different acupoints are located (e.g., the junction of thick muscle and tendon). It is the determinant of the deformation gradient tensor, used to characterize the volume change of tissue under pressure; This is an incompressibility parameter.

[0027] Furthermore, after the finite element simulation solver calculates the soft tissue deformation field for each frame, the first principal strain and equivalent stress in the acupoint center region can be extracted by differentiating the strain energy density function. Finally, these data are used as the target Ground Truth (i.e., the preset biomechanical stimulation threshold standard) and jointly serialized and encapsulated with the corresponding action frames and meridian data, stored in a standard action knowledge base. The encapsulated data structure in this action knowledge base can be represented as follows: ,in This represents the skeletal posture at time t; This represents the meridian path at time t; The first principal strain represents the central region of the acupoint. This indicates the equivalent stress in the central region of the acupoint; This indicates the total amount of data.

[0028] It should be noted that, through the above processing, the standard movement knowledge base can break away from the traditional video-based practice method. It not only provides standard spatial movement trajectories but also quantifies the biomechanical essence of traditional Chinese medicine's guiding exercises for relaxing muscles and tendons. This provides a strong data-supported objective diagnostic benchmark for the downstream interactive correction and analysis module to determine whether the user's movements are in place and whether they have achieved a genuine health-preserving effect.

[0029] It should be noted that after obtaining the user's three-dimensional human skeletal model and constructing a standard movement knowledge base containing TCM theoretical parameters, the system faces the challenge of transcending superficial limb movements to analyze the actual manifestations of the user's movement postures in TCM therapeutic effects (i.e., meridian traction and acupoint stimulation). Based on this problem, the present invention constructs an interactive correction analysis module.

[0030] Furthermore, this interactive deviation correction analysis module is configured to match the user's 3D human skeletal model with data in a standard movement knowledge base, and based on a preset evaluation algorithm, to quantitatively analyze the deviation between the user's posture and the standard posture in terms of kinematic morphology, meridian path conformity, and acupoint activation status, generating deviation correction diagnostic information. In this process, since pure joint space coordinate comparison cannot reflect the unity of intention, qi, form, and spirit in traditional Chinese medicine guiding exercises, this module preferably adopts a three-level quantitative evaluation architecture.

[0031] Specifically, at the first level: kinematic morphology evaluation, the system first uses the Dynamic Time Warping (DTW) algorithm to align the user with the standard action in the temporal dimension, and calculates the basic three-dimensional coordinate deviation and rotation angle deviation of each key joint in Euclidean space.

[0032] Specifically, at the second level—the assessment of meridian path conformity—since traditional Chinese medicine guiding exercises emphasize using momentum to guide qi, the extension of the limbs must be able to smoothly stretch the target meridians. Therefore, the preset algorithm for this assessment first uses the user's three-dimensional human skeletal model as the driving force. Utilizing the pre-set meridian-skeletal skin weight parameters in the aforementioned standard knowledge base, it deforms the three-dimensional meridian circulation path model attached to the user's skeleton, thereby calculating the three-dimensional meridian curve path in the user's current real-world state in real time. Subsequently, to measure the morphological similarity between the user's current meridian path and the standard meridian path corresponding to the standard movement, this invention also introduces the discrete Friesian distance algorithm for geometric deviation calculation. It is important to emphasize that, unlike the traditional point-to-point Euclidean distance, the Friesian distance is often figuratively described as "the minimum leash length for a person to lead a dog," making it very suitable for assessing meridian qi and blood circulation paths with continuous direction and temporal characteristics (because meridians are often complex spatial curves with a specific flow direction). Its discretization calculation formula is defined as follows: and in, and These represent the user's current set of meridian control points. and standard meridian control point set The two polygonal curve paths formed by these two curves. This represents the spatial Euclidean distance between control point pairs; This represents the calculated Frescher distance deviation value. The smaller the deviation value, the more it matches the meridian pathways formed in the user's body with the standard TCM Qi and blood circulation trajectory. Conversely, if the value suddenly increases, it indicates that the user may have caused the meridian (such as the Lung Meridian of Hand-Taiyin) to fold or become blocked at some point due to stiff movements or incorrect twisting angles.

[0033] It needs to be explained that in the discretization calculation formula above, the role of min is to select from three valid methods for each node pair (i, j) in the path (i.e., ...). These correspond to "human stops, dog walks", "dog stops, human walks", and "human and dog walk simultaneously" respectively. The goal is to find the minimum rope length required to reach the current node, which is the optimal action synchronization scheme. The purpose of `max` is to lock the maximum rope length during the "human and dog walks" process, as the rope length changes dynamically.

[0034] Specifically, in the third level—acupoint activation status assessment—the concept of "deqi" in Traditional Chinese Medicine (TCM) relies on effective mechanical stimulation of the acupoint. However, performing complex finite element soft tissue stress analysis (FEA) in real-time on user terminal devices is impractical, leading to severe lag. Therefore, this invention introduces a surrogate model in the interactive correction analysis module to quickly estimate the biomechanical stimulation at key acupoint locations. Specifically, this surrogate model is a lightweight neural network based on a multilayer perceptron (MLP), designed to circumvent the high computational overhead of terminal devices. The MLP structure includes an input layer, at least three fully connected hidden layers (using the LeakyReLU activation function), and an output layer. In the model building and training phase, firstly, the training input tensor is extracted and constructed. This input tensor includes the relative rotation matrix (flattened into a one-dimensional vector) of the associated joints around the target acupoint and the instantaneous angular velocity vector of the corresponding limb. Secondly, the first principal strain estimate and equivalent stress estimate of the deep tissue of the target acupoint calculated by offline finite element analysis (FEA) are concatenated into the true label tensor. Finally, supervised training is performed using mean squared error (MSE) as the loss function, and the network weights are continuously updated through the Adam optimizer until the loss converges. In the actual inference phase, the surrogate model uses the relative rotation matrix of the joints around the target acupoint and the limb angular velocity extracted from the kinematic morphology assessment as input features. Through forward propagation, it predicts the first principal strain estimate and equivalent stress estimate of the deep tissue of the target acupoint under the current posture. Subsequently, the system compares the calculated actual stimulus with the key acupoint biomechanical response target Ground Truth stored in the standard movement knowledge base to construct an acupoint activation evaluation formula. in, and To adjust the hyperparameters for the weights of tensile strain and compressive stress; To activate tolerance variance. When When the value is close to 1, it indicates that the user's current movement amplitude precisely reaches the mechanical threshold of stimulating the acupoint, achieving effective intervention in the sense of traditional Chinese medicine. This represents the estimated first principal strain value of the deep tissue of the target acupoint. This represents the estimated equivalent stress of the deep tissues of the target acupoint.

[0035] Furthermore, after completing the quantification at the three levels mentioned above, this multimodal data needs to be transformed into root causes of deviations that users can understand. For example, if the system finds low activation at the Zusanli acupoint, the cause could be insufficient knee flexion or insufficient forward lean. To uncover such complex nonlinear causal relationships, the interactive deviation correction analysis module is also equipped with a diagnostic network based on a multi-head attention mechanism.

[0036] Specifically, the diagnostic network is configured to concatenate temporal kinematic deviation feature vectors, meridian path conformity feature vectors, and acupoint activation feature vectors, mapping them to a traditional query, key, and value matrix. This matrix is ​​then weighted and fused using a classic multi-head attention calculation formula. It's important to note that during the calculation, the multi-head attention formula generates an attention weight distribution matrix. The system extracts the key weight columns corresponding to queries with abnormal meridian path conformity or acupoint activation, and identifies the kinematic node with the highest weight coefficient (i.e., the highest attention score). For example, when the system detects insufficient deformation of the Hand Shaoyin Heart Meridian, if the kinematic node with the highest attention score is the shoulder joint internal rotation angle, the system uses this weight extreme value mapping to identify and output the excessive shoulder joint internal rotation angle deviation as the main root cause of the user's action deviation.

[0037] Furthermore, in this invention, the interactive correction analysis module also integrates a fuzzy logic inference engine. This engine receives the main root cause output by the diagnostic network and, through a fuzzification interface, maps precise numerical deviations to preset fuzzy sets (e.g., mapping an elbow angle error of +15° to a fuzzy membership degree of over-extension). Subsequently, the engine calls a preset TCM theory rule base (containing IF-THEN rules formulated by TCM experts, such as: IF (elbow over-extension) AND (low stimulation of Quchi acupoint) THEN (strategy = slightly bend the elbow, guiding Qi to Quchi with intention)) to perform fuzzy inference and defuzzification operations. This transforms the data conclusions into a correction strategy in natural language that conforms to TCM guidance practices, and packages and distributes the diagnostic information to the multimodal feedback module for voice rendering and holographic visual interaction.

[0038] It should be noted that traditional motion-assistive software typically uses simple skeletal wireframe overlays or preset mechanical voice prompts. This approach lacks interactive warmth and fails to intuitively express abstract concepts such as Qi, blood, and meridians in Traditional Chinese Medicine. Therefore, this invention constructs a multimodal feedback module to overcome this interaction bottleneck.

[0039] Furthermore, the multimodal feedback module includes a virtual digital employee configured to receive the correction diagnosis information generated by the aforementioned interactive correction analysis module, and to provide users with real-time correction guidance that combines traditional Chinese medicine theory by graphically rendering the information on the display interface and outputting voice commands.

[0040] Specifically, in the graphical rendering stage at the visual feedback level, in order to visualize the process of guiding qi through form in traditional Chinese medicine, this invention employs a particle system inspired by computational fluid dynamics (CFD) for visualization. This multimodal feedback module is configured to visualize and render the meridian pathways in the form of dynamic particle flow on the user's three-dimensional virtual avatar (or a similar digital twin mirror image).

[0041] Furthermore, to ensure that the dynamic particle flow accurately reflects the user's movement accuracy, this system establishes a real-time dynamic mapping between particle kinematic characteristics, meridian path conformity, and acupoint activation status. This mapping algorithm pre-defined the particle's position at time... Basic velocity of flow along meridian curves and particle emission density Its dynamic update formula is as follows: in, This represents the baseline velocity constant of Qi and blood at rest; when the user's action height is standardized (discrete Fréchet distance approaches 0), When the value approaches 1, the particle velocity is greatly enhanced; conversely, if the action deviation is large, the particle velocity will be significantly reduced, presenting a visual manifestation of qi stagnation and blood stasis in traditional Chinese medicine theory. This is the velocity gain sensitivity parameter. Maximum particle rendering density; and To control the scaling and bias parameters of the activation threshold; This is an activation function used to smoothly map linear values ​​to the range of 0 to 1.

[0042] In addition to speed and density, the system also changes the particle color space in real time based on the discrete Fréchet distance (e.g., using the HSV color model). When movements are standard and blood circulation is smooth, the particle flow presents a bright and continuous blue-green hue (representing vitality and smooth flow); when a joint bends at an incorrect angle, causing meridian folding and blockage, the particle flow at that point instantly turns into a highly saturated red, accompanied by random scattering of local particles (turbulence effect). This multi-dimensional visual rendering allows users to receive intuitive biofeedback simply by observing the flow of light effects representing their virtual avatar on the screen, without needing extensive knowledge of traditional Chinese medicine anatomy, thus actively fine-tuning their force application and joint posture.

[0043] Furthermore, in the auditory feedback stage of voice command output, relying solely on visual feedback can easily distract users during practice. Therefore, the voice commands output by the virtual digital employee in this multimodal feedback module are not traditional trigger-based recording and playback, but are driven by an integrated language model (LLM).

[0044] Specifically, this language model aims to generate personalized and contextualized guidance utterances. In terms of the model's input settings, the system uses the natural language form of the correction strategy output by the aforementioned fuzzy logic inference engine (e.g., slightly bend the elbow to guide the Qi to the Quchi point) as the core intent; at the same time, it extracts the user's historical interaction data (including the frequency of historical error-prone actions, user age, physical condition identification tags, and the duration of practice today) as the contextual background.

[0045] In addition, to enable the language model to accurately understand specific scenarios, the multimodal feedback module is also equipped with a prompt word construction unit, which is used to serialize and concatenate the above heterogeneous data into a structured joint input condition sequence, and its mapping template is as follows: [System Role]: You are a professional and empathetic traditional Chinese medicine coach; [User Status]: Age {User Age}, Physical Condition {Physical Condition Identification Tag}, Today's Practice Time {Practice Duration}, Fatigue Level Assessment {Fatigue Index}; [Current Deviation]: Historical high-frequency error actions {frequency of historical error-prone actions}, the root cause of the current action deviation is {the main root cause of the diagnostic network output}; [Traditional Chinese Medicine Correction Strategy]: {Correction strategy output by the fuzzy logic engine}; [Generation Requirements]: Please generate a sentence of no more than 30 characters that is conversational and has a calming and encouraging tone.

[0046] Furthermore, this language model employs an autoregressive generative architecture, given the aforementioned joint input condition sequence. Predict word by word and generate the optimal response sequence. Its core conditional probability distribution formula is expressed as: in, The word or phrase indicating a prediction at the current moment; This indicates a previously generated sequence of utterances; This is the model weight matrix obtained after fine-tuning the language model on massive amounts of ancient Chinese medicine texts, health guidance corpora, and anthropomorphic customer service dialogue data.

[0047] It's important to emphasize that this architecture allows the virtual digital employee's communication style to be highly adaptable to different scenarios. For example, when the system detects that an elderly user repeatedly performs the same movement (e.g., looking backwards to relieve neck strain), the language model, based on historical data, determines that the user may have neck stiffness. Instead of outputting a harsh instruction to increase the range of motion, it generates guidance such as, "Uncle Li, we've detected that your neck and shoulders may be a bit tense today. Don't push yourself; turn until you feel a slight stretch. Focusing your attention on the Dazhui acupoint will help you relax," providing more emotional and gradual guidance. This enhances the system's interactive friendliness and user engagement.

[0048] Furthermore, traditional Chinese medicine health preservation is not a one-off correction, but a systematic project that requires long-term commitment and individualized approach. To extend the feedback from a single correction into long-term health management, this invention also includes a personalized course planning module.

[0049] Furthermore, the module integrates a reinforcement learning model to track user performance data recorded by the interaction correction analysis module over a long period of time, and autonomously learns to generate optimal teaching strategies, dynamically adjusting the training content and difficulty of the model.

[0050] Specifically, this reinforcement learning model models the entire long-term health education process as a Markov decision process (MDP). In this reinforcement learning model, the system's component definitions and internal relationships are set as follows: state space ( ): The system in the first The user's comprehensive feature vector is extracted from each training cycle. This vector integrates the user's average achievement rate of actions over the past week, the average flow rate index of each meridian, the estimated value of physical exertion, and the fatigue index actively reported by the user. Action space ( ): Course adjustment strategies that virtual digital employees, as intelligent agents, can adopt. These include, but are not limited to: switching the type of guiding exercise (downgrading from Baduanjin to the more soothing Liuzijue), increasing or decreasing the number of repetitions of specific movements, adjusting the tempo of background music (to guide breathing frequency), or inserting targeted joint warm-up modules before formal training.

[0051] Reward function ( This function constructs a multi-dimensional composite reward function: in, This indicates the increase in acupoint activation and movement accuracy during practice after the course was adjusted. This indicates the consistency of user check-ins and the completion rate of training (to prevent users from giving up due to overly difficult courses). This is a penalty item for potential sports injury risk calculated based on the frequency of non-standard movements; These are the weighting coefficients corresponding to each item.

[0052] Policy Update: Proximal Policy Optimization (PPO) is the preferred algorithm for updating agent network parameters. The PPO algorithm ensures the smoothness of course difficulty adjustments by limiting the magnitude of each policy update. The objective function for the pruning agent is as follows: in, It is an empirical expectation operator for the samples sampled in the t-th training period, used to calculate the statistical average of the pruning surrogate objective function in the PPO algorithm in the corresponding sample space; This represents the probability ratio of implementing specific course adjustments under both the old and new strategies. This is the advantage function, used to evaluate how much better the current adjustment strategy is compared to the average strategy; This is a hyperparameter (usually set to 0.2) used for the cropping ratio to prevent the model from making radical decisions to drastically increase or decrease the difficulty of the course due to a single instance of extremely good or poor user performance. It is a numerical limiting clipping function used to constrain the action probability ratio of the new and old strategies in the PPO algorithm within a preset range, so as to avoid the strategy update amplitude being too large at one time and ensure the stability of the health course adjustment.

[0053] It should be noted that through the application of this reinforcement learning model, the digital employee is no longer a player mechanically executing a fixed schedule, but a TCM therapist with accumulated experience. It can continuously trial and error, and approach the most suitable health intervention path for each individual user, based on their daily physical fluctuations and long-term musculoskeletal adaptation capabilities, thereby achieving intelligent and multimodal interaction in TCM health maintenance assistance. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A digital employee system based on intelligent multimodal interactive traditional Chinese medicine health preservation, characterized in that: include: A 3D pose reconstruction module is configured to receive a monocular video stream of a user's body and reconstruct a 3D human skeleton model representing the user's pose in real time based on the monocular video stream. A standard movement knowledge base, wherein the knowledge base pre-stores standard three-dimensional movement sequences corresponding to traditional Chinese medicine guiding techniques, and each frame of the standard three-dimensional movement sequence is associated with a digitized three-dimensional meridian circulation path model and preset key acupoint biomechanical response data; An interactive correction analysis module is configured to match the user's real-time three-dimensional human skeleton model with the data in the standard action knowledge base, and based on a preset evaluation algorithm, to quantitatively analyze the deviation between the user's posture and the standard posture from the aspects of kinematic morphology, meridian path conformity and acupoint activation status, and generate correction diagnosis information. A multimodal feedback module includes a virtual digital employee configured to receive the corrective diagnostic information and provide users with real-time corrective guidance that integrates with traditional Chinese medicine theory by graphically rendering the information on a display interface and outputting voice commands.

2. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, The 3D pose reconstruction module integrates a spatiotemporal graph convolutional network, which is used to extract and generate a preliminary 3D human skeleton coordinate sequence from the monocular video stream. The 3D pose reconstruction module also includes a physical constraint optimization unit, which is used to apply human biomechanical constraints to the preliminary 3D human skeleton coordinate sequence and output a final 3D human skeleton model that conforms to physical laws.

3. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, The biomechanical response data of the key acupoints include: Finite element analysis was performed using a standard human body model to obtain target values ​​for stress and strain at specific acupoints under standard movements.

4. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, When evaluating the conformity of the meridian path based on a preset evaluation algorithm, the following are included: Using the user's real-time 3D human skeleton model as the driving force, the 3D meridian circulation path model attached to it is deformed to obtain the user's current meridian path. The discrete Frescher distance algorithm is used to calculate the geometric deviation between the user's current meridian path and the corresponding standard meridian path in the standard action knowledge base.

5. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, When evaluating the activation status of the acupoints based on a preset evaluation algorithm, the following is included: Using a proxy model, the biomechanical stimulation of key acupoints is quickly estimated based on the user's three-dimensional human skeletal model, and the stimulation is compared with the biomechanical response data of the key acupoints stored in the standard movement knowledge base.

6. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, The interactive correction analysis module also includes a diagnostic network based on a multi-head attention mechanism, which is configured as follows: The analysis results of the kinematic morphology, meridian path conformity, and acupoint activation status are weighted and processed to identify and output the main root causes of user action deviations.

7. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 6, characterized in that, The system also includes a fuzzy logic reasoning engine, which is configured as follows: This is used to receive the main root cause output by the diagnostic network, and to convert the main root cause into a correction strategy in natural language form according to a preset TCM theory rule base, for use by the multimodal feedback module.

8. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, The multimodal feedback module is configured as follows during graphical rendering: On the user's three-dimensional virtual avatar, the meridian path is visualized and rendered in the form of a dynamic particle flow. The color, speed, or density of the particle flow changes in real time according to the meridian path conformity result calculated by the interactive correction analysis module.

9. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, The virtual digital employee in the multimodal feedback module outputs voice commands driven by a language model. The language model receives the corrective diagnostic information and combines it with the user's historical interaction data to generate personalized and contextualized guidance.

10. The intelligent multimodal interactive TCM health preservation auxiliary digital employee system as described in claim 1, characterized in that, The system also includes a personalized course planning module, which integrates a reinforcement learning model to track user performance data recorded by the interaction correction analysis module over a long period of time, and autonomously learn and generate an optimal teaching strategy based on the user performance data, dynamically adjusting the training content and difficulty of the model.