Multi-modal pose evaluation method and device

Through the multimodal position evaluation method, combined with the data of video, pressure sensors and fiber optic sensors, the problem that a single sensor in the prior art is difficult to evaluate the multidimensional position of the human body, and accurate evaluation and real-time feedback of the human position are achieved.

CN120093285AActive Publication Date: 2025-06-06SUZHOU INST OF BIOMEDICAL ENG & TECH CHINESE ACADEMY OF SCI +1

Patent Information

Application Number
CN202510291763.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-06
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing posture monitoring system only relies on a single sensor, making it difficult to comprehensively evaluate the multi-dimensional characteristics of human body posture, resulting in the inability to accurately evaluate the person's posture in different scenarios.

Method used

The multimodal pose evaluation method is used to collect video data through the camera, the pressure sensor collects pose-related pressure distribution data, and the optical fiber sensor measures the curved data of the spine, and perform data fusion, feature extraction and deep learning model training to achieve weighted fusion and evaluation of multiple features.

Benefits of technology

It can accurately evaluate people's poses in different scenarios, improve the accuracy and robustness of the assessment, and provide real-time pose feedback and improvement suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120093285A_ABST
    Figure CN120093285A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal pose evaluation method and device, and belongs to the field of health monitoring, and the method comprises the steps: collecting a dynamic video of a human body through a camera, and obtaining video data; collecting pressure distribution data related to poses through pressure sensors arranged on the cushion and the vest; an optical fiber sensor arranged on the vest is attached to the spine to measure bending data of the spine; carrying out unified time resolution and spatial registration on the data, carrying out feature extraction on the three kinds of data, and carrying out weighted fusion on the extracted features; and designing and training a deep learning model: inputting the multiple features after weighted fusion into the deep learning model, and outputting an evaluation result, and through the steps, the poses of the person in different scenes can be accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of health monitoring, and in particular to a multi-modal posture assessment method and device. Background Art

[0002] As society pays more and more attention to children's health, problems such as scoliosis and posture disorders caused by bad standing, sitting and exercise habits have become the focus of public health. At present, there are some posture monitoring systems on the market, but these systems often rely on only a single sensor, such as a camera or a pressure sensor, which makes it difficult to comprehensively evaluate the multi-dimensional characteristics of posture, and therefore cannot accurately evaluate the posture of people in different scenarios. Summary of the invention

[0003] In order to overcome the deficiencies of the prior art, one of the objectives of the present invention is to provide a multimodal posture assessment method that can accurately assess the posture of a person in different scenarios.

[0004] In order to overcome the deficiencies of the prior art, a second object of the present invention is to provide a multimodal posture assessment device that can accurately assess the posture of a person in different scenarios.

[0005] One of the purposes of the present invention is achieved by the following technical solution:

[0006] A multimodal posture assessment method comprises the following steps:

[0007] Data collection: The camera collects dynamic videos of the human body to obtain video data; the pressure sensors installed on the seat cushion and vest collect posture-related pressure distribution data; the optical fiber sensor installed on the vest fits the spine to measure the curvature data of the spine;

[0008] Data fusion: Use multiple spline interpolation to time-align the video data and pressure distribution data, downsample the bending data to match the pressure distribution data to unify the data time resolution of the camera, pressure sensor and fiber optic sensor; establish a global posture coordinate system, use the affine transformation method to map the pressure sensor grid to the skeleton key point topology structure, calculate the force conditions of the corresponding area, optimize the alignment error between the coordinate systems of the camera, pressure sensor and fiber optic sensor, and improve the spatial matching accuracy;

[0009] Feature extraction: Use deep neural networks to estimate the posture of video data, extract the coordinates of key points of the human body, calculate the inclination angle of the trunk, the offset of the head, and the movement frequency; calculate the pressure center coordinates, asymmetry index, longitudinal pressure gradient, and transverse pressure fluctuation based on the pressure distribution data; calculate the curvature of each position of the spine and the spinal curvature angle based on the spinal curvature data; calculate the attention weights of different extracted features, and perform weighted fusion of the features;

[0010] Design and train deep learning models: Use composite loss functions to build deep learning models, optimize classification and regression tasks, and prevent modal collapse; collect multiple posture data including incorrect postures and normal postures, combine video key point detection, pressure distribution map analysis and fiber curvature signal calculation to annotate the data using multimodal fusion, and use a staged training method to train the model;

[0011] Result evaluation: The weighted fusion of multiple features is input into the deep learning model, and the evaluation results are output.

[0012] Furthermore, in the data fusion step, it also includes adopting a time window sliding mechanism to perform linear interpolation on the video data, the pressure distribution data and the spinal curvature data within a short time scale to reduce mutation errors.

[0013] Furthermore, in the data fusion step, a global posture coordinate system is established specifically as follows: the seventh cervical vertebra is taken as the origin, the spine direction is the Z axis, the shoulder direction is the X axis, and the front-back direction is the Y axis.

[0014] Furthermore, in the feature extraction step, the trunk inclination angle are the coordinates of the left hip key point, is the coordinate of the right hip key point; head offset P nose is the coordinate point of nose tip, is the coordinate of the center point of the shoulder; the motion frequency is the main motion mode frequency extracted by analyzing the displacement spectrum of the key points of the shoulder.

[0015] Furthermore, in the feature extraction step, the pressure center coordinates p i represents the pressure value measured by the i-th pressure sensor; x i ,y i Indicates the coordinates of the sensor; asymmetry index Where ∑p left is the total pressure of the left pressure sensor, ∑p right is the total pressure of the right pressure sensor, ∑p total is the total pressure; the longitudinal pressure gradient ΔP is the pressure difference between the upper and lower pressure sensors, Δh is the vertical distance between the two pressure sensors; lateral pressure fluctuation p i is the pressure value of the i-th sensor, μ is the mean pressure of all sensors, and N is the number of sensors.

[0016] Furthermore, in the feature extraction step, the curvature of each position of the spine Δλ B is the drift of the central Bragg wavelength of FBG, λ B is the initial Bragg wavelength of FBG, S is the strain sensitivity factor; the spine bending angle α = ∫κ(s)ds, s is the arc length coordinate along the length direction of the spine, and κ(s) is the local curvature of the spine at different positions.

[0017] Furthermore, in the step of designing and training a deep learning model, three sub-models are designed. The three sub-models correspond to dynamic video, pressure sensor, and fiber optic sensor, respectively. The three sub-models respectively output the trust distribution of posture categories, and decision fusion is performed through DS synthesis rules.

[0018] Furthermore, in the step of designing and training the deep learning model, the composite loss function is L = λ 1 L cls +λ 2 L reg +λ 3 L ortho , where λ 1 is the weight parameter of classification loss, L cls is the classification loss, λ 2 is the weight parameter of regression loss, L reg is the regression loss, λ 3 is the weight parameter of the orthogonal loss, L ortho is the orthogonal loss.

[0019] Furthermore, in the step of designing and training the deep learning model, the model is trained by a phased training method as follows: the first phase performs single-modal pre-training, in which the video branch is initialized using HRNet to improve the accuracy of key point detection; the second phase fixes the backbone network for feature extraction of each modality, and only optimizes the fusion layer to ensure cross-modal information alignment; the third phase performs end-to-end fine-tuning, and adjusts the learning rate to 1e -4 , in order to balance the convergence speed and optimization accuracy.

[0020] Furthermore, in the result evaluation step, the user's real-time posture is dynamically displayed through a 3D human body model or bone point projection, color coding is used to mark the parts of bad posture, and the pressure distribution is displayed through a heat map to highlight areas with excessive pressure. The long-term posture change trend is tracked in combination with a time series curve, and the spinal curvature change curve is displayed in real time.

[0021] The second object of the present invention is achieved by adopting the following technical solution:

[0022] A multi-modal posture assessment device, used to implement any of the above multi-modal posture assessment methods, the multi-modal posture assessment device comprising

[0023] The camera collects dynamic videos of the human body and obtains video data;

[0024] Pressure sensors, which are arranged on the seat cushion and the vest to collect posture-related pressure distribution data;

[0025] An optical fiber sensor is disposed on the vest and fits the spine to measure the curvature data of the spine;

[0026] The processor analyzes the collected video data, pressure distribution data, and bending data to evaluate the user's posture.

[0027] Compared with the prior art, the multimodal posture assessment method of the present invention acquires video data by capturing dynamic video of the human body through a camera; acquires posture-related pressure distribution data through pressure sensors arranged on a seat cushion and a vest; measures spinal curvature data by fitting the spinal column through an optical fiber sensor arranged on the vest; performs unified temporal resolution and spatial registration on the data, extracts features from the three types of data, and performs weighted fusion on the extracted features; designs and trains a deep learning model: inputs the weighted fused multiple features into the deep learning model, and outputs an assessment result. Through the above steps, the posture of a person in different scenarios can be accurately assessed. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flow chart of the multi-modal posture assessment method of the present invention;

[0029] Figure 2 is a schematic diagram of a vest in the multi-modal posture assessment method of the present invention;

[0030] Figure 3 is a schematic diagram of a seat cushion in the multi-modal posture assessment method of the present invention;

[0031] Figure 4 Schematic diagram of skeleton key point extraction in the multimodal pose assessment method of the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0034] See also Figure 1 The present invention provides a multi-modal posture assessment method, comprising the following steps:

[0035] Data collection: The camera collects dynamic videos of the human body to obtain video data; the pressure sensors installed on the seat cushion and vest collect posture-related pressure distribution data; the optical fiber sensor installed on the vest fits the spine to measure the curvature data of the spine;

[0036] Data fusion: Use multiple spline interpolation to time-align the video data and pressure distribution data, downsample the bending data to match the pressure distribution data to unify the data time resolution of the camera, pressure sensor and fiber optic sensor; establish a global posture coordinate system, use the affine transformation method to map the pressure sensor grid to the skeleton key point topology structure, calculate the force conditions of the corresponding area, optimize the alignment error between the coordinate systems of the camera, pressure sensor and fiber optic sensor, and improve the spatial matching accuracy;

[0037] Feature extraction: Use deep neural networks to estimate the posture of video data, extract the coordinates of key points of the human body, calculate the inclination angle of the trunk, the offset of the head, and the movement frequency; calculate the pressure center coordinates, asymmetry index, longitudinal pressure gradient, and transverse pressure fluctuation based on the pressure distribution data; calculate the curvature of each position of the spine and the spinal curvature angle based on the spinal curvature data; calculate the attention weights of different extracted features, and perform weighted fusion of the features;

[0038] Design and train deep learning models: Use composite loss functions to build deep learning models, optimize classification and regression tasks, and prevent modal collapse; collect multiple posture data including incorrect postures and normal postures, combine video key point detection, pressure distribution map analysis and fiber curvature signal calculation to annotate the data using multimodal fusion, and use a staged training method to train the model;

[0039] Result evaluation: The weighted fusion of multiple features is input into the deep learning model, and the evaluation results are output.

[0040] Specifically, the data collection steps are as follows:

[0041] A binocular camera is used to capture dynamic video of the human body in order to extract skeleton key points. In this embodiment, the user is a child, and evaluating the child's posture can help discover and evaluate health conditions such as scoliosis and posture disorders in children. The current video frame is obtained from the binocular camera and the video frame sequence is updated; the video frame sequence is sent to the human posture estimation and behavior recognition module, and the 3D coordinates of the key points are obtained by three-dimensional reconstruction of the left and right lens images to obtain video data.

[0042] Please continue reading Figure 2 as well as Figure 3 The pressure sensors are set on the back of the vest and the seat cushion. Due to different sitting postures, the pressure at different positions of the back and the seat cushion will be different. Collecting the pressure at different positions of the back and the seat cushion can more accurately evaluate the posture of the human body.

[0043] The fiber optic sensor is set in the middle of the correction vest, stacked in a double S shape, fitting the spine, and is used to measure the curvature of the spine.

[0044] The multimodal posture assessment method of the present invention further includes a data preprocessing step, which is located after the data acquisition step and before the data fusion step. The data preprocessing step is specifically: collecting posture-related pressure distribution data through pressure sensors set on the seat cushion and the vest; video data is processed by background elimination and denoising, skeleton point extraction (such as Figure 4 As shown in the figure), time series data generation and image standardization are used to extract dynamic trajectory information of key parts; pressure data is detected through signal filtering, partition pressure analysis and time series trend analysis to detect changes in pressure distribution in different regions and their potential adverse posture effects. For spinal curvature monitoring based on fiber optic sensors, the system additionally introduces fiber Bragg grating (FBG) sensing technology to obtain spinal curvature changes in real time. By fusing fiber optic data with video bone point information, using adaptive filtering and nonlinear regression algorithms, the detection accuracy of spinal deformation is further improved, providing more comprehensive support for posture correction.

[0045] The data fusion steps include unifying the temporal resolution and spatial registration.

[0046] Since different sensors have different sampling rates, the time resolution needs to be unified:

[0047] The video data (30fps) and pressure data (100Hz) are time-aligned using cubic spline interpolation to improve data continuity and accuracy. The fiber data (200Hz) is downsampled to 100Hz to match the pressure sensor data, and anti-aliasing filtering is used to reduce information loss. A timing window sliding mechanism is used to perform linear interpolation within a short time scale (250ms) to reduce mutation errors.

[0048] Spatial registration: In order to ensure that different modal data are fused and analyzed in the same coordinate system, a global posture coordinate system is established, with the seventh cervical vertebra (C7) as the origin, the longitudinal direction of the vest (spine direction) is defined as the Z axis, the shoulder direction is the X axis, and the front-to-back direction is the Y axis. The affine transformation method is used to map the pressure sensor grid to the skeleton key point topology structure and calculate the force situation of the corresponding area. The ICP algorithm is used to optimize the alignment error between different sensor coordinate systems to improve the spatial matching accuracy.

[0049] The feature extraction step includes feature extraction of video data, feature extraction of pressure distribution data, and feature extraction of spinal curvature data.

[0050] When extracting features from video data, an improved HRNet-W48 deep neural network is used for posture estimation, extracting 26 key points of the human body. The key point coordinate data is low-pass filtered to reduce the impact of jitter noise on feature calculation. The trunk inclination angle, head offset, and movement frequency features are calculated. The trunk inclination angle measures the alignment of the spine, the head offset reflects the asymmetry of the head posture, and the movement frequency features are used to evaluate long-term posture stability. Trunk inclination angle are the coordinates of the left hip key point, is the coordinate of the right hip key point; head offset P nose is the coordinate point of nose tip, is the coordinate of the center point of the shoulder; the movement frequency uses fast Fourier transform (FFT) to analyze the displacement spectrum of the key points of the shoulder and extract the main movement mode frequency.

[0051] The purpose of extracting the characteristics of pressure distribution data is to use the pressure sensor data in the seat cushion and vest to evaluate the characteristics of force distribution. The extraction of pressure distribution data includes calculating the pressure center coordinates, asymmetry index, longitudinal pressure gradient, and transverse pressure fluctuation. The pressure center coordinates are used to analyze the stability of the center of gravity in the sitting posture, the asymmetry index is used to measure the balance of the force distribution on the left and right sides, the longitudinal pressure gradient is used to evaluate the changes in the back pressure along the spine, and the transverse pressure fluctuation is used to measure the force fluctuations in different areas of the back. Pressure center coordinates p i represents the pressure value measured by the i-th pressure sensor; x i ,y i Indicates the coordinates of the sensor; asymmetry index Where ∑p left is the total pressure of the left pressure sensor, ∑p right is the total pressure of the right pressure sensor, ∑p total is the total pressure; the longitudinal pressure gradient ΔP is the pressure difference between the upper and lower pressure sensors, Δh is the vertical distance between the two pressure sensors; lateral pressure fluctuation p i is the pressure value of the i-th sensor, μ is the mean of all sensor pressures, and N is the number of sensors. Wavelet Transform is used to analyze the pressure data and decompose the signals in different frequency bands to identify long-term bad posture patterns.

[0052] When extracting the features of the spinal curvature data, the optical fiber sensor is a fiber Bragg grating (FBG) sensor. The FBG optical fiber demodulation algorithm is used to calculate the curvature of each position of the spine and the spinal curvature angle. Δλ B is the drift of the central Bragg wavelength of the FBG, reflecting the strain effect caused by the deformation of the optical fiber; B is the initial Bragg wavelength of FBG, S is the strain sensitivity factor, S depends on the fiber material and the writing method; the spine bending angle α=∫κ(s)ds, s is the arc length coordinate along the length of the spine, and κ(s) is the local curvature of the spine at different positions. Based on the curvature-torsion model, the three-dimensional curve of the spine is constructed. The local coordinate system of the spine is solved by the Frenet-Serret formula to analyze the spine morphology. The video skeleton features and pressure distribution features are complementary, so the self-attention mechanism is used for feature fusion to enhance the association between multimodal information. The attention weights of different modal features are calculated for weighted fusion. The MLP (multi-layer perceptron) uses a two-layer fully connected network with ReLU as the activation function to extract cross-modal information.

[0053] The specific steps for designing and training a deep learning model are:

[0054] MTL-GCN (Multi-Task Graph Convolutional Network) is used to model the human skeleton relationship through graph neural network, and combined with cross-modal feature learning to achieve high-precision posture classification and spinal morphology prediction.

[0055] Using a composite loss function to optimize classification and regression tasks and prevent mode collapse:

[0056] L=λ 1 L cls +λ 2 L reg +λ 3 L ortho

[0057] λ 1 is the weight parameter of classification loss, λ 1 Control the impact of classification loss on the overall optimization. L clsFor classification loss, the Focal Loss method is used to improve classification accuracy. 2 is the weight parameter of regression loss, λ 2 Controls the contribution of regression loss. L reg Huber Loss (δ = 1.0) is used as the regression loss to improve the robustness of the spinal curvature regression. 3 is the weight parameter of the orthogonal loss, which controls the influence of the orthogonal loss. ortho Orthogonality loss: It is used to prevent modal collapse, ensure that different modalities (such as sEMG signals and visual features) remain independent in the feature space, and avoid information loss caused by redundant information. Through the orthogonality constraint, the features of different modalities are forced to remain orthogonal mathematically, thereby improving the effectiveness of multimodal fusion.

[0058] Three sub-models are designed, corresponding to dynamic video, pressure sensor and fiber optic sensor respectively. The three sub-models output the trust distribution of posture categories respectively, and the decision fusion is performed through DS synthesis rules.

[0059] The model training scheme is

[0060] In terms of data, the data set includes posture data of 200 adolescents aged 10-16, covering 7 typical incorrect postures and normal postures. Among them, the definition of incorrect posture categories is based on biomechanical and clinical standards, such as neck forward tilt angle greater than 15°, scoliosis Cobb angle greater than 10°, pelvic posterior tilt caused by sacral angle less than 25°, etc. Data annotation adopts multimodal fusion method, combined with video key point detection, pressure distribution map analysis and fiber curvature signal calculation to ensure accurate classification of posture categories.

[0061] In order to improve the generalization ability of the model, data enhancement is performed on different modal data. For video data, a random occlusion strategy is used to simulate view occlusion, and Gaussian noise (σ=0.01) is added to enhance robustness, while time series interpolation is performed to simulate frame rate changes; for pressure data, finite element analysis is used to simulate elastic deformation caused by different body shapes, and random sensor failure simulation is introduced to improve system robustness; for optical fiber data, curvature perturbation (±5%) is added to simulate measurement errors.

[0062] The training strategy adopts a phased training method to ensure the stability and effectiveness of the model when fusing multimodal information. The first stage is single-modal pre-training, in which the video branch is initialized using HRNet to improve the accuracy of key point detection; the second stage fixes the backbone network for feature extraction of each modality and only optimizes the fusion layer to ensure cross-modal information alignment; the third stage performs end-to-end fine-tuning and adjusts the learning rate to 1e -4, in order to balance the convergence speed and optimization accuracy.

[0063] In terms of model compression, knowledge distillation and quantization techniques are used to improve the deployment efficiency of the model. During the knowledge distillation process, the teacher model is a complete MTL-GCN (18.7M parameters), and the student model is a lightweight LightGCN (4.2M parameters). The distillation temperature is set to 3, and the Kullback-Leibler (KL) divergence loss function is used for training to ensure that the student model can effectively inherit the knowledge of the teacher model. In addition, during the quantization deployment stage, the model is converted from FP32 to INT8, the quantization calibration set contains 200 samples, and is finally deployed to the TensorRT engine to improve the inference speed and energy efficiency.

[0064] The classification task uses classification accuracy as the main measurement criterion, and the spinal posture regression task uses the mean square error (MAE) of the spinal curvature as the evaluation indicator. At the same time, the end-to-end delay is taken into account, and the goal is to make the real-time performance of the entire system less than 200ms. In addition, the number of model parameters and power consumption (mJ / inference) are used as auxiliary evaluation indicators to ensure that the system is suitable for actual deployment. OpenPose+Random Forest, MediaPipe+LSTM, and unimodal GCN are selected as baseline models for comparison to verify the superiority of this method. At the same time, ablation experiments are carried out to analyze the contribution of different modules.

[0065] The specific result evaluation steps are as follows: input the weighted fusion of multiple features into the deep learning model and output the evaluation results. In the final decision stage, the classification results of different modalities may be uncertain. The Dempster-Shafer (DS) evidence theory is used for multimodal decision fusion to improve the credibility of the final prediction.

[0066] Feedback posture analysis results to users in a friendly and intuitive way, and provide specific suggestions and long-term trend tracking functions to help users effectively improve posture health. In terms of data visualization, the module dynamically displays the user's real-time posture through 3D human body models or bone point projections, and uses color coding (such as green for normal and red for abnormal) to intuitively mark the parts of bad posture. At the same time, the pressure distribution is displayed through heat maps, highlighting the areas with excessive pressure, and combining time series curves to track the trend of long-term posture changes. Display the spinal curvature change curve in real time. In addition, the historical data review function supports daily, weekly or monthly posture score trend charts to provide users with a comprehensive understanding of the improvement effect. In terms of user operation functions, the system automatically generates personalized suggestions based on the evaluation results, such as "adjusting sitting posture to reduce back pressure", and presents them in various forms such as text prompts, voice broadcasts or video demonstrations; users can also set improvement goals (such as "maintaining a good posture for more than 80% of the time"), and the system will track the completion of the goals and provide feedback. To enhance the user experience, the system issues reminders through pop-ups, prompts or vibrations, and users can customize the frequency and mode of reminders according to personal needs.

[0067] A multimodal posture assessment method also includes a correction step, which is located after the result evaluation step. In terms of real-time posture guidance, the system reminds the user in real time through voice feedback, such as prompting "Please sit up straight", and supports multi-language mode to meet the needs of users with different language backgrounds. At the same time, the interface will synchronously display the action correction animation, such as showing in detail how to adjust the shoulder or waist position, and provide relevant popular science knowledge to help users understand the hazards of bad posture and the importance of correction, and enhance the motivation to improve. In addition, the system has designed a vibration reminder function in key areas (such as seats or wearable devices). When a bad posture is detected, a slight vibration is triggered to gently remind the user to make timely adjustments.

[0068] In terms of equipment adjustment, the correction execution module works in conjunction with intelligent hardware. For example, through the electric adjustment function of the smart seat, the seat height, back angle and lumbar support are automatically adjusted to help users achieve an ideal ergonomic sitting posture. In addition, the module supports the integration of wearable correction devices, such as smart correction straps or pressure-sensing clothing, which can monitor the user's posture in real time and provide physical support or correction force as needed. The support strength can be dynamically adjusted according to the user's movement changes to ensure the comfort and effectiveness of the correction process.

[0069] In order to enhance the user's initiative and health awareness, the system also provides personalized exercise suggestions and guidance. When the user maintains the same posture for a long time, the system will push a reminder and suggest that the user do appropriate stretching exercises to relieve muscle tension. Combined with mobile phones or wearable devices, users can learn the correct stretching and exercise methods through the guidance videos played. Through the combination of these functions, the correction execution module realizes comprehensive support from real-time feedback to physical intervention, helping users gradually develop good posture habits and maintain long-term health.

[0070] The multimodal posture assessment method of the present application uses the dynamic video of the camera, the pressure distribution data of the pressure sensor, and the spinal curvature data of the optical fiber sensor. It is analyzed based on multi-dimensional data, and a hierarchical fusion architecture is adopted. The data level, feature level and decision level fusion are combined to improve the complementarity of multimodal information and enhance the robustness of posture recognition. At the data level, the signals of the pressure sensor and the fiber Bragg grating (FBG) sensor are fused in a unified coordinate system to enhance the spatial consistency of spinal curvature information. The vest coordinate system is adopted, and the seventh cervical vertebra (C7) is taken as the origin to establish a local spinal coordinate system. The spinal curvature detected by the optical fiber sensor and the longitudinal pressure gradient detected by the pressure sensor are weighted and fused. The video skeleton features and the pressure distribution features are complementary, so the self-attention mechanism is used for feature fusion to enhance the association between multimodal information. The attention weights of different modal features are calculated and weighted fusion is performed. The MLP (multi-layer perceptron) adopts a two-layer fully connected network with the activation function of ReLU to extract cross-modal information. In the final decision stage, the classification results of different modalities may be uncertain. The Dempster-Shafer (DS) evidence theory is used to perform multimodal decision fusion to improve the credibility of the final prediction. Three sub-models (video, pressure, and fiber) are set to output the trust distribution of posture categories respectively, and decision fusion is performed through the DS synthesis rule. Bayesian optimization is used to adjust the confidence threshold of the fusion strategy to improve the fusion accuracy. Through the above steps, the posture of people in different scenarios can be accurately evaluated.

[0071] The present application also relates to a multi-modal posture evaluation device, which is used to implement the above-mentioned multi-modal posture evaluation method. The multi-modal posture evaluation device includes

[0072] The camera collects dynamic videos of the human body and obtains video data;

[0073] Pressure sensors are installed on the seat cushion and vest to collect posture-related pressure distribution data;

[0074] Fiber optic sensor: The fiber optic sensor is set on the vest, and the fiber optic sensor is attached to the spine to measure the curvature data of the spine;

[0075] The processor analyzes the collected video data, pressure distribution data, and bending data to evaluate the user's posture.

[0076] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which are equivalent modifications and improvements made to the above embodiments based on the essential technology of the present invention, and all of them belong to the protection scope of the present invention.

Claims

1. A multimodal posture assessment method, characterized in that: The following steps are involved: Data collection: The camera collects dynamic videos of the human body to obtain video data; the pressure sensors installed on the seat cushion and vest collect posture-related pressure distribution data; the optical fiber sensor installed on the vest fits the spine to measure the curvature data of the spine; Data fusion: The video data and the pressure distribution data are time-aligned using multiple spline interpolation, and the bending data is downsampled to match the pressure distribution data to unify the time resolution of the camera, pressure sensor, and optical fiber sensor data; Establish a global posture coordinate system, use the affine transformation method to map the pressure sensor grid to the skeleton key point topology structure, calculate the force situation of the corresponding area, optimize the alignment error between the coordinate systems of the camera, pressure sensor and fiber optic sensor, and improve the spatial matching accuracy; Feature extraction: Use deep neural networks to estimate the posture of video data, extract the coordinates of key points of the human body, calculate the inclination angle of the trunk, the offset of the head, and the movement frequency; calculate the pressure center coordinates, asymmetry index, longitudinal pressure gradient, and transverse pressure fluctuation based on the pressure distribution data; calculate the curvature of each position of the spine and the spinal curvature angle based on the spinal curvature data; Calculate the attention weights of different extracted features and perform weighted fusion of the features; Design and train deep learning models: Use composite loss functions to build deep learning models, optimize classification and regression tasks, and prevent modal collapse; collect multiple posture data including incorrect postures and normal postures, combine video key point detection, pressure distribution map analysis and fiber curvature signal calculation to annotate the data using multimodal fusion, and use a staged training method to train the model; Result evaluation: The weighted fusion of multiple features is input into the deep learning model, and the evaluation results are output.

2. The multimodal posture assessment method according to claim 1, characterized in that: The data fusion step also includes adopting a time window sliding mechanism to perform linear interpolation on the video data, the pressure distribution data and the spinal curvature data within a short time scale to reduce mutation errors.

3. The multimodal posture assessment method according to claim 1, characterized in that: In the data fusion step, the global posture coordinate system is established specifically as follows: the seventh cervical vertebra is taken as the origin, the spine direction is the Z axis, the shoulder direction is the X axis, and the front-back direction is the Y axis.

4. The multimodal posture assessment method according to claim 1, characterized in that: In the feature extraction step, the trunk inclination angle are the coordinates of the left hip key point, is the coordinate of the right hip key point; head offset P nose is the coordinate point of nose tip, is the coordinate of the center point of the shoulder; the motion frequency is the main motion mode frequency extracted by analyzing the displacement spectrum of the key points of the shoulder.

5. The multimodal posture assessment method according to claim 1, characterized in that: In the feature extraction step, the pressure center coordinates p i represents the pressure value measured by the i-th pressure sensor; x i ,y i Indicates the coordinates of the sensor; Asymmetric Index Where ∑p left is the total pressure of the left pressure sensor, ∑p right is the total pressure of the right pressure sensor, ∑p total is the total pressure; the longitudinal pressure gradient ΔP is the pressure difference between the upper and lower pressure sensors, Δh is the vertical distance between the two pressure sensors; lateral pressure fluctuation p i is the pressure value of the i-th sensor, μ is the mean pressure of all sensors, and N is the number of sensors.

6. The multimodal posture assessment method according to claim 1, characterized in that: In the feature extraction step, the curvature of each position of the spine Δλ B is the drift of the central Bragg wavelength of FBG, λ B is the initial Bragg wavelength of FBG, S is the strain sensitivity factor; the spine bending angle α = ∫κ(s)ds, s is the arc length coordinate along the length direction of the spine, and κ(s) is the local curvature of the spine at different positions.

7. The multimodal posture assessment method according to claim 1, characterized in that: In the step of designing and training the deep learning model, three sub-models are designed. The three sub-models correspond to dynamic video, pressure sensor and fiber optic sensor respectively. The three sub-models respectively output the trust distribution of posture categories, and decision fusion is performed through DS synthesis rules.

8. The multimodal posture assessment method according to claim 1, characterized in that: In the step of designing and training the deep learning model, the composite loss function is L = λ1L cls +λ2L reg +λ3L ortho , where λ1 is the weight parameter of classification loss, L cls is the classification loss, λ2 is the weight parameter of regression loss, L reg is the regression loss, λ3 is the weight parameter of the orthogonal loss, L ortho is the orthogonal loss.

9. The multimodal posture assessment method according to claim 1, characterized in that: In the step of designing and training the deep learning model, the model is trained by a phased training method as follows: in the first phase, single-modal pre-training is performed, in which the video branch is initialized using HRNet to improve the accuracy of key point detection; in the second phase, the backbone network for feature extraction of each modality is fixed, and only the fusion layer is optimized to ensure cross-modal information alignment; The third stage is end-to-end fine-tuning, and the learning rate is adjusted to 1e -4 , in order to balance the convergence speed and optimization accuracy.

10. The multimodal posture assessment method according to claim 1, characterized in that: In the result evaluation step, the user's real-time posture is dynamically displayed through a 3D human body model or bone point projection, color coding is used to mark the parts of bad posture, and the pressure distribution is displayed through a heat map to highlight the areas with excessive pressure. The long-term posture change trend is tracked in combination with a time series curve, and the spinal curvature change curve is displayed in real time.

11. A multi-modal posture assessment device, used to implement the multi-modal posture assessment method according to any one of claims 1 to 10, characterized in that: The multi-modal posture assessment device comprises The camera collects dynamic videos of the human body and obtains video data; Pressure sensors, which are arranged on the seat cushion and the vest to collect posture-related pressure distribution data; An optical fiber sensor is disposed on the vest and fits the spine to measure the curvature data of the spine; The processor analyzes the collected video data, pressure distribution data, and bending data to evaluate the user's posture.

Citation Information

Patent Citations

  • Deep learning sitting posture measurement and detection method based on monocular camera

    CN116469174A

  • Sitting posture recognition method and system based on deep learning

    CN116645721A

  • Office chair sitting posture health detection method and system based on pressure distribution

    CN118975794A

  • Systems and methods for evaluation of scoliosis and kyphosis

    US20190239797A1

  • Human posture detection

    US20240193981A1

Cited By

  • Intelligent quantitative evaluation system for teenager scoliosis fused with multi-modal imaging

    CN120381281A

  • Method and device for generating thermodynamic diagram of scoliosis orthosis

    CN121587894A

  • Spine curvature monitoring method of intelligent seat and vehicle-mounted health cloud system

    CN121590559A

  • A spinal curvature monitoring method of an intelligent seat and a vehicle-mounted health cloud system

    CN121590559B

  • Intelligent assessment method for cervical lateral deviation based on multi-modal fusion and deep learning

    CN121746355A