Portable posture detection and gait analysis method and system

Through a portable attitude detection system, combined with multi-spectral camera array and embedded heterogeneous computing architecture, the problems of heavy equipment, poor environmental adaptability and unstable target tracking are solved, efficient and accurate posture and gait analysis are achieved, and intuitive visual reports are generated, suitable for multi-scene applications.

CN120108043BActive Publication Date: 2025-09-02BEIJING SCI & TECH PATENT OFFICE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510586640.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-02
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing attitude detection and gait analysis system has bulky equipment, poor environmental adaptability, unstable target tracking, single feedback method, slow processing speed, difficult to achieve real-time analysis, and low detection accuracy in complex environments.

Method used

The foldable multispectral camera array, embedded heterogeneous computing architecture (NPU+GPU collaboration), dynamic model cropping strategy, dual-channel sensing fusion technology, spatiotemporal attention mask predictor and improved Social-LSTM model are adopted, and combined with Graph-TCN hybrid network, precise pose evaluation and precise gait feature extraction in multiple scenarios.

Benefits of technology

It realizes lightweight and portable design, high-efficiency computing, maintains detection accuracy of 92.4%, and is stable tracking under strong backlight and occlusion scenarios. It generates visual reports, supports accurate posture evaluation and gait feature extraction in multiple scenarios, improving the equipment's environmental adaptability and target tracking stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108043B_ABST
    Figure CN120108043B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of human posture detection and gait analysis, and discloses a portable posture detection and gait analysis method and system. The method includes: building a hardware system, the hardware system including an embedded vision module, a data processing module, and an interaction module, and collecting a human motion video stream through the embedded vision module; using a background difference method and an adaptive illumination compensation algorithm to extract a dynamic foreground from the collected video stream, and output an illumination-compensated image; based on the output illumination-compensated image, using the data processing module, adding a timing constraint loss to the YoLoV8 model, locating key points of the human body, and constructing a three-dimensional skeletal motion model; using a dynamic occlusion processing engine and a multi-target decoupling mechanism, and tracking the constructed three-dimensional skeletal motion model based on a hierarchical correlation metric function. The present invention can collect human motion data in real time and achieve accurate posture assessment and gait feature extraction in multiple scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of human posture detection and gait analysis, and in particular to a portable posture detection and gait analysis method and system. Background Art

[0002] Existing posture detection and gait analysis systems typically employ a cumbersome device architecture, such as using traditional servers or mainframe computers as data processing modules and ordinary cameras for data collection. Existing technologies for posture detection and gait analysis suffer from the following objective shortcomings:

[0003] (1) The test environment has strict requirements. For example, there can only be one tester in the field of view of the camera, and no other people. There are also strict requirements on lighting conditions, and there cannot be strong light. In actual use, it is difficult to find a venue that meets such conditions. During large-scale screening, if other people appear in the field of view, data errors may occur. This is because the existing technology lacks effective target recognition and environmental adaptability, cannot accurately distinguish between the tester and other people, and cannot cope with the impact of light changes on detection accuracy.

[0004] (2) There are deficiencies in target tracking and recognition. When occlusion occurs, the target is easily lost and the tester's motion trajectory cannot be accurately tracked. In multi-target scenarios, different testers cannot be effectively distinguished. ID switching errors are prone to occur, leading to data confusion. This is because the target tracking algorithm of the existing technology is relatively simple and lacks in-depth understanding and analysis of the target, and cannot maintain a stable tracking effect in complex situations.

[0005] (3) The equipment in the existing technology is usually bulky and inconvenient to carry and use, which limits its application in different scenarios. At the same time, the equipment cannot automatically adapt to different test distances and angles and requires manual adjustment, which increases the complexity and workload of the operation. In addition, when facing different environmental interference factors, such as reflections, the equipment in the existing technology also lacks effective suppression measures, which affects the accuracy of the test results.

[0006] (4) Using traditional CPU for data processing, the efficiency is low when running complex posture detection and gait analysis algorithms, resulting in slow processing speed of the entire system and difficulty in achieving real-time analysis. When processing high-definition video streams, there may be freezes and analysis results cannot be given in time, affecting user experience and actual application effects.

[0007] (5) The feedback method is relatively simple, usually only displaying the analysis results through simple text or charts, lacking intuitive visualization effects. Moreover, detailed reports cannot be automatically generated based on the analysis results. Doctors or researchers need to manually record and organize data, which increases the workload and the risk of errors.

[0008] Therefore, the present invention provides a portable posture detection and gait analysis method and system. Summary of the Invention

[0009] This application provides a portable posture detection and gait analysis method and system for real-time collection of human motion data, achieving accurate posture assessment and gait feature extraction in multiple scenarios.

[0010] In a first aspect, the present application provides a portable posture detection and gait analysis method, the method comprising:

[0011] Step S1: Building a hardware system, wherein the hardware system includes an embedded vision module, a data processing module, and an interaction module, and collecting a human motion video stream through the embedded vision module;

[0012] Step S2: extracting dynamic foreground from the captured video stream using background difference method and adaptive illumination compensation algorithm, and outputting illumination compensated image;

[0013] Step S3: Based on the output illumination compensation image, the data processing module is used to add a timing constraint loss to the YoLoV8 model, locate the key points of the human body, and construct a three-dimensional skeletal motion model;

[0014] Step S4: Using a dynamic occlusion processing engine and a multi-target decoupling mechanism, the constructed three-dimensional skeletal motion model is tracked based on a hierarchical correlation metric function to obtain posture data, and a Graph-TCN hybrid network is used to analyze the tracked posture data, obtain the key point displacements between consecutive frames, and calculate gait data;

[0015] Step S5: Based on the acquired posture data and gait data, the posture assessment score and correction suggestions are displayed in real time through the interactive module, and an analysis report including a gait parameter trend graph and a risk assessment matrix is ​​generated.

[0016] In conjunction with the first aspect, in a first implementation of the first aspect of the present application, the portable posture detection and gait analysis method includes: the hardware system is further configured with an environment adaptation module, and the environment adaptation module includes:

[0017] Multi-scale feature fusion network: used for adaptive detection distance;

[0018] Anti-occlusion compensation algorithm: used to complete the data by predicting the motion trajectory when part of the limb is occluded during the tracking of the 3D skeletal motion model in step S4;

[0019] Reflection suppression unit: A solution combining polarizing filters and software is used to reduce the impact of reflections on detection.

[0020] In combination with the first aspect, in the second implementation method of the first aspect of the present application, the portable posture detection and gait analysis method includes: the embedded vision module integrates a retractable multispectral camera array, the multispectral camera array is a wide-angle camera group with an adjustable angle, and the embedded vision module adopts a mode of fusion of a visible light camera and an infrared ToF depth sensor, and automatically switches according to ambient lighting conditions.

[0021] In combination with the first aspect, in a third implementation of the first aspect of the present application, the data processing module is equipped with an Nvidia Jetson Nano computing module, adopts an NPU+GPU dual-core heterogeneous architecture, and the visual processing core and the motion analysis core work in parallel.

[0022] In combination with the first aspect, in a fourth implementation of the first aspect of the present application, the data processing module adopts a dynamic model clipping strategy for automatically switching between the full network and clipped network modes according to the complexity of the scene;

[0023] The interactive module uses a touch screen to perform interface display.

[0024] In combination with the first aspect, in a fifth implementation of the first aspect of the present application, the portable posture detection and gait analysis method includes: in step S3:

[0025] The timing constraint loss function is:

[0026]

[0027] Among them, J is the number of key points and T is the length of the time window;

[0028] is: the coordinates of the jth key point in the tth frame.

[0029] In combination with the first aspect, in a sixth implementation of the first aspect of the present application, the portable posture detection and gait analysis method includes: in step S4:

[0030] The hierarchical association metric function is:

[0031]

[0032] Among them, S IoU is the intersection-union similarity; is the motion similarity;

[0033] is the appearance similarity, and α, β, and γ are weight coefficients.

[0034] Among them, the motion similarity Smotion introduces an improved Social-LSTM interaction model.

[0035] In combination with the first aspect, in a seventh implementation of the first aspect of the present application, in step S5:

[0036] The Graph-TCN hybrid network satisfies:

[0037]

[0038] Among them, A is the adjacency matrix of the skeleton connection graph; is the feature matrix of the lth layer;

[0039] is the weight matrix of the graph convolution layer; It is a temporal convolutional network; is the activation function.

[0040] In combination with the first aspect, in the eighth implementation of the first aspect of the present application, the portable posture detection and gait analysis method includes: the dynamic occlusion processing engine includes a mask predictor based on spatiotemporal attention, and through the mask predictor, a motion probability heat map is constructed using historical trajectories. When occlusion occurs, the LSTM-based trajectory prediction module maintains continuous tracking of less than or equal to frames, and the position prediction error is less than or equal to 7.2 pixels.

[0041] In a second aspect, the present application provides a portable posture detection and gait analysis system, the system comprising:

[0042] A hardware module is used to build a hardware system, which includes an embedded vision module, a data processing module, and an interaction module. The embedded vision module is used to collect human motion video streams;

[0043] Dynamic foreground extraction module, used to extract dynamic foreground from the collected video stream using background difference method and adaptive illumination compensation algorithm, and output illumination compensated image;

[0044] A skeleton model construction module is used to add a timing constraint loss to the YoLoV8 model based on the output illumination compensation image, locate the key points of the human body, and construct a three-dimensional skeleton motion model using the data processing module;

[0045] The tracking module is used to track the constructed 3D skeletal motion model based on the hierarchical correlation metric function using the dynamic occlusion processing engine and multi-target decoupling mechanism to obtain posture data. The tracked posture data is analyzed using the Graph-TCN hybrid network to obtain the key point displacement between consecutive frames and calculate the gait data.

[0046] The feedback module is used to display the posture assessment score and correction suggestions in real time through the interactive module based on the acquired posture data and gait data, and generate an analysis report including a gait parameter trend graph and a risk assessment matrix.

[0047] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0048] In the technical solution provided in this application, the hardware adopts a foldable multi-spectral camera array and an embedded heterogeneous computing architecture (NPU+GPU collaboration), combined with a dynamic model clipping strategy, which reduces the computing load by 38% while ensuring detection accuracy (AP greater than or equal to 85%), and realizes lightweight and portable design (the volume after folding is less than or equal to 220mm×200mm×100mm) and high-performance computing (greater than or equal to 30fps real-time processing), solving the problems of bulky and high latency of traditional equipment. Secondly, in response to complex environmental interference, the system integrates dual-channel sensor fusion technology (visible light + infrared ToF depth sensor) and illumination invariant feature extraction network, decouples illumination and shadow through the adversarial generation network, and still maintains a detection accuracy of 92.4% in strong backlight scenes; at the same time, it introduces spatiotemporal attention mask prediction The detector and the improved Social-LSTM model achieve less than or equal to 15 frames of continuous tracking (with an error of less than or equal to 7.2 pixels) in occluded or multi-person scenes, effectively avoiding ID switching errors. In addition, the 3D skeletal motion modeling based on the improved YoLoV8 model and the Graph-TCN hybrid network accurately extract parameters such as step length and joint angle, achieving accurate posture assessment and gait feature extraction in multiple scenarios, and generating real-time visual reports (including gait trend graphs and risk assessment matrices) through a multimodal interactive interface, providing high-value data support for medical diagnosis, rehabilitation training and other fields. These technologies work together to solve the pain points of existing systems such as poor environmental adaptability, unstable target tracking and single feedback function, and are suitable for multiple scenarios such as outdoor screening, dense crowd analysis and dynamic rehabilitation monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 This is a flow chart of the portable posture detection and gait analysis method in an embodiment of the present application;

[0051] Figure 2 It is a structural diagram of the portable posture detection and gait analysis system in an embodiment of the present application. DETAILED DESCRIPTION

[0052] Embodiments of the present application provide a portable posture detection and gait analysis method and system. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0053] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the portable posture detection and gait analysis method of the present application, the method includes:

[0054] Step S1: Building a hardware system, wherein the hardware system includes an embedded vision module, a data processing module, and an interaction module, and collecting a human motion video stream through the embedded vision module.

[0055] Specifically, the embedded vision module is responsible for high-precision acquisition of human motion video streams and adapting to complex environments. The data processing module processes video stream data in real time, performs key point detection, three-dimensional modeling and gait analysis. The interactive module provides a user interface through a touch screen, provides real-time feedback on analysis results and generates reports, while also visualizing them in real time through software.

[0056] Step S2: extract the dynamic foreground from the captured video stream using a background difference method and an adaptive illumination compensation algorithm, and output an illumination-compensated image.

[0057] Specifically, the background difference method is combined with the adaptive illumination compensation algorithm to eliminate the impact of ambient illumination changes on detection accuracy. This combination can effectively improve the robustness of foreground extraction, especially in scenes with large illumination changes. For example, in a complex crowd environment, the test subject can be accurately identified and tracked, and even if other people appear in the camera's field of view, it will not interfere with the detection of the test subject. In the case of light changes, the illumination compensation parameters can be automatically adjusted to ensure image quality, thereby improving the accuracy and reliability of detection.

[0058] For the collected video stream, set each frame image to I t , the background model is B t , the image after illumination compensation is :

[0059] Background modeling:

[0060] The initial background model is the first frame image:

[0061] in, is the initial background model; The first frame image.

[0062]

[0063] in,

[0064] is the weight coefficient.

[0065] Foreground detection:

[0066]

[0067] in, is the difference image of the t-th frame.

[0068]

[0069] in, is the binary foreground mask, 255 is foreground and 0 is background;

[0070] threshold indicates the threshold value;

[0071] Introducing spatiotemporal adaptive weight coefficients :

[0072]

[0073] in Represents the variance between the current frame and the background model;

[0074] is the dynamic threshold; k is the adjustment factor;

[0075] Multi-model background modeling:

[0076]

[0077] in, is the weight of the i-th background model at time t; is the i-th background model.

[0078] The mixed Gaussian model and the color constancy model are used for parallel calculation, and the weights are dynamically adjusted through KL divergence. i ;

[0079] Light compensation:

[0080]

[0081] in, For the Frame original image;

[0082] GammaCorrection is a two-dimensional gamma correction function that uses the illumination component and the adaptive target mean for adjustment;

[0083] Lighting compensation enhancement solution:

[0084]

[0085] in, is the reflection component; Γ is the adaptive gamma correction function;

[0086] Separate illumination components via dual-domain filtering.

[0087] Deep learning-assisted compensation:

[0088] Build a lightweight U-Net network, input the original image and HSV spatial features, and output the illumination compensated image:

[0089] .

[0090] in, Light compensation image predicted by U-Net; Compensate images for real lighting;

[0091] and is the weight coefficient of the loss function;

[0092] SSIM is a structural similarity index that measures the consistency of image structure.

[0093] Step S3: Based on the output illumination compensation image, the data processing module is used to add a timing constraint loss to the YoLoV8 model, locate the key points of the human body, and construct a three-dimensional skeletal motion model.

[0094] Specifically, the YoLoV8 model is improved to locate key points of the human body and construct a 3D skeletal motion model. This model enables more accurate assessment of human posture and extraction of gait features. For example, when analyzing the gait cycle, common gait parameters such as stride length, cadence, and toe-off height can be more accurately determined, providing a more reliable data foundation for subsequent gait analysis and diagnosis.

[0095] Step S4: Using the dynamic occlusion processing engine and the multi-target decoupling mechanism, the constructed three-dimensional skeletal motion model is tracked based on the hierarchical correlation metric function to obtain posture data. The tracked posture data is analyzed using the Graph-TCN hybrid network to obtain the key point displacement between consecutive frames and calculate the gait data.

[0096] Specifically, the dynamic occlusion processing engine and multi-target decoupling mechanism ensure the reliability of target tracking. The multi-target decoupling mechanism is based on a hierarchical association measurement function. The motion similarity Smotion introduces an improved Social-LSTM interaction model to solve the ID switching problem in dense crowds. During large-scale screening, multiple testers may appear in the camera's field of view at the same time. The multi-target decoupling mechanism can accurately distinguish different testers and avoid ID confusion. For example, when conducting gait screening in crowded places such as schools or enterprises, the device can accurately identify and track each tester, ensuring that the gait data of each tester can be accurately recorded and analyzed.

[0097] Step S5: Based on the acquired posture data and gait data, the posture assessment score and correction suggestions are displayed in real time through the interactive module, and an analysis report including a gait parameter trend graph and a risk assessment matrix is ​​generated.

[0098] Specifically, the interactive module provides a multimodal interactive interface, which calculates the posture stability score in real time based on indicators such as key point offset and joint range of motion. The visual interface displays the real-time posture assessment score and correction suggestions, and generates a PDF analysis report including a gait parameter trend chart and a risk assessment matrix. These functions enable users to intuitively understand their gait conditions and conduct targeted corrections and training based on the analysis report. For example, for rehabilitation patients, by viewing the gait parameter trend chart, they can clearly understand the progress of their gait recovery, and doctors can also adjust the treatment plan in a timely manner based on the risk assessment matrix.

[0099] In a specific embodiment, the hardware system is further configured with an environment adaptation module, and the environment adaptation module includes:

[0100] Multi-scale feature fusion network: used for adaptive detection distance, automatically adapting to the detection distance of 1-5 meters (usually the test distance is less than 3 meters);

[0101] Anti-occlusion compensation algorithm: used to complete the data by predicting the motion trajectory when part of the limb is occluded during the tracking of the 3D skeletal motion model in step S4;

[0102] Reflection suppression unit: A solution combining polarizing filters and software is used to reduce the impact of reflections on detection.

[0103] Specifically, in different test sites and scenarios, the equipment can be quickly deployed and the detection distance and angle can be automatically adjusted to adapt to different test needs; when encountering reflection interference, the reflection suppression unit can effectively reduce the impact of reflection on the detection results and improve the accuracy and reliability of the detection.

[0104] The anti-occlusion compensation algorithm works together with the dynamic occlusion processing engine and multi-target decoupling mechanism. During long-term gait tracking, even in the presence of brief occlusions or the simultaneous presence of multiple people, the device can accurately track the target and provide continuous and reliable gait data.

[0105] The anti-occlusion compensation algorithm is specifically as follows:

[0106] Let the t-th frame detection set be , the tracker set is , where each tracker Include:

[0107] State vector

[0108] in, is the target center coordinate; is the aspect ratio;

[0109] is the height; The rate of change of the corresponding parameter;

[0110] Appearance feature vector ;

[0111] Covariance matrix Describe state uncertainty;

[0112] Input: Target ID k , current frame tracking list

[0113] Output: Target position ( )

[0114] The output target does not exist in the tracking list When the occlusion processing principle is triggered:

[0115] N consecutive frames do not match:

[0116] For short-term loss, that is, when N≤3, the Kalman filter is used to predict the trajectory:

[0117]

[0118] Among them, F is the state transfer matrix; B is the control input matrix;

[0119] is the state vector of target k at time t (including the center coordinates , aspect ratio γ, height h and its rate of change);

[0120] is the control vector;

[0121] In case of long-term loss, that is, when N>3, the system prompts the tester to take his place.

[0122] In a specific embodiment, the embedded vision module integrates a retractable multispectral camera array, which is a wide-angle camera group with adjustable angles. The embedded vision module adopts a mode that integrates a visible light camera and an infrared ToF depth sensor, and automatically switches according to ambient lighting conditions.

[0123] Specifically, the retractable multispectral camera array adopts a foldable structure design, which forms an adjustable field of view of 60°-120° when unfolded, and the volume when folded is less than or equal to 220mm×200mm×100mm. This design makes the device easy to carry and store. Whether it is for large-scale screening in remote areas or transferred between different medical institutions, it can be easily carried, greatly improving the flexibility of the device.

[0124] The purpose of illumination adaptation is achieved by integrating a visible light camera and an infrared ToF depth sensor. An illumination invariant feature extraction network (LIF-Net) is designed, and an illumination-shadow decoupling representation is constructed through a generative adversarial network. The detection accuracy is maintained at 92.4% in strong backlighting scenarios. In actual use, lighting conditions are often complex and changeable. For example, during outdoor screening, strong direct sunlight or shadows may be encountered. The present invention can automatically adapt to these lighting changes to ensure the accuracy of the detection results. For example, in outdoor screening in the early morning or evening, even if the light is weak and there is backlighting, the device can still accurately detect human posture and gait characteristics.

[0125] In a specific embodiment, the data processing module is equipped with an Nvidia Jetson Nano computing module, adopts an NPU+GPU dual-core heterogeneous architecture, and the visual processing core and the motion analysis core work in parallel.

[0126] Specifically, the Nvidia Jetson Nano computing module of the data processing module can be powered by either AC power or lithium battery. It adopts a portable architecture with embedded edge computing and adaptive inference. Through the NPU+GPU dual-core heterogeneous architecture, the NPU specializes in fixed-point acceleration of lightweight posture detection networks, and the GPU is responsible for the timing convolution operations of gait analysis. Compared with traditional pure CPU solutions, it achieves an 11.6-fold improvement in energy efficiency. For example, with the same power supply, this device can continue to work for a longer time, which is especially important for large-scale screening scenarios. It does not require frequent charging, which improves work efficiency. When processing high-definition video streams, it can quickly complete data processing and analysis, and provide accurate detection results in a timely manner to meet the needs of real-time analysis, improve work efficiency and user experience.

[0127] In addition, the real-time processing speed is ensured to be no less than 30fps. In the gait analysis process, real-time performance is very important. For example, in clinical diagnosis, doctors need to obtain the patient's gait data in a timely manner for diagnosis. The high-speed real-time processing capability of the present invention can meet this demand, quickly provide accurate gait analysis results, and provide timely basis for the doctor's diagnosis.

[0128] In a specific embodiment, the data processing module adopts a dynamic model clipping strategy for automatically switching between full network and clipped network modes according to scene complexity;

[0129] The interactive module uses a touch screen to perform interface display.

[0130] Specifically, dynamic model cropping can automatically switch the network depth (full network / cropped network) based on the complexity of the scene, reducing the computing load by 38% while ensuring that AP is greater than or equal to 85%. In different scenarios, it can intelligently adjust the allocation of its own computing resources to avoid unnecessary computing waste. For example, in a simple gait analysis scenario, the device can quickly switch to the cropped network mode to speed up processing, while ensuring sufficient computing accuracy in complex scenarios.

[0131] In a specific embodiment, in step S3:

[0132] The timing constraint loss function is:

[0133]

[0134] Among them, J is the number of key points and T is the length of the time window;

[0135] is: the coordinates of the jth key point in the tth frame.

[0136] In a specific embodiment, in step S4:

[0137] The hierarchical association metric function is:

[0138]

[0139] Among them, S IoU is the intersection-union similarity; is the motion similarity;

[0140] is the appearance similarity, α, β, γ are weight coefficients;

[0141] Smotion introduces an improved Social-LSTM interaction model.

[0142] Specifically, comprehensive geometry ( ), motion (improved Social-LSTM) and appearance (128-dimensional feature vector) similarity ( ), improve the robustness of multi-target tracking and solve the ID switching problem in dense crowds.

[0143] In a specific embodiment, in step S5:

[0144] The Graph-TCN hybrid network satisfies:

[0145]

[0146] Among them, A is the adjacency matrix of the skeleton connection graph; is the feature matrix of the lth layer;

[0147] is the weight matrix of the graph convolution layer; It is a temporal convolutional network; is the activation function.

[0148] In a specific embodiment, the dynamic occlusion processing engine includes a mask predictor based on spatiotemporal attention. Through the mask predictor, a motion probability heat map is constructed using historical trajectories. When occlusion occurs, the LSTM-based trajectory prediction module maintains continuous tracking of less than or equal to 15 frames, and the position prediction error is less than or equal to 7.2 pixels.

[0149] Specifically, in actual gait analysis scenarios, the test subject may be obscured by surrounding objects or people. The dynamic occlusion processing engine of the present invention can effectively solve this problem. For example, when conducting gait screening in crowded places, when the test subject is briefly obscured by other people, the device can still accurately track his or her movement trajectory, ensuring the continuity and accuracy of the data.

[0150] The portable posture detection and gait analysis method in the embodiment of the present application is described above. The portable posture detection and gait analysis system in the embodiment of the present application is described below. Figure 2 In one embodiment of the present application, a portable posture detection and gait analysis system includes:

[0151] A hardware module is used to build a hardware system, wherein the hardware system includes an embedded vision module, a data processing module, and an interaction module, and is used to collect a human motion video stream through the embedded vision module;

[0152] Dynamic foreground extraction module, used to extract dynamic foreground from the collected video stream using background difference method and adaptive illumination compensation algorithm, and output illumination compensated image;

[0153] A skeleton model construction module is used to add a timing constraint loss to the YoLoV8 model based on the output illumination compensation image, locate the key points of the human body, and construct a three-dimensional skeleton motion model using the data processing module;

[0154] The tracking module is used to track the constructed 3D skeletal motion model based on the hierarchical correlation metric function using the dynamic occlusion processing engine and multi-target decoupling mechanism to obtain posture data. The tracked posture data is analyzed using the Graph-TCN hybrid network to obtain the key point displacement between consecutive frames and calculate the gait data.

[0155] The feedback module is used to display the posture assessment score and correction suggestions in real time through the interactive module based on the acquired posture data and gait data, and generate an analysis report including a gait parameter trend graph and a risk assessment matrix.

[0156] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A portable posture detection and gait analysis method, characterized in that: include: S1. Build a hardware system, which includes an embedded vision module, a data processing module, and an interaction module. The embedded vision module is used to collect human motion video streams. S2. Use background difference method and adaptive illumination compensation algorithm to extract dynamic foreground from the collected video stream and output illumination compensation image; including background modeling, foreground detection and introduction of spatiotemporal adaptive weight coefficient for the collected video stream. : Represents the variance between the current frame and the background model; is the dynamic threshold; k is the adjustment factor; Multi-model background modeling: is the weight of the i-th background model at time t; is the i-th background model, and the background model is B t ; The mixed Gaussian model and the color constancy model are used for parallel calculation, and the weights are dynamically adjusted through KL divergence. i ; Light compensation: For the Frame original image; GammaCorrection is a two-dimensional gamma correction function that uses the illumination component and the adaptive target mean for adjustment; Lighting compensation enhancements: is the reflection component; Γ is the adaptive gamma correction function; Separate illumination components through dual-domain filtering; Deep learning-assisted compensation: Build a lightweight U-Net network, input the original image and HSV spatial features, and output the illumination compensated image: The illumination compensation image predicted by U-Net; Compensate images for real lighting; and is the weight coefficient of the loss function; SSIM is a structural similarity index that measures the consistency of image structure; S3. Based on the output illumination compensation image, the data processing module is used to add the timing constraint loss on the basis of the YoLoV8 model, locate the key points of the human body, and construct a 3D skeletal motion model. S4. Using the dynamic occlusion processing engine and multi-target decoupling mechanism, the constructed 3D skeletal motion model is tracked based on the hierarchical correlation metric function to obtain posture data. The tracked posture data is analyzed using the Graph-TCN hybrid network to obtain the key point displacement between consecutive frames and calculate the gait data. S5. Based on the acquired posture data and gait data, the posture assessment score and correction suggestions are displayed in real time through the interactive module, and an analysis report including a gait parameter trend graph and a risk assessment matrix is ​​generated.

2. The method according to claim 1, characterized in that The hardware system is also equipped with an environmental adaptation module, which includes: Multi-scale feature fusion network: used for adaptive detection distance; Anti-occlusion compensation algorithm: used to complete the data by predicting the motion trajectory when part of the limb is occluded during the tracking of the 3D skeletal motion model in step S4; Reflection suppression unit: A solution combining polarizing filters and software is used to reduce the impact of reflections on detection.

3. The method according to claim 1, characterized in that The embedded vision module integrates a retractable multispectral camera array. The multispectral camera array is a wide-angle camera group with adjustable angles. The embedded vision module adopts a fusion mode of visible light camera and infrared ToF depth sensor, which automatically switches according to the ambient lighting conditions.

4. The method according to claim 1, wherein The data processing module is equipped with an Nvidia Jetson Nano computing module and adopts an NPU+GPU dual-core heterogeneous architecture, with the visual processing core and motion analysis core working in parallel.

5. The method according to claim 1, characterized in that The data processing module adopts a dynamic model pruning strategy to automatically switch between full network and pruning network modes according to the complexity of the scene; The interactive module uses a touch screen for interface display.

6. The method according to claim 1, characterized in that In step S3: The timing constraint loss function is: Among them, J is the number of key points and T is the length of the time window; is: the coordinates of the jth key point in the tth frame.

7. The method according to claim 1, characterized in that In step S4: The hierarchical association metric function is: Among them, S IoU is the intersection-union similarity; is the motion similarity; is the appearance similarity, α, β, γ are weight coefficients; Smotion introduces an improved Social-LSTM interaction model.

8. The method according to claim 1, characterized in that In step S5: The Graph-TCN hybrid network satisfies: Among them, A is the adjacency matrix of the skeleton connection graph; is the feature matrix of the lth layer; is the feature matrix of the l+1th layer; is the weight matrix of the graph convolution layer; It is a temporal convolutional network; is the activation function.

9. The method according to claim 1, characterized in that The dynamic occlusion processing engine includes a mask predictor based on spatiotemporal attention. Through the mask predictor, historical trajectories are used to construct a motion probability heat map. When occlusion occurs, the LSTM-based trajectory prediction module maintains continuous tracking of less than or equal to 15 frames, and the position prediction error is less than or equal to 7.2 pixels.

10. A portable posture detection and gait analysis system for implementing the method according to any one of claims 1 to 9, characterized in that: Includes the following modules: Hardware module, used to build a hardware system. The hardware system includes an embedded vision module, a data processing module, and an interaction module. The embedded vision module is used to collect human motion video streams. Dynamic foreground extraction module, used to extract dynamic foreground from the collected video stream using background difference method and adaptive illumination compensation algorithm, and output illumination compensated image; The skeleton model construction module is used to compensating the output illumination image, using the data processing module to add timing constraint loss to the YoLoV8 model, locate the key points of the human body, and construct a 3D skeleton motion model; The tracking module is used to track the constructed 3D skeletal motion model based on the hierarchical correlation metric function using the dynamic occlusion processing engine and multi-target decoupling mechanism to obtain posture data. The tracked posture data is analyzed using the Graph-TCN hybrid network to obtain the key point displacement between consecutive frames and calculate the gait data. The feedback module is used to display the posture assessment score and correction suggestions in real time through the interactive module based on the acquired posture data and gait data, and generate an analysis report including a gait parameter trend graph and a risk assessment matrix.

Citation Information

Patent Citations

  • Maneuvering multi-target tracking method based on combination of kernel adaptive filtering and YOLOX detection

    CN114972418A

  • Multi-person gait recognition system based on three-dimensional human body skeleton

    CN117173792A

  • Concrete construction process and worker health monitoring system based on computer vision

    CN119810913A