Human body tumble detection method and system based on multi-sensor fusion

By collaboratively acquiring multimodal data from thermal imaging sensors and millimeter-wave radar, and analyzing spatiotemporal graph convolutional networks, the problems of high false alarm rate and limited applicability in existing fall detection technologies have been solved, achieving efficient and reliable fall detection in complex scenarios.

CN120918637APending Publication Date: 2025-11-11广东奥莱敏控技术有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510999029.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies cannot accurately identify human fall behavior in complex scenarios, resulting in a high false alarm rate and limited applicability, making it difficult to achieve efficient and reliable fall detection and timely response.

Method used

The system uses thermal imaging sensors and millimeter-wave radar to collect human body temperature distribution and three-dimensional motion information. By fusing thermodynamic and kinematic spatiotemporal gradient features, a spatiotemporal graph structure is constructed. The spatiotemporal graph convolutional network is used to perform human fall feature evolution, and a fall reminder is executed when a preset threshold is triggered.

Benefits of technology

It achieves accurate identification and reliable response to human fall behavior in complex scenarios, reduces false alarm rate, and improves detection accuracy and timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120918637A_ABST
    Figure CN120918637A_ABST
Patent Text Reader

Abstract

The invention discloses a human body tumble detection method and system based on multi-sensor fusion, and the method comprises the steps: collecting human body temperature distribution and three-dimensional point cloud data through a thermal imaging sensor and a millimeter wave radar, respectively extracting a thermodynamic feature tensor and a three-dimensional motion feature tensor, and capturing the contour and motion track information of a human body; through space-time gradient calculation, body temperature change rate and posture sudden change intensity characteristics are deeply excavated; fusing the two types of features to construct a space-time diagram structure, representing joint thermodynamic features by nodes, and representing joint dynamic association by edges; then realizing tumble feature evolution and outputting a risk probability by virtue of multi-level interaction such as space aggregation, multi-scale time sequence modeling and an attention mechanism of a space-time diagram convolutional network; when the probability exceeds a threshold value, reminding is triggered, and high-precision fall detection driven by multi-sensor data is achieved. According to the invention, the human body tumble behavior in a complex scene can be accurately identified, and efficient and reliable tumble detection and timely response are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for detecting human falls based on multi-sensor fusion. Background Technology

[0002] Currently, fall detection technologies primarily rely on single sensors or simple multi-sensor data stitching. For example, some systems use wearable accelerometers or gyroscopes to collect human motion information, analyzing parameters such as posture angles and acceleration changes to determine if a fall has occurred; others use cameras for image recognition, leveraging computer vision technology to extract human motion features. However, the former is susceptible to interference from daily activities, leading to a high false alarm rate, and also suffers from inconvenience when worn; the latter is limited by lighting conditions and privacy concerns, restricting its applicability to various scenarios. Therefore, current technologies cannot accurately identify human fall behavior in complex scenarios, making it difficult to achieve efficient and reliable fall detection and timely response. Summary of the Invention

[0003] This invention provides a method and system for human fall detection based on multi-sensor fusion, which can accurately identify human fall behavior in complex scenarios and achieve efficient and reliable fall detection and timely response.

[0004] One embodiment of the present invention provides a human fall detection method based on multi-sensor fusion, comprising:

[0005] The thermal imaging sensor detects the temperature distribution of the surrounding human body in real time, generates a thermal imaging image of the human body, and simultaneously transmits high-frequency electromagnetic waves to the surrounding human body through millimeter-wave radar and receives reflected signals to generate three-dimensional point cloud data of the human body.

[0006] Human contours are extracted from the thermal imaging image to generate a thermodynamic feature tensor of the human body, and human motion trajectory is reconstructed from the three-dimensional point cloud data to generate a three-dimensional motion feature tensor of the human body.

[0007] Thermodynamic spatiotemporal gradient calculation is performed on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features that reflect the rate of change of body temperature distribution of the human body, and kinematic spatiotemporal gradient calculation is performed on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features that reflect the intensity of abrupt changes in the posture of the human body.

[0008] The spatiotemporal gradient features of the thermodynamics and the spatiotemporal gradient features are fused to construct the spatiotemporal graph structure of human motion;

[0009] The spatiotemporal graph structure is input into the spatiotemporal graph convolutional network, and the human fall feature evolution is performed through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network to output the fall risk probability.

[0010] When the probability of falling exceeds a preset threshold, a human fall warning operation is executed.

[0011] Another embodiment of the present invention provides a human fall detection system based on multi-sensor fusion, comprising:

[0012] The acquisition module allows users to detect the temperature distribution of the surrounding human body in real time through a thermal imaging sensor, generate a thermal imaging image of the human body, and simultaneously transmit high-frequency electromagnetic waves to the surrounding human body through a millimeter-wave radar and receive reflected signals to generate three-dimensional point cloud data of the human body.

[0013] The feature extraction module is used to extract the human body contour from the thermal imaging image, generate the thermodynamic feature tensor of the human body, and reconstruct the human body motion trajectory from the three-dimensional point cloud data, generating the three-dimensional motion feature tensor of the human body.

[0014] The feature calculation module is used to perform thermodynamic spatiotemporal gradient calculation on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features that reflect the rate of change of body temperature distribution of the human body, and to perform kinematic spatiotemporal gradient calculation on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features that reflect the intensity of abrupt changes in the posture of the human body.

[0015] A construction module is used to construct a spatiotemporal graph structure of human motion based on the fusion of the thermodynamic spatiotemporal gradient features and the kinematic spatiotemporal gradient features;

[0016] The prediction module is used to input the spatiotemporal graph structure into the spatiotemporal graph convolutional network, perform human fall feature evolution through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network, and output the fall risk probability.

[0017] The reminder module is used to perform a human fall reminder operation when the probability of falling exceeds a preset threshold.

[0018] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0019] By collaboratively acquiring multimodal data from thermal imaging sensors and millimeter-wave radar, the body's temperature distribution and three-dimensional motion information are obtained. The thermal imaging images are used to extract the human contour and generate a thermodynamic feature tensor. Simultaneously, the millimeter-wave point cloud data is used to reconstruct the motion trajectory and generate a three-dimensional motion feature tensor, achieving parallel extraction of physiological and motion features. Spatiotemporal gradient calculations are performed on both types of feature tensors to quantify the rate of change in body temperature distribution and the intensity of abrupt posture changes, thereby capturing the biomechanical anomalies of the fall process and obtaining dual-modal gradient features of thermodynamic and kinematic spatiotemporal gradients. These dual-modal gradient features are fused to construct a spatiotemporal graph structure. The hierarchical feature interaction of the spatiotemporal graph convolutional network enables dynamic evolution recognition of fall features, ultimately triggering an alert based on a probability threshold. In summary, this invention, through the fusion of thermodynamic and kinematic spatiotemporal features and the structured representation capabilities of graph neural networks, effectively solves the technical problems of high false alarm rates from single sensors and the inability of simple multi-sensor stitching to capture the essential features of falls in the background technology, achieving accurate recognition and reliable response to human fall behavior in complex scenarios. Therefore, the embodiments of the present invention, through the fusion of multimodal spatiotemporal features of thermal imaging and millimeter-wave radar and graph structure evolution analysis, can accurately identify human fall behavior in complex scenarios, and achieve efficient and reliable fall detection and timely response. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a human fall detection method based on multi-sensor fusion according to an embodiment of the present invention.

[0021] Figure 2 This is a schematic diagram of a human fall detection system based on multi-sensor fusion provided in an embodiment of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] See Figure 1 This is a flowchart illustrating a human fall detection method based on multi-sensor fusion according to an embodiment of the present invention. The human fall detection method based on multi-sensor fusion includes the following steps:

[0024] S10: The thermal imaging sensor detects the temperature distribution of the surrounding human body in real time, generates a thermal imaging image of the human body, and simultaneously transmits high-frequency electromagnetic waves to the surrounding human body through millimeter-wave radar and receives reflected signals to generate three-dimensional point cloud data of the human body.

[0025] S11, extract the human body contour from the thermal imaging image to generate the thermodynamic feature tensor of the human body, and reconstruct the human body motion trajectory from the three-dimensional point cloud data to generate the three-dimensional motion feature tensor of the human body.

[0026] S12, perform thermodynamic spatiotemporal gradient calculation on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features that reflect the rate of change of body temperature distribution of the human body, and perform kinematic spatiotemporal gradient calculation on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features that reflect the intensity of abrupt changes in the posture of the human body.

[0027] S13, construct a spatiotemporal graph structure of human motion based on the fusion of the thermodynamic spatiotemporal gradient features and the kinematic spatiotemporal gradient features;

[0028] S14, input the spatiotemporal graph structure into the spatiotemporal graph convolutional network, perform human fall feature evolution through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network, and output the fall risk probability;

[0029] S15, when the probability of falling exceeds a preset threshold, execute a human fall reminder operation.

[0030] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0031] By collaboratively acquiring multimodal data from thermal imaging sensors and millimeter-wave radar, the body's temperature distribution and three-dimensional motion information are obtained. The thermal imaging images are used to extract the human contour and generate a thermodynamic feature tensor. Simultaneously, the millimeter-wave point cloud data is used to reconstruct the motion trajectory and generate a three-dimensional motion feature tensor, achieving parallel extraction of physiological and motion features. Spatiotemporal gradient calculations are performed on both types of feature tensors to quantify the rate of change in body temperature distribution and the intensity of abrupt posture changes, thereby capturing the biomechanical anomalies of the fall process and obtaining dual-modal gradient features of thermodynamic and kinematic spatiotemporal gradients. These dual-modal gradient features are fused to construct a spatiotemporal graph structure. The hierarchical feature interaction of the spatiotemporal graph convolutional network enables dynamic evolution recognition of fall features, ultimately triggering an alert based on a probability threshold. In summary, this invention, through the fusion of thermodynamic and kinematic spatiotemporal features and the structured representation capabilities of graph neural networks, effectively solves the technical problems of high false alarm rates from single sensors and the inability of simple multi-sensor stitching to capture the essential features of falls in the background technology, achieving accurate recognition and reliable response to human fall behavior in complex scenarios. Therefore, the embodiments of the present invention, through the fusion of multimodal spatiotemporal features of thermal imaging and millimeter-wave radar and graph structure evolution analysis, can accurately identify human fall behavior in complex scenarios, and achieve efficient and reliable fall detection and timely response.

[0032] As an improvement to the above embodiment, the step of detecting the temperature distribution of the surrounding human body in real time using a thermal imaging sensor to generate a thermal imaging image of the human body, and simultaneously transmitting high-frequency electromagnetic waves to the surrounding human body using a millimeter-wave radar and receiving reflected signals to generate three-dimensional point cloud data of the human body, includes the following sub-steps:

[0033] Adaptive noise filtering is applied to the raw temperature data acquired by the thermal imaging sensor to generate a noise-reduced temperature distribution matrix.

[0034] A multi-scale segmentation algorithm is used to identify human target regions in the denoised temperature distribution matrix, generating an accurate binary mask that separates the human body from the background.

[0035] Based on the binary mask, spatial domain filtering is performed on the original temperature data to extract the temperature distribution data of the target area of ​​the human body and generate a thermal imaging image of the human body.

[0036] The direction of arrival (DOA) of the reflected signal received by the millimeter-wave radar is estimated and multipath interference is eliminated to generate an initial set of spatial point clouds.

[0037] The initial set of spatial point clouds is dynamically separated using a density clustering algorithm to extract three-dimensional point cloud data corresponding to human targets.

[0038] In this embodiment, data acquired by thermal imaging sensors and millimeter-wave radar are processed separately to obtain effective information. For thermal imaging data, the raw temperature data is first processed by adaptive noise filtering to suppress noise interference from the environment and equipment, resulting in a denoised temperature distribution matrix. Then, a multi-scale segmentation algorithm is used to identify the human target region from the denoised data, generating a binary mask. Finally, spatial domain filtering is applied to the raw data based on the mask to accurately extract the human body temperature distribution data, forming a thermal imaging image. For millimeter-wave radar data, the direction of arrival of the reflected signal is first estimated and multipath interference is eliminated to reduce signal interference, generating an initial set of spatial point clouds. Then, a density clustering algorithm is used to separate the three-dimensional point cloud data corresponding to the human target from the set. Through the above series of steps, a high-quality data source is provided for subsequent human fall detection based on multi-sensor fusion. In summary, this embodiment effectively improves the accuracy and reliability of the data through step-by-step optimization processing of thermal imaging and millimeter-wave radar data. The application of algorithms such as adaptive noise filtering and multi-scale segmentation solves the problems of noise interference and inaccurate target recognition in thermal imaging data; techniques such as direction-of-arrival estimation and density clustering eliminate multipath interference and non-human target interference in millimeter-wave radar data. Compared with traditional single processing methods, this embodiment provides a cleaner and more targeted data foundation for human fall detection, thereby improving the accuracy of multi-sensor fusion in recognizing human fall behavior in complex scenarios.

[0039] To facilitate understanding of this embodiment, the following examples are provided:

[0040] As an example, thermal imaging sensors and millimeter-wave radar (such as 60G millimeter-wave radar) work together to acquire human body data. The thermal imaging sensor, based on the principle of infrared radiation, incorporates an uncooled microbolometer and can detect the infrared radiation energy emitted by the human body in the surrounding environment in real time at a frequency of 25 frames per second. The sensor converts the received infrared radiation into electrical signals, which are then preprocessed through analog-to-digital conversion and non-uniformity correction to map infrared radiation of different intensities to corresponding temperature values, ultimately generating a 640×480 resolution thermal imaging image. This image visually presents the temperature distribution of the human body in grayscale or pseudo-color, clearly outlining the human body's contours. Simultaneously, the millimeter-wave radar continuously emits high-frequency electromagnetic waves at 77GHz into the surrounding space. These electromagnetic waves are reflected upon encountering the human body, and the radar module's receiving antenna array captures the reflected signals. The received reflected signals contain information such as the distance, speed, and angle of the human target. After low-noise amplification, mixing, filtering, and other RF front-end processing, they enter the signal processing unit. In the signal processing unit, the reflected signal is analyzed and demodulated using algorithms such as Fast Fourier Transform (FFT) and Doppler processing. The processed information is then converted into point cloud data in three-dimensional space. Each point cloud data point contains the coordinates (x, y, z) of the corresponding human body part in space as well as velocity information, providing raw data support for subsequent feature extraction and fall detection.

[0041] Specifically, firstly, adaptive noise filtering is performed on the raw temperature data acquired by the thermal imaging sensor. An improved adaptive weighted median filtering algorithm is used, assuming the raw temperature data is T = {t...} ij} where i represents the row index and j represents the column index. This algorithm calculates the temperature difference between each pixel and its neighboring pixels, assigning different weights based on the magnitude of the difference: pixels with smaller differences have higher weights, and pixels with larger differences have lower weights. A weight threshold ω is set. th When the sum of the weights of a pixel and its neighboring pixels is less than ω th At that time, a stronger filtering process is applied to the pixel to suppress noise, resulting in a denoised temperature distribution matrix.

[0042] Next, a multi-scale segmentation algorithm is used to refine the denoised temperature distribution matrix T. n Human target region recognition is performed. This algorithm is based on multi-resolution analysis, decomposing the temperature distribution matrix into multiple sub-matrices of different scales. At each scale, a region-growing-based segmentation method is used to select representative seed points, based on temperature stability and neighborhood similarity. Starting from the seed points, pixels with a temperature difference less than a set threshold Δt and satisfying spatial adjacency are merged into the same region, gradually growing the human target region. After multi-scale fusion processing, a precise binary mask M for separating the human body from the background is generated, where M...ij =1 represents the number of pixels in the target area of ​​the human body, M ij =0 indicates background pixels.

[0043] Then, spatial domain filtering is performed on the original temperature data T based on a binary mask M. This is achieved by constructing a filtering matrix F, where M... ij When F = 1, ij =1, when M ij When F = 0, ij =0. Multiply the filter matrix F with the original temperature data T element-wise, i.e., T = 0. h =F⊙T, where ⊙ represents element-wise multiplication of matrices, thereby removing background temperature data interference, extracting temperature distribution data of the target human body region, and generating a thermal imaging image T of the human body. h .

[0044] For the reflected signal received by millimeter-wave radar, direction-of-arrival (DOA) estimation and multipath interference cancellation are performed first. An improved ODA algorithm based on compressed sensing and spatiotemporal joint processing is adopted. Let the received signal vector be r(t), and a joint sparse dictionary D containing spatial and temporal information is constructed, where the spatial dimension corresponds to different directions of the radar antenna array, and the temporal dimension corresponds to different sampling times. The ODA problem is transformed into a sparse signal reconstruction problem under the joint sparse dictionary D. The solution is obtained by solving the optimization problem min... α ||α||1s.tr(t)=Dα, obtaining the sparse representation coefficients α of the signal, and thus determining the direction of arrival (DOA) of the signal. For multipath interference, a deep learning-based multipath signal prediction model is used. This model takes historical received signals and environmental features as input, predicts multipath signal components through a trained neural network, and subtracts the predicted multipath signal components from the original received signal to generate an initial set P of spatial point clouds. init .

[0045] Finally, the initial set P of spatial point clouds is processed using a density clustering algorithm. init Dynamic target separation is performed. An improved dynamic density peak clustering algorithm is adopted, and an adaptive density threshold calculation method is introduced. For a point p in the point cloud set, its local density ρ is calculated. p , Where d(p, q) represents the Euclidean distance between points p and q, and σ is the distance scale parameter. A density threshold calculation function τ(ρ) = τ0 + β·ρ is defined, where τ0 is the base threshold and β is an adjustment coefficient. When the local density ρ at a certain point... p Greater than τ(ρ) p Furthermore, if the distance from a point to a point with higher density meets certain conditions, that point and its neighboring points are grouped into the same cluster. This algorithm traverses the entire initial set of spatial point clouds and ultimately extracts the 3D point cloud data P corresponding to the human target.human .

[0046] As an improvement to the above embodiment, the step of extracting the human body contour from the thermal imaging image to generate the thermodynamic feature tensor of the human body, and reconstructing the human body motion trajectory from the three-dimensional point cloud data to generate the three-dimensional motion feature tensor of the human body, includes the following sub-steps:

[0047] The human contour extraction algorithm based on thermal gradient is used to extract the human contour from the thermal imaging image and generate a human contour vector image containing joint nodes.

[0048] The thermal distribution model of key parts of the human body contour vector image is performed to generate a joint temperature distribution matrix;

[0049] The joint temperature distribution matrix is ​​tensorized and recombined according to the time series to generate the thermodynamic feature tensor of the human body.

[0050] Feature point matching and motion compensation are performed on the three-dimensional point cloud data to generate a discrete path point set of human motion trajectory;

[0051] The discrete path point set is smoothly interpolated using a non-uniform B-spline curve fitting algorithm to generate a continuous three-dimensional motion feature tensor of the human body.

[0052] In this embodiment, human thermodynamic and kinematic features are extracted from thermal imaging images and 3D point cloud data, respectively. For thermal imaging images, a human contour extraction algorithm based on thermal gradients, combined with adaptive gradient thresholding and human anatomy knowledge, is first used to accurately extract a human contour vector map containing joint nodes. Then, thermal distribution modeling is performed on key parts of the contour vector map, and a joint temperature distribution matrix is ​​generated using an improved Gaussian mixture model combined with spatial constraints. Next, the matrix is ​​tensorized and reorganized according to a time series, incorporating temporal difference and acceleration calculations to obtain an enhanced thermodynamic feature tensor. For 3D point cloud data, feature point matching is performed using an improved FPFH descriptor, and motion compensation is achieved using an iterative nearest-point algorithm with a weighted matrix to generate a discrete path point set. Finally, a non-uniform B-spline curve fitting algorithm is used to adjust the node spacing according to motion speed and acceleration constraints, generating a continuous 3D motion feature tensor, thus providing multi-dimensional human feature data for subsequent fall detection. Therefore, this embodiment deeply mines human feature information from thermal imaging and point cloud data through a series of improved algorithms and processing flows. Contour extraction based on thermal gradients and an improved Gaussian mixture model enhance the accuracy of human thermodynamic feature extraction. Improved feature point matching and non-uniform B-spline fitting enhance the accuracy and continuity of human motion trajectories. Compared to traditional methods, this embodiment achieves refined extraction and modeling of human thermodynamic and kinematic features, effectively overcoming the limitations of single-sensor data. It provides richer and more accurate feature representations for multi-sensor fusion-based human fall detection, significantly improving the reliability and timeliness of fall behavior recognition in complex scenarios.

[0053] For ease of understanding, this embodiment will be described in detail below:

[0054] First, a human contour extraction algorithm based on thermal gradients is used to extract the human contour from the thermal imaging image. The original thermal imaging image T acquired by the thermal imaging sensor... h (i,j), where i represents the row coordinate of the image, ranging from 1 to the total number of rows H; j represents the column coordinate of the image, ranging from 1 to the total number of columns W. Since the original image is susceptible to interference from factors such as ambient temperature fluctuations and device noise, it is preprocessed first. Multi-scale Gaussian filtering is used, employing Gaussian kernels with standard deviations of σ1 = 0.8, σ2 = 1.5, and σ3 = 2.5 respectively. Perform convolution operation on the original image This yields smoothed images at different scales. This effectively highlights the temperature difference between the human body and the background, reducing the impact of noise on subsequent contour extraction.

[0055] Next, an improved anisotropic gradient operator is used to calculate the temperature gradient. Traditional gradient calculation methods struggle to accurately capture human contour details when processing thermal imaging images. The improved operator calculates the gradient using different weights in the horizontal and vertical directions based on the directionality of the local temperature distribution in the image. For pixel (i, j), the horizontal gradient... Vertical gradient Among them, the weighting coefficient w x (k, l) and w y (k, l) will be adaptively adjusted according to the temperature difference between adjacent pixels, specifically: Here σ dir =σ0+α·var(N) ij ), σ0 is the baseline standard deviation, set to 5. It is an empirical value obtained based on a large amount of thermal imaging image data, used to ensure the calculation benchmark under normal circumstances; α is an adjustment coefficient, with a value of 0.2, used to fine-tune the standard deviation according to the actual image conditions; var(N ij represents the 8-neighbor variance of pixel (i, j), reflecting the degree of temperature variation within that neighborhood. Based on the calculated horizontal and vertical gradients, through... Calculate the gradient magnitude by Calculate the gradient direction to more accurately represent the temperature change trend of the human body contour in various directions.

[0056] Then, through adaptive thresholds Determine the contour points. Here, τ0 is the basic threshold, set to 12, which is the initial judgment standard determined through analysis and experimentation of a large amount of thermal imaging human body image data; α1 and α2 are weighting coefficients, with values ​​of 0.7 and 0.3 respectively, used to balance the influence of average gradient and information entropy on the threshold. This represents the average gradient within the neighborhood of pixel (i, j), reflecting the overall gradient level of that region; entropy(N ijThe entropy of the neighborhood information is used to measure the degree of disorder in the temperature distribution of the region. When the gradient magnitude G(i,j) of a pixel is greater than the adaptive threshold τ(i,j), the pixel is determined to be a contour point. During contour tracking, an improved Freeman chain code is used to record the contour direction. Compared with traditional methods, it adds refined rules for direction judgment and can more accurately reflect the contour direction. At the same time, combined with human anatomy knowledge, the temperature characteristics and contour morphology of human joints are analyzed in depth. For example, the temperature distribution and contour shape of the knee joint change significantly when bending and extending. By extracting and analyzing these features, the positions of 15 joint nodes, including the shoulder, elbow, and knee joints, are identified. Finally, a human contour vector map V containing the joint nodes is generated. Each joint node in this vector map is represented by precise three-dimensional coordinates (x, y, z), where the z coordinate is obtained by mapping the depth information of the thermal imaging device.

[0057] Next, heat distribution modeling of key areas is performed on the human body contour vector image. Based on the physiological structure and movement characteristics of the human body, the human body contour vector image is meticulously divided into 16 key area regions, including the head, torso, left and right upper arms, and left and right forearms. m (m = 1, 2, ..., 16). An improved Gaussian mixture model is used for each region. Modeling is performed. Where K... m For region R m The number of mixed components is set to 4 for the head due to its more complex temperature distribution, 5 for the torso, and 3 for the limbs; π m,k It is the weight of the k-th component; μ m,k The mean vector represents the center of the temperature distribution; ∑ m,k This is a diagonal covariance matrix, reflecting the dispersion of the temperature distribution. To further improve the accuracy of temperature distribution modeling in the joint region, a spatial constraint term is introduced. Where λ is the constraint strength coefficient, used to control the strength of spatial constraints; μ joint This corresponds to the temperature value of the joint; d k,joint Let be the distance from the center of the k-th Gaussian component to the joint; σ is the spatial influence range parameter, with a value of 10 pixels. Through multiple iterations using the Expectation-Maximization (EM) algorithm, the model parameters are continuously optimized to better reflect the actual temperature distribution of each key component, ultimately generating a 15×5 joint temperature distribution matrix M. temp In this matrix, each row corresponds to a joint, and each column represents the average temperature, temperature standard deviation, temperature gradient, temperature change rate, and temperature entropy, respectively, providing a quantitative description of the thermal distribution at the joint from multiple dimensions.

[0058] Next, the joint temperature distribution matrix is ​​tensor-reconstructed according to the time series. Within T consecutive time frames, a joint temperature distribution matrix M is acquired at each time point t. temp (t). To effectively capture the temporal characteristics of temperature changes, a time window mechanism is introduced, with a window size of W = 7 frames and a step size of 1 frame. For time points satisfying t ≥ W, according to formula T... thermo (t)[:,:,w]=M temp (t-W+w+1), where w = 0, 1, ..., W-1, generates a thermodynamic characteristic tensor. Subsequently, the tensor is subjected to time-domain difference processing to obtain ΔT. thermo (t)[:,:,w]=T thermo (t)[:,:,w]-T thermo (t)[:,:,w-1], w=1,...,W-1, to obtain temperature change information between adjacent time frames. Then, the acceleration due to temperature change is further calculated, i.e. Finally, the original temperature tensor, difference tensor, and acceleration tensor are concatenated to obtain the enhanced thermodynamic characteristic tensor. Its dimensions are 15×5×(3W-4), and this tensor can comprehensively and meticulously present the dynamic changes in temperature of human joints over time.

[0059] For 3D point cloud data, feature point matching and motion compensation are performed. First, for two consecutive frames of point cloud P... t and P t-1 Feature point extraction. A curvature- and normal-direction-based method is employed, selecting points with significant curvature changes and stable normal directions as feature points. This is because during human movement, the curvature and normal direction of the point cloud at key locations such as joints undergo significant changes, allowing for the accurate extraction of key feature points closely related to human movement. In the feature point matching stage, an improved FPFH (Fast Point Feature Histograms) descriptor is used. (Improved...) SPFH stands for Simplified Point Feature Histogram, used to describe the local geometric features of a single point; N(p i Let p be a point. i The k-nearest neighbors, where k takes the value 10; For distance weights, σ d The distance attenuation parameter is set to 0.1 meters. This descriptor integrates information from a single point and its nearest neighbors, adjusting the contribution of neighbors by distance weights, thus significantly improving the accuracy of feature point matching. Motion compensation employs an improved Iterative Closest Point (ICP) algorithm, introducing a weight matrix W, and minimizing... Solve for the rotation matrix R and the translation vector t. Where p i ∈Pt q i ∈P t-1 To match point pairs; w i The credibility weights are determined by both descriptor similarity and geometric constraints. η1 and η2 are weighting coefficients, d desc To describe the sub-distance, d geo The geometric distance is used. Through the above improvements, the point cloud offset problem caused by sensor motion or changes in human posture can be effectively addressed, generating a discrete path point set of the human motion trajectory. Each point accurately represents the position of the human body's center of mass in three-dimensional space.

[0060] Finally, a non-uniform B-spline curve fitting algorithm is used to smoothly interpolate the discrete path point set. Traditional B-spline curves often fail to achieve ideal fitting results when processing human motion trajectories due to the non-uniformity of human motion speed. The improved non-uniform B-spline curve fitting algorithm dynamically adjusts the node spacing based on the motion speed, and the curve equation is... Where d i To control the vertices, N i,k (u) is a k-th degree B-spline basis function, and the node vector U = [u0, u1, ..., u] is a k-th degree B-spline basis function. n+k+1 ]. Node spacing Δu i The calculation method is as follows Where v i =||s i+1 -s i || / Δt represents the motion velocity, Δ is the velocity influence factor, and β is the minimum node spacing. Simultaneously, an acceleration constraint term is introduced. Ensure curve smoothness, where a i For the curve in parameter u i acceleration at that point Let λ be the average acceleration and λ be the constraint strength. This is achieved by minimizing the energy function E = Ω. data +Ω acc Solve for the optimal control vertex, where Ω data This is used to measure how well a curve fits a discrete set of path points. The final result is a continuous 3D human motion feature tensor T with dimensions T×15×3. kinematic Each slice accurately represents the three-dimensional coordinates of 15 joints at a specific moment, thus presenting the changes in the human body's movement trajectory in three-dimensional space in a precise and smooth manner.

[0061] As an improvement to the above embodiment, the step of performing thermodynamic spatiotemporal gradient calculation on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features reflecting the rate of change of body temperature distribution of the human body, and performing kinematic spatiotemporal gradient calculation on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features reflecting the intensity of abrupt changes in posture of the human body, includes the following sub-steps:

[0062] Perform a first-order difference operation on the thermodynamic feature tensor in the time dimension to generate the temperature change rate tensor at each key point.

[0063] The temperature change rate tensor at each joint point is spatially gradient convolved using the three-dimensional Sobel operator to extract the thermodynamic spatiotemporal gradient features that reflect the rate of change of the body temperature distribution of the human body.

[0064] The three-dimensional motion feature tensor is decomposed and synthesized into velocity vectors to generate temporal data of motion components for each limb segment;

[0065] The sliding window standard deviation detection algorithm is used to perform acceleration mutation analysis on the time series data of the motion components, and generate kinematic spatiotemporal gradient features that reflect the intensity of the posture mutation of the human body.

[0066] In this embodiment, based on the acquired thermodynamic feature tensor and three-dimensional motion feature tensor, the dynamic changes in human body temperature and posture are deeply explored from the spatiotemporal dimensions. For the thermodynamic feature tensor, first-order difference operations are performed in the time dimension to obtain the temperature change rate tensor at each joint point, capturing the temperature change over time. Then, a spatial gradient convolution is performed on the temperature change rate tensor using an improved three-dimensional Sobel operator to further analyze the temperature change trend from the spatial dimension, thereby generating thermodynamic spatiotemporal gradient features that comprehensively reflect the rate of change in human body temperature distribution. For the three-dimensional motion feature tensor, velocity vector decomposition and synthesis are performed first to decompose the overall motion into temporal data of motion components of each limb segment, refining the motion features. Then, an improved sliding window standard deviation detection algorithm, combined with a dynamic threshold adjustment mechanism, is used to perform acceleration mutation analysis on the temporal data of the motion components, accurately capturing the abrupt changes in human posture and generating kinematic spatiotemporal gradient features that reflect the intensity of posture mutations. Through the above steps, the deep extraction of thermodynamic and kinematic dynamic changes in human movement is achieved, providing key information for subsequent fusion analysis. In summary, this embodiment overcomes the limitations of traditional methods that analyze data from only a single dimension by calculating the spatiotemporal gradients of thermodynamic and kinematic characteristics. The combination of first-order difference and the three Sobel operators effectively enhances the spatiotemporal perception of changes in human body temperature distribution; velocity vector decomposition combined with an improved sliding window standard deviation detection algorithm significantly improves the accuracy of capturing sudden changes in human posture. Compared with existing technologies, this embodiment can more meticulously and accurately characterize the changes in thermodynamic and kinematic dynamic features during human movement, providing more discriminative feature basis for human fall detection based on multi-sensor fusion, improving the ability to identify and detect fall behavior in complex scenarios, and reducing the probability of false alarms and false negatives.

[0067] Specifically, the following examples illustrate this embodiment:

[0068] First, a first-order difference operation is performed on the thermodynamic characteristic tensor along the time dimension to generate the temperature change rate tensor at each key point. (Thermodynamic characteristic tensor) It contains 15 key points, 5 temperature feature dimensions, and multiple frames of data within a time window. For each element in the tensor... Where j represents the keypoint index (j∈[1, 15]), f represents the temperature feature dimension index (f∈[1, 5]), and w represents the frame index within the time window (w∈[1, 3W-4]), the formula is used to... (When w > 1) Perform first-order difference operations. In this formula, The tensor element values ​​corresponding to the current frame key point j and the temperature feature dimension f are: The element value at the corresponding position in the previous frame is used as the reference point; the difference between the two values ​​yields the temperature change at the keypoint within that temperature feature dimension. By iterating through the entire thermodynamic feature tensor along the time dimension, a temperature change rate tensor is generated. This tensor reflects the temperature change of each joint point over time.

[0069] Next, spatial gradient convolution is performed on the temperature change rate tensor of each joint point using a 3D Sobel operator to extract thermodynamic spatiotemporal gradient features. The traditional Sobel operator is mostly used for 2D images; here, it is extended to 3D space. A 3D Sobel operator kernel S is constructed with three directions (x, y, z) corresponding to the temperature feature dimension, joint point, and time. x S y S z Taking the x-direction as an example, its operator kernel performs convolution operations along the temperature feature dimension. For the temperature change rate tensor ΔT... rate Each element in the matrix is ​​convolved with its corresponding Sobel kernel in its 3D neighborhood. Taking the calculation of the gradient in the x-direction as an example... Where S x (i, k, l) are the coefficients of the three-dimensional Sobel operator kernel in the x-direction at position (i, k, l), ΔT rate (j+i, f+k, w+l) represents the element values ​​at the corresponding neighborhood positions of the temperature change rate tensor. Similarly, the gradients G in the y and z directions are calculated. y G z Then through the formula Calculate the gradient magnitude to obtain the thermodynamic spatiotemporal gradient characteristics that reflect the rate of change in human body temperature distribution. This feature comprehensively describes the rate and trend of temperature changes at various joints of the human body from a spatiotemporal perspective.

[0070] Then, velocity vector decomposition and synthesis are performed on the three-dimensional motion feature tensor to generate temporal data of the motion components of each limb segment. (Three-dimensional motion feature tensor) The three-dimensional coordinate information of 15 key points at T time points was recorded. For the coordinate changes of each key point at adjacent time points, the coordinates (x, y) of key point j at times t and t+1 are recorded. j,t y j,t , z j,t ) and (x j,t+1 y j,t+1 , z j,t+1 For example, using the formula Calculate the velocity vector, where Δt is the time interval. Decompose the velocity vector into components in the x, y, and z directions in the Cartesian coordinate system. Since the human limbs are composed of multiple joints, in order to obtain the motion components of each limb segment, the velocity components of relevant joints are synthesized based on the human anatomical structure. For example, for the upper arm segment, the velocity components in the corresponding directions of the shoulder and elbow joints are weighted and summed according to certain weights (determined based on limb length and motion correlation) to obtain the time-series data of the upper arm's motion components in three directions. By performing the above operations on all limb segments, a set V of motion component time-series data for each limb segment is generated. segments This collection comprehensively records the movement speed information of various limbs of the human body at different times.

[0071] Finally, a sliding window standard deviation detection algorithm is used to perform acceleration abrupt change analysis on the motion component time series data to generate kinematic spatiotemporal gradient features. The sliding window size is set to m (based on the general temporal characteristics of human motion, a value of 5-8 time points is taken). For each directional component sequence (e.g., ...) in the motion component time series data of each limb segment... ), within the window, through formulas Calculate the standard deviation, where This represents the average value of the velocity components within the window. This represents the velocity component values ​​at each time point within the window. The standard deviation reflects the dispersion of the velocity data within the window. When a person experiences a sudden change in posture, such as a fall, the velocity change intensifies, and the standard deviation increases significantly. To more accurately detect sudden acceleration changes, a threshold adjustment mechanism is introduced, setting a dynamic threshold. This represents the average standard deviation of the component in that direction over a certain time period. The standard deviation is the mean of the standard deviations up to the current time, and k is an adjustment coefficient (ranging from 1.5 to 2.5). When the calculated standard deviation... Exceeding the dynamic threshold At that time, a sudden acceleration change was considered to have occurred. By performing the above detection on the temporal data of the motion components in three directions for each limb segment, a kinematic spatiotemporal gradient feature G reflecting the intensity of the sudden change in human posture was generated. kinematic This feature can effectively capture sudden changes in human posture during movement, providing crucial information for subsequent fall detection.

[0072] As an improvement to the above embodiments, the step of constructing a spatiotemporal graph structure of human motion based on the fusion of the thermodynamic spatiotemporal gradient features and the kinematic spatiotemporal gradient features includes the following sub-steps:

[0073] The thermodynamic spatiotemporal gradient features are mapped to a graph node attribute set, wherein each node in the graph node attribute set corresponds to a specific human joint.

[0074] The kinematic spatiotemporal gradient features are converted into a graph edge weight set, where each edge of the graph edge weight set represents the strength of the dynamic association between joints;

[0075] Construct an initial topological skeleton based on the standard human skeletal connections;

[0076] The graph node attribute set is injected into the nodes of the initial topological graph skeleton, and the graph edge weight set is injected into the corresponding edges of the initial topological graph skeleton to generate a spatiotemporal graph structure of human motion.

[0077] In this embodiment, based on thermodynamic and kinematic spatiotemporal gradient features, a multi-dimensional representation of human motion features is achieved through graph structure modeling. First, the thermodynamic spatiotemporal gradient features are mapped to graph node attributes corresponding to human joints, giving each node attribute information reflecting joint temperature changes. Then, based on the dynamic relationships between joints, the kinematic spatiotemporal gradient features are converted into graph edge weights, quantifying the degree of mutual influence between joint movements. Next, an initial topological graph skeleton is built based on the standard human skeletal connection relationships, constructing the basic framework for human joint connections. Finally, node attributes and edge weights are injected into the initial topological graph to form a complete spatiotemporal graph structure for human motion, organically integrating the thermodynamic and kinematic features of human motion into a unified graph model, providing a structured data foundation for subsequent feature analysis and fall detection. Therefore, this embodiment achieves deep integration and efficient expression of thermodynamic and kinematic features through spatiotemporal graph structure construction. Compared to traditional data fusion methods, transforming features into graph node attributes and edge weights not only intuitively reflects the individual characteristics of human joints and the inter-joint relationships, but also leverages the inherent advantages of graph structures, facilitating subsequent feature extraction and analysis using deep learning models such as graph convolutional networks. This structured data representation significantly improves the efficiency of multimodal feature fusion and information utilization, enhances the ability to describe complex human movement patterns, and thus provides a superior data format for accurately identifying fall behavior, effectively improving the accuracy and reliability of fall detection.

[0078] The following is a detailed description of this embodiment:

[0079] First, the thermodynamic spatiotemporal gradient features are mapped to a set of graph node attributes. (Thermodynamic spatiotemporal gradient features) It includes 15 key points, 5 temperature feature dimensions, and multiple frames of data within a time window. For each key point, its thermodynamic spatiotemporal gradient feature vector within the time window is extracted. Taking key point j as an example, its feature vector attr... j =[G thermo (j, 1, 1), G thermo (j, 2, 1), ..., G thermo[j, 5, 3W-5], this vector integrates the temperature change rate information of this joint point under different temperature feature dimensions and time points. Since each node corresponds to a specific human joint, the feature vectors of the 15 joint points are arranged sequentially to form a graph node attribute set Attr = [attr1, attr2, ..., attr 15 This allows each node to have an attribute description that reflects the changes in the body temperature distribution of the corresponding joint.

[0080] Next, the kinematic spatiotemporal gradient features are transformed into a graph edge weight set. Kinematic spatiotemporal gradient feature G kinematic This reflects the correlation of postural abrupt changes between various joints in the human body. Based on the connection relationships between the joints, joint pairs (j1, j2) are identified (e.g., shoulder and elbow joints). For each joint pair, graph edge weights are determined using an improved correlation strength calculation method. Let the kinematic spatiotemporal gradient feature vectors of joints j1 and j2 be respectively... and Using formula The dynamic correlation strength between joints is calculated, where dist(·) is the Euclidean distance function, used to measure the difference between two feature vectors; α is an adjustment coefficient (ranging from 0.5 to 1.5), used to adjust the influence of distance on the weights. This formula results in larger edge weights for feature vectors that are closer together, indicating stronger dynamic correlation. All joint pairs are traversed to generate a graph edge weight set, where the weight of each edge represents the dynamic correlation strength between the corresponding joints.

[0081] Then, an initial topological graph skeleton is constructed based on the standard human skeletal connection relationships. The human skeleton has a fixed connection structure, such as the shoulder joint connecting to the upper arm bones, and the elbow joint connecting the upper arm bones and forearm bones. Based on this characteristic, nodes represent joints, and edges represent the connections between joints, constructing an undirected graph G0 = (V, E). The node set V contains 15 nodes, corresponding to 15 human joints; the edge set E is determined according to the standard human skeletal connection relationships, for example, containing edge (1,2) (assuming 1 is the shoulder joint and 2 is the joint point of the upper arm near the shoulder joint). This initial topological graph skeleton only reflects the joint connections and does not yet include specific kinematic and thermodynamic characteristics.

[0082] Finally, the graph node attribute set is injected into the nodes of the initial topological graph skeleton, and the graph edge weight set is injected into the corresponding edges, generating the spatiotemporal graph structure of human motion. For each node v of the initial topological graph skeleton G0... j ∈V, the key attribute vector attr corresponding to the graph node attribute set. j This node is assigned a thermodynamic spatiotemporal gradient characteristic description; for each edge Set the weights corresponding to the graph edge weights This edge is assigned to reflect the strength of the dynamic connection between joints. After this operation, a complete spatiotemporal graph structure of human motion, G = (V, E, Attr, Weight), is formed. This graph structure integrates thermodynamic and kinematic feature information from two levels: node attributes and edge weights. It can comprehensively characterize the spatiotemporal characteristics of the human body during movement, providing a structured data foundation for subsequent fall feature analysis through graph convolutional networks.

[0083] As an improvement to the above embodiment, the step of inputting the spatiotemporal graph structure into a spatiotemporal graph convolutional network, performing human fall feature evolution through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network, and outputting the fall risk probability includes the following sub-steps:

[0084] The spatial graph convolutional layer of the spatiotemporal graph convolutional network is used to aggregate neighborhood features of the spatiotemporal graph structure to generate a spatially enhanced node feature matrix.

[0085] The node feature matrix is ​​modeled in a multi-scale temporal manner by the dilated temporal convolutional layer of the spatiotemporal graph convolutional network to extract the spatiotemporal joint feature vector.

[0086] The spatiotemporal joint feature vector is dynamically reweighted using a learnable attention mechanism to generate a fall feature vector that focuses on key joints.

[0087] The fully connected neural network layer of the spatiotemporal graph convolutional network performs nonlinear mapping on the fall feature vector and outputs the original fall risk probability value.

[0088] The original fall risk probability values ​​within a continuous time window are processed by median filtering to generate the final fall risk probability.

[0089] In this embodiment, a spatiotemporal graph convolutional network is used to perform deep feature mining and analysis on the constructed spatiotemporal graph structure of human movement. First, through a spatial graph convolutional layer, based on an adaptive weight aggregation method, neighborhood features are aggregated according to the connection weights between nodes in the graph structure to enhance the feature expression of nodes in the spatial dimension and generate a spatially enhanced node feature matrix. Next, with the variable dilation rate mechanism of the dilated temporal convolutional layer, the node feature matrix is ​​convolved at multiple scales in the temporal dimension to extract temporal information at different time scales and fuse it with spatial features to form a spatiotemporal joint feature vector. Then, a learnable multi-head attention mechanism is used to dynamically reweight the spatiotemporal joint feature vector from multiple angles, focusing on key joint features related to falls to generate a fall feature vector. Then, a fully connected neural network layer is used to perform nonlinear mapping on the fall feature vector to initially output the original fall risk probability value. Finally, median filtering is used to process the original probability value within a continuous time window to eliminate fluctuation interference and obtain the final stable fall risk probability, realizing a complete process from spatiotemporal graph structure to fall risk quantitative assessment. In summary, this embodiment significantly improves the extraction and analysis capabilities of human fall features through a multi-level feature interaction mechanism. The adaptive aggregation of the spatial graph convolutional layer and the multi-scale modeling of the dilated temporal convolutional layer effectively integrate the spatiotemporal information of human movement; the learnable attention mechanism precisely focuses on key joints, enhancing the sensitivity to fall features; and median filtering improves the stability of the risk probability output.

[0090] To facilitate understanding of this embodiment, the following example is provided:

[0091] First, neighborhood features of the spatiotemporal graph structure are aggregated through the spatial graph convolutional layer of the spatiotemporal graph convolutional network to generate a spatially enhanced node feature matrix. The spatial graph convolutional layer employs an improved adaptive weight aggregation method, which aggregates weights for each node v in the spatiotemporal graph structure G = (V, E, Attr, Weight). i ∈V, its initial feature vector is attr i During the aggregation process, based on node v i Its neighboring node v j (j∈N(i), N(i) is node v) i The weight w of the edges between the neighboring nodes of the set) ij ∈Weight and the node's own characteristics, through the formula Perform aggregation calculations. Where W... (l) b is the weight matrix of the l-th spatial graph convolutional layer, used to perform linear transformation on the features of neighboring nodes; (l) σ is the corresponding bias vector; σ is the activation function, which uses the LeakyReLU function to avoid the gradient vanishing problem during network training. To normalize the weights and ensure that the contribution ratio of each neighboring node to the feature update of the central node is reasonable, the formula is used to calculate the weights of node v. i It can dynamically aggregate neighborhood information based on the features and connection weights of neighboring nodes to generate updated feature vectors. Arrange the updated feature vectors of all nodes in order to form a spatially augmented node feature matrix. Where |V| is the number of nodes, d (l) Let be the dimension of the feature vector of the l-th layer.

[0092] Next, multi-scale temporal modeling of the node feature matrix is ​​performed using dilated temporal convolutional layers in a spatiotemporal graph convolutional network to extract spatiotemporal joint feature vectors. The dilated temporal convolutional layers introduce a variable dilation rate mechanism to capture human motion features at different time scales. For the spatially enhanced node feature matrix H... (l) Convolution is performed in the time dimension. Let the input feature matrix sequence be... Where T is the number of time steps. For the feature matrix at time step t... Through convolution kernel K (m) (where m represents different convolution kernels, corresponding to different dilation rates) Convolution calculation is performed, using the formula: Where *d represents the dilated convolution operation, and r is the dilation rate. Different dilation rates are set according to different values ​​of m (such as r1=1, r2=2, r3=4) so ​​that the convolution operation can span different number of time steps and capture multi-scale temporal information. This is the feature vector after convolution by the m-th convolution kernel. The feature vectors obtained from convolutions with different dilation rates are concatenated, i.e. (M is the total number of convolutional kernels), and after linear transformation and activation function processing, the spatiotemporal joint feature vector z is obtained. t The spatiotemporal joint feature vectors of all time steps are combined to form the feature sequence Z = [z1; z2; ...; z...]. T ].

[0093] Then, a learnable attention mechanism is used to dynamically reweight the spatiotemporal joint feature vectors, generating fall feature vectors focused on key joints. The learnable attention mechanism employs a multi-head attention structure to analyze the spatiotemporal joint feature vectors from multiple perspectives. For each spatiotemporal joint feature vector z in the feature sequence Z... t First, a query vector is generated through multiple linear transformations. key vector Sum value vector (k = 1, 2, ..., K, where K is the number of heads), that is in This is a learnable weight matrix. Then, the attention score is calculated. in and Let d be the i-th and j-th query vector and key vector, respectively. k Let be the dimension of the key vector, and · denote the vector dot product. The attention scores are normalized using the Softmax function to obtain the attention weights. Finally, the value vectors are weighted and summed according to the attention weights to obtain the output of each head. The outputs of multiple heads are concatenated and subjected to a linear transformation, i.e. Obtain the fall feature vector f focusing on key joints t Aggregate the fall feature vectors from all time steps to form the final feature representation F = [f1; f2; ...; f T ].

[0094] Next, a fully connected neural network layer based on a spatiotemporal graph convolutional network performs a nonlinear mapping on the fall feature vector, outputting the original fall risk probability value. The fully connected neural network layer consists of multiple fully connected layers, inputting the aggregated feature representation F into the fully connected neural network. First, feature transformation is performed through the first fully connected layer h1 = σ(W1F + b1), where W1 is the weight matrix, b1 is the bias vector, and σ is the activation function (using the ReLU function). Then, further feature extraction and transformation are performed through several similar fully connected layers, finally outputting through the output layer p = Sigmoid(W1F + b1). o h n +b o Output the original fall risk probability value p, where the Sigmoid function maps the output value to the interval [0,1], representing the probability of falling.

[0095] Finally, median filtering is applied to the raw fall risk probability values ​​within the continuous time window to generate the final fall risk probability. To reduce the fluctuation of the raw fall risk probability values ​​and improve the stability of the detection results, a time window size of S (e.g., S=5) is set. For each time point t, the raw fall risk probability values ​​p of the previous S-1 time points and the current time point are taken. t-S+1 p t-S+2 , ..., p t Arrange these probability values ​​in ascending order and take the middle value as the final probability of falling. Median filtering effectively removes unreasonable probability values ​​caused by noise or transient anomalies, making the final output fall risk probability more accurate and stable, and providing a reliable basis for subsequent fall judgment.

[0096] As an improvement to the above embodiment, the step of performing a human fall warning operation when the probability of falling exceeds a preset threshold includes the following sub-steps:

[0097] The probability of falling is compared with a preset threshold. If the probability of falling exceeds the preset threshold, an alert signal is generated that includes at least one of vibration command, flashing command, and voice command, based on the preset device type and user preference settings.

[0098] The generated reminder signal is sent to the preset terminal device to provide an alert.

[0099] In this embodiment, firstly, the fall risk probability output by the deep residual network is compared with a preset threshold to determine whether an alarm condition is triggered, ensuring that the triggering of the reminder signal has a clear quantitative basis. When the fall risk probability is detected to exceed the threshold, a reminder signal containing at least one of vibration, flashing light, and voice command is selectively generated based on the preset device type (such as mobile phone, smart bracelet, home alarm, etc.) and user personalized preferences to meet different usage scenarios and user needs. Finally, the generated reminder signal is sent to a preset terminal device, completing a complete closed loop from detection to reminder, and realizing timely notification of fall events.

[0100] To facilitate understanding of this embodiment, the following examples are provided:

[0101] After the system outputs the fall risk probability based on the deep residual network, it first compares this probability value with a pre-set threshold. The pre-set threshold can be flexibly adjusted according to the actual application scenario and user needs. For example, in a home-based elderly care scenario, it can be set to 0.7 to balance false alarms and missed alarms. If the fall risk probability exceeds the threshold, the system will immediately activate the alert signal generation mechanism. This mechanism operates based on pre-stored device type and user preference settings: for smartphone terminals, if the user preference is set to vibration priority, the system will generate vibration command codes of specific frequency and intensity; if the user preference is flashing alerts, it will generate commands to control the flashing frequency and color of the phone's flashlight; if the user selects voice alerts, the system will retrieve pre-set warning voice clips from the voice library, such as "Fall risk detected, please check immediately," and supports multi-language switching. These commands can be generated individually or in combination according to user settings. After generating the alert signal, the system sends the signal data packet containing vibration, flashing, and voice commands to the preset terminal device through a wireless network communication module, such as Wi-Fi, 4G / 5G, etc. After receiving the data packet, the terminal device parses the instruction code within it and drives the corresponding hardware module to perform reminder operations, such as activating the phone's vibration motor, controlling the flashlight to blink, and playing a warning voice, thereby promptly notifying relevant personnel to pay attention to potential fall risks.

[0102] See Figure 2 This is a schematic diagram of a human fall detection system based on multi-sensor fusion according to an embodiment of the present invention. The human fall detection system based on multi-sensor fusion includes:

[0103] The acquisition module 10 allows the user to detect the temperature distribution of the surrounding human body in real time through a thermal imaging sensor, generate a thermal imaging image of the human body, and simultaneously transmit high-frequency electromagnetic waves to the surrounding human body through a millimeter-wave radar and receive reflected signals to generate three-dimensional point cloud data of the human body.

[0104] The feature extraction module 11 is used to extract the human body contour from the thermal imaging image, generate the thermodynamic feature tensor of the human body, and reconstruct the human body motion trajectory from the three-dimensional point cloud data, generating the three-dimensional motion feature tensor of the human body.

[0105] The feature calculation module 12 is used to perform thermodynamic spatiotemporal gradient calculation on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features that reflect the rate of change of body temperature distribution of the human body, and to perform kinematic spatiotemporal gradient calculation on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features that reflect the intensity of abrupt changes in the posture of the human body.

[0106] Module 13 is used to construct a spatiotemporal graph structure of human motion based on the fusion of the thermodynamic spatiotemporal gradient features and the kinematic spatiotemporal gradient features.

[0107] Prediction module 14 is used to input the spatiotemporal graph structure into the spatiotemporal graph convolutional network, perform human fall feature evolution through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network, and output the fall risk probability.

[0108] The reminder module 15 is used to perform a human fall reminder operation when the probability of falling exceeds a preset threshold.

[0109] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0110] By collaboratively acquiring multimodal data from thermal imaging sensors and millimeter-wave radar, the body's temperature distribution and three-dimensional motion information are obtained. The thermal imaging images are used to extract the human contour and generate a thermodynamic feature tensor. Simultaneously, the millimeter-wave point cloud data is used to reconstruct the motion trajectory and generate a three-dimensional motion feature tensor, achieving parallel extraction of physiological and motion features. Spatiotemporal gradient calculations are performed on both types of feature tensors to quantify the rate of change in body temperature distribution and the intensity of abrupt posture changes, thereby capturing the biomechanical anomalies of the fall process and obtaining dual-modal gradient features of thermodynamic and kinematic spatiotemporal gradients. These dual-modal gradient features are fused to construct a spatiotemporal graph structure. The hierarchical feature interaction of the spatiotemporal graph convolutional network enables dynamic evolution recognition of fall features, ultimately triggering an alert based on a probability threshold. In summary, this invention, through the fusion of thermodynamic and kinematic spatiotemporal features and the structured representation capabilities of graph neural networks, effectively solves the technical problems of high false alarm rates from single sensors and the inability of simple multi-sensor stitching to capture the essential features of falls in the background technology, achieving accurate recognition and reliable response to human fall behavior in complex scenarios. Therefore, the embodiments of the present invention, through the fusion of multimodal spatiotemporal features of thermal imaging and millimeter-wave radar and graph structure evolution analysis, can accurately identify human fall behavior in complex scenarios, and achieve efficient and reliable fall detection and timely response.

[0111] As an improvement to the above embodiments, the acquisition module is specifically used for:

[0112] Adaptive noise filtering is applied to the raw temperature data acquired by the thermal imaging sensor to generate a noise-reduced temperature distribution matrix.

[0113] A multi-scale segmentation algorithm is used to identify human target regions in the denoised temperature distribution matrix, generating an accurate binary mask that separates the human body from the background.

[0114] Based on the binary mask, spatial domain filtering is performed on the original temperature data to extract the temperature distribution data of the target area of ​​the human body and generate a thermal imaging image of the human body.

[0115] The direction of arrival (DOA) of the reflected signal received by the millimeter-wave radar is estimated and multipath interference is eliminated to generate an initial set of spatial point clouds.

[0116] The initial set of spatial point clouds is dynamically separated using a density clustering algorithm to extract three-dimensional point cloud data corresponding to human targets.

[0117] As an improvement to the above embodiments, the feature extraction module is specifically used for:

[0118] The human contour extraction algorithm based on thermal gradient is used to extract the human contour from the thermal imaging image and generate a human contour vector image containing joint nodes.

[0119] The thermal distribution of key parts of the human body contour vector image is modeled to generate a joint temperature distribution matrix.

[0120] The joint temperature distribution matrix is ​​tensorized and recombined according to the time series to generate the thermodynamic feature tensor of the human body.

[0121] Feature point matching and motion compensation are performed on the three-dimensional point cloud data to generate a discrete path point set of human motion trajectory;

[0122] The discrete path point set is smoothly interpolated using a non-uniform B-spline curve fitting algorithm to generate a continuous three-dimensional motion feature tensor of the human body.

[0123] As an improvement to the above embodiments, the feature calculation module is specifically used for:

[0124] Perform a first-order difference operation on the thermodynamic feature tensor in the time dimension to generate the temperature change rate tensor at each key point.

[0125] The temperature change rate tensor at each joint point is spatially gradient convolved using the three-dimensional Sobel operator to extract the thermodynamic spatiotemporal gradient features that reflect the rate of change of the body temperature distribution of the human body.

[0126] The three-dimensional motion feature tensor is decomposed and synthesized into velocity vectors to generate temporal data of motion components for each limb segment;

[0127] The sliding window standard deviation detection algorithm is used to perform acceleration mutation analysis on the time series data of the motion components, and generate kinematic spatiotemporal gradient features that reflect the intensity of the posture mutation of the human body.

[0128] As an improvement to the above embodiments, the construction module is specifically used to: map the thermodynamic spatiotemporal gradient features into a graph node attribute set, wherein each node in the graph node attribute set corresponds to a specific human joint;

[0129] The kinematic spatiotemporal gradient features are converted into a graph edge weight set, where each edge of the graph edge weight set represents the strength of the dynamic association between joints;

[0130] Construct an initial topological skeleton based on the standard human skeletal connections;

[0131] The graph node attribute set is injected into the nodes of the initial topological graph skeleton, and the graph edge weight set is injected into the corresponding edges of the initial topological graph skeleton to generate a spatiotemporal graph structure of human motion.

[0132] As an improvement to the above embodiments, the prediction module is specifically used to: perform neighborhood feature aggregation on the spatiotemporal graph structure through the spatial graph convolutional layer of the spatiotemporal graph convolutional network to generate a spatially enhanced node feature matrix;

[0133] The node feature matrix is ​​modeled in a multi-scale temporal manner by the dilated temporal convolutional layer of the spatiotemporal graph convolutional network to extract the spatiotemporal joint feature vector.

[0134] The spatiotemporal joint feature vector is dynamically reweighted using a learnable attention mechanism to generate a fall feature vector that focuses on key joints.

[0135] The fully connected neural network layer of the spatiotemporal graph convolutional network performs nonlinear mapping on the fall feature vector and outputs the original fall risk probability value.

[0136] The original fall risk probability values ​​within a continuous time window are processed by median filtering to generate the final fall risk probability.

[0137] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0138] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for human fall detection based on multi-sensor fusion, characterized in that, include: The thermal imaging sensor detects the temperature distribution of the surrounding human body in real time, generates a thermal imaging image of the human body, and simultaneously transmits high-frequency electromagnetic waves to the surrounding human body through millimeter-wave radar and receives reflected signals to generate three-dimensional point cloud data of the human body. Human contours are extracted from the thermal imaging image to generate a thermodynamic feature tensor of the human body, and human motion trajectory is reconstructed from the three-dimensional point cloud data to generate a three-dimensional motion feature tensor of the human body. Thermodynamic spatiotemporal gradient calculation is performed on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features that reflect the rate of change of body temperature distribution of the human body, and kinematic spatiotemporal gradient calculation is performed on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features that reflect the intensity of abrupt changes in the posture of the human body. The spatiotemporal gradient features of the thermodynamics and the spatiotemporal gradient features are fused to construct the spatiotemporal graph structure of human motion; The spatiotemporal graph structure is input into the spatiotemporal graph convolutional network, and the human fall feature evolution is performed through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network to output the fall risk probability. When the probability of falling exceeds a preset threshold, a human fall warning operation is executed.

2. The human fall detection method based on multi-sensor fusion as described in claim 1, characterized in that, The process involves using a thermal imaging sensor to detect the temperature distribution of the surrounding human body in real time, generating a thermal imaging image of the human body, and simultaneously transmitting high-frequency electromagnetic waves to the surrounding human body via millimeter-wave radar and receiving reflected signals to generate three-dimensional point cloud data of the human body. This includes the following sub-steps: Adaptive noise filtering is applied to the raw temperature data acquired by the thermal imaging sensor to generate a noise-reduced temperature distribution matrix. A multi-scale segmentation algorithm is used to identify human target regions in the denoised temperature distribution matrix, generating an accurate binary mask that separates the human body from the background. Based on the binary mask, spatial domain filtering is performed on the original temperature data to extract the temperature distribution data of the target area of ​​the human body and generate a thermal imaging image of the human body. The direction of arrival (DOA) of the reflected signal received by the millimeter-wave radar is estimated and multipath interference is eliminated to generate an initial set of spatial point clouds. The initial set of spatial point clouds is dynamically separated using a density clustering algorithm to extract three-dimensional point cloud data corresponding to human targets.

3. The human fall detection method based on multi-sensor fusion as described in claim 2, characterized in that, The process of extracting human contours from the thermal imaging image to generate a thermodynamic feature tensor of the human body, and reconstructing human motion trajectory from the 3D point cloud data to generate a 3D motion feature tensor of the human body, includes the following sub-steps: The human contour extraction algorithm based on thermal gradient is used to extract the human contour from the thermal imaging image and generate a human contour vector image containing joint nodes. The thermal distribution model of key parts of the human body contour vector image is performed to generate a joint temperature distribution matrix; The joint temperature distribution matrix is ​​tensorized and recombined according to the time series to generate the thermodynamic feature tensor of the human body. Feature point matching and motion compensation are performed on the three-dimensional point cloud data to generate a discrete path point set of human motion trajectory; The discrete path point set is smoothly interpolated using a non-uniform B-spline curve fitting algorithm to generate a continuous three-dimensional motion feature tensor of the human body.

4. The human fall detection method based on multi-sensor fusion as described in claim 3, characterized in that, The step of performing thermodynamic spatiotemporal gradient calculation on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features reflecting the rate of change of body temperature distribution of the human body, and performing kinematic spatiotemporal gradient calculation on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features reflecting the intensity of abrupt changes in the posture of the human body, includes the following sub-steps: Perform a first-order difference operation on the thermodynamic feature tensor in the time dimension to generate the temperature change rate tensor at each key point. The temperature change rate tensor at each joint point is spatially gradient convolved using the three-dimensional Sobel operator to extract the thermodynamic spatiotemporal gradient features that reflect the rate of change of the body temperature distribution of the human body. The three-dimensional motion feature tensor is decomposed and synthesized into velocity vectors to generate temporal data of motion components for each limb segment; The sliding window standard deviation detection algorithm is used to perform acceleration mutation analysis on the time series data of the motion components, and generate kinematic spatiotemporal gradient features that reflect the intensity of the posture mutation of the human body.

5. The human fall detection method based on multi-sensor fusion as described in claim 4, characterized in that, The process of constructing a spatiotemporal graph structure for human motion based on the fusion of the thermodynamic spatiotemporal gradient features and the kinematic spatiotemporal gradient features includes the following sub-steps: The thermodynamic spatiotemporal gradient features are mapped to a graph node attribute set, wherein each node in the graph node attribute set corresponds to a specific human joint. The kinematic spatiotemporal gradient features are converted into a graph edge weight set, where each edge of the graph edge weight set represents the strength of the dynamic association between joints; Construct an initial topological skeleton based on the standard human skeletal connections; The graph node attribute set is injected into the nodes of the initial topological graph skeleton, and the graph edge weight set is injected into the corresponding edges of the initial topological graph skeleton to generate a spatiotemporal graph structure of human motion.

6. The human fall detection method based on multi-sensor fusion as described in claim 5, characterized in that, The step of inputting the spatiotemporal graph structure into a spatiotemporal graph convolutional network, performing human fall feature evolution through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network, and outputting the fall risk probability includes the following sub-steps: The spatial graph convolutional layer of the spatiotemporal graph convolutional network is used to aggregate neighborhood features of the spatiotemporal graph structure to generate a spatially enhanced node feature matrix. The node feature matrix is ​​modeled in a multi-scale temporal manner by the dilated temporal convolutional layer of the spatiotemporal graph convolutional network to extract the spatiotemporal joint feature vector. The spatiotemporal joint feature vector is dynamically reweighted using a learnable attention mechanism to generate a fall feature vector that focuses on key joints. The fully connected neural network layer of the spatiotemporal graph convolutional network performs nonlinear mapping on the fall feature vector and outputs the original fall risk probability value. The original fall risk probability values ​​within a continuous time window are processed by median filtering to generate the final fall risk probability.

7. A human fall detection system based on multi-sensor fusion, characterized in that, include: The acquisition module allows users to detect the temperature distribution of the surrounding human body in real time through a thermal imaging sensor, generate a thermal imaging image of the human body, and simultaneously transmit high-frequency electromagnetic waves to the surrounding human body through a millimeter-wave radar and receive reflected signals to generate three-dimensional point cloud data of the human body. The feature extraction module is used to extract the human body contour from the thermal imaging image, generate the thermodynamic feature tensor of the human body, and reconstruct the human body motion trajectory from the three-dimensional point cloud data, generating the three-dimensional motion feature tensor of the human body. The feature calculation module is used to perform thermodynamic spatiotemporal gradient calculation on the thermodynamic feature tensor to generate thermodynamic spatiotemporal gradient features that reflect the rate of change of body temperature distribution of the human body, and to perform kinematic spatiotemporal gradient calculation on the three-dimensional motion feature tensor to generate kinematic spatiotemporal gradient features that reflect the intensity of abrupt changes in the posture of the human body. A construction module is used to construct a spatiotemporal graph structure of human motion based on the fusion of the thermodynamic spatiotemporal gradient features and the kinematic spatiotemporal gradient features; The prediction module is used to input the spatiotemporal graph structure into the spatiotemporal graph convolutional network, perform human fall feature evolution through the multi-level feature interaction mechanism of the spatiotemporal graph convolutional network, and output the fall risk probability. The reminder module is used to perform a human fall reminder operation when the probability of falling exceeds a preset threshold.

8. The human fall detection system based on multi-sensor fusion as described in claim 7, characterized in that, The acquisition module is specifically used for: Adaptive noise filtering is applied to the raw temperature data acquired by the thermal imaging sensor to generate a noise-reduced temperature distribution matrix. A multi-scale segmentation algorithm is used to identify human target regions in the denoised temperature distribution matrix, generating an accurate binary mask that separates the human body from the background. Based on the binary mask, spatial domain filtering is performed on the original temperature data to extract the temperature distribution data of the target area of ​​the human body and generate a thermal imaging image of the human body. The direction of arrival (DOA) of the reflected signal received by the millimeter-wave radar is estimated and multipath interference is eliminated to generate an initial set of spatial point clouds. The initial set of spatial point clouds is dynamically separated using a density clustering algorithm to extract three-dimensional point cloud data corresponding to human targets.

9. The human fall detection system based on multi-sensor fusion as described in claim 8, characterized in that, The feature extraction module is specifically used for: The human contour extraction algorithm based on thermal gradient is used to extract the human contour from the thermal imaging image and generate a human contour vector image containing joint nodes. The thermal distribution of key parts of the human body contour vector image is modeled to generate a joint temperature distribution matrix. The joint temperature distribution matrix is ​​tensorized and recombined according to the time series to generate the thermodynamic feature tensor of the human body. Feature point matching and motion compensation are performed on the three-dimensional point cloud data to generate a discrete path point set of human motion trajectory; The discrete path point set is smoothly interpolated using a non-uniform B-spline curve fitting algorithm to generate a continuous three-dimensional motion feature tensor of the human body.

10. The human fall detection system based on multi-sensor fusion as described in claim 9, characterized in that, The feature calculation module is specifically used for: Perform a first-order difference operation on the thermodynamic feature tensor in the time dimension to generate the temperature change rate tensor at each key point. The temperature change rate tensor at each joint point is spatially gradient convolved using the three-dimensional Sobel operator to extract the thermodynamic spatiotemporal gradient features that reflect the rate of change of the body temperature distribution of the human body. The three-dimensional motion feature tensor is decomposed and synthesized into velocity vectors to generate temporal data of motion components for each limb segment; The sliding window standard deviation detection algorithm is used to perform acceleration mutation analysis on the time series data of the motion components, and generate kinematic spatiotemporal gradient features that reflect the intensity of the posture mutation of the human body.