A millimeter wave radar fall detection method and system based on a fusion neural network

CN122780656APending Publication Date: 2026-09-18WEIFU INTELLIGENT SENSE (WUXI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610806514.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

此类方法数据凝练、抗干扰能力较强,但现有模型多采用单一网络结构或简单的串联架构,难以充分兼顾局部运动细节与长程动态规律的协同表达

Benefits of technology

(1)数据处理定制化:本发明采用垂直降权DBSCAN进行点云聚类,以跌倒检测的任务精度为目标,经消融实验确定垂直方向权重,用于解决跌倒与下蹲/弯腰在垂直维度上的特征混淆问题,优化目标、权重确定依据、技术效果预期深度适配跌倒检测任务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780656A_ABST
    Figure CN122780656A_ABST
Patent Text Reader

Abstract

This invention relates to a fall detection method and system based on a fusion neural network using millimeter-wave radar. The invention includes acquiring raw point cloud data detected by millimeter-wave radar; retaining moving points that meet motion discrimination conditions; clustering and tracking the moving points; outputting target attribute data based on the clustering and tracking; obtaining trajectory data; constructing a fixed-length time-series sample; dividing it into a training set, a validation set, and a test set; constructing a dual-branch parallel fusion neural network; training the fusion neural network using the training set; selecting and saving the optimal model using the validation set; evaluating the generalization performance of the optimal model using the test set; and performing real-time fall detection using the optimal model after generalization performance. This invention enables the detection and recognition of human falls by constructing a parallel dual-branch fusion neural network using convolutional neural networks and gated recurrent units, combined with millimeter-wave radar target trajectories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of millimeter-wave radar target detection technology, and in particular to a millimeter-wave radar fall detection method and system based on a fused neural network. Background Technology

[0002] With the increasing number of elderly people living alone, the risks of falls and sudden illnesses in the home environment are rising, highlighting the growing importance of home health monitoring systems. Similarly, in industrial settings (such as factory workshops, substations, and construction sites), falls due to fatigue, complex environments, or improper machinery operation are frequent, often accompanied by more serious secondary hazards such as mechanical injuries, electric shocks, or chemical leaks. Therefore, fall detection technology that combines privacy protection, environmental adaptability, and real-time response capabilities has broad application prospects in both home-based elderly care and industrial safety.

[0003] Existing fall detection methods based on millimeter-wave radar can be broadly classified into three categories: (1) Methods based on micro-Doppler spectra classify radar echoes generated by human movements by analyzing their time-frequency characteristics. Commonly used algorithms include short-time Fourier transform combined with convolutional neural networks, Stockwell transform combined with convolutional neural networks, and Doppler spectra combined with long short-term memory networks. These methods are sensitive to the overall Doppler changes of movements, but they are difficult to capture fine spatial structure information, and the features are not obvious when the direction of human movement is perpendicular to the radar radial direction.

[0004] (2) Point cloud data-based methods utilize three-dimensional spatial point clouds generated by radar virtual arrays to represent human posture. Commonly used models include PointNet and PointNet++, PointLSTM, and combinations of PointNet++ and Long Short-Term Memory networks. These methods provide rich spatial information, but suffer from severe problems of sparse point clouds or even missing points under multi-channel radar configurations, resulting in insufficient detection robustness.

[0005] (3) Target trajectory-based methods directly utilize the temporal information such as position, velocity, and acceleration output by radar tracking. Commonly used models include two-layer long short-term memory networks, hybrid architectures of convolutional neural networks and long short-term memory networks, and anomaly detection frameworks based on variational autoencoders. These methods have strong data condensation and anti-interference capabilities, but existing models mostly adopt a single network structure or a simple serial architecture, making it difficult to fully consider the coordinated expression of local motion details and long-range dynamic laws. Summary of the Invention

[0006] Therefore, this invention provides a fall detection method and system based on a fusion neural network for millimeter-wave radar. It can build a parallel dual-branch fusion neural network by using a convolutional neural network (CNN) and a gated recurrent unit (GRU), and combine it with the target trajectory of millimeter-wave radar to realize the detection and recognition of human falls.

[0007] To address the aforementioned technical problems, this invention provides a millimeter-wave radar fall detection method based on a fused neural network, comprising: Acquire raw point cloud data detected by millimeter-wave radar, wherein the raw point cloud data includes at least one of the following: distance, radial velocity, horizontal angle, elevation angle, radar cross section, and echo power of the point target; The original point cloud data is subjected to motion discrimination based on the absolute value of the radial velocity of the point target, and the moving points that meet the motion discrimination conditions are retained. Clustering and tracking of the moving point includes: introducing an anisotropic weighting strategy in the distance metric of density clustering, assigning a weight factor to the vertical direction that is lower than that to the horizontal direction, so that the clustering decision focuses more on the spatial continuity of the target in the horizontal plane than the vertical direction, and predicting the state vector of the existing track in the current frame based on the Kalman filter. Target attribute data is output based on clustering tracking; The target trajectory formed by the target attribute data is preprocessed to obtain trajectory data; Construct a fixed-length time series sample based on the trajectory data; The time-series samples are labeled with categories and divided into training set, validation set and test set; A dual-branch parallel fusion neural network is constructed, comprising a one-dimensional convolutional branch, a gated recurrent unit branch, and a feature fusion and decision layer. The one-dimensional convolutional branch extracts the temporal samples and outputs local transient change features in the fall action. The gated recurrent unit branch extracts and outputs a global feature vector showing the transition of the human body from a normal posture to a fall state. The feature fusion and decision layer concatenates the local transient change features with the global feature vector by adjusting the feature dimension to obtain a fused feature vector. After Softmax normalization of the fused feature vector, the classification probabilities of the fall category and the non-fall category are output. The fusion neural network is trained using the training set; The validation set is used to select a model and save the optimal model. The generalization performance of the optimal model is evaluated using the test set. Real-time fall detection is performed using the optimal model after generalization performance.

[0008] In one embodiment of the present invention, the anisotropic weighting strategy employs the following weighted distance metric: ; in, This represents the weighted distance between the i-th detection point and the j-th detection point. , , Let x, y, x, and y represent the horizontal, vertical, and y coordinates of the i-th detection point in the three-dimensional coordinate system, respectively. , , Let x represent the horizontal, vertical, and y coordinates of the j-th detection point in the three-dimensional coordinate system, respectively, and let α represent the vertical weight coefficient, where 0 < α < 1.

[0009] In one embodiment of the present invention, the current frame prediction of the state vector of an existing track based on a Kalman filter includes: The track state vector of the Kalman filter is: ; Where x, y, and z represent the position components of the target in the three-dimensional coordinate system. , , These represent the velocity components of the target in the three-dimensional coordinate system. , , Let T represent the acceleration components of the target in the three-dimensional coordinate system, and T denote the matrix transpose. The Kalman filter performs state prediction and covariance prediction on an existing track. The expression is: ; ; in, This represents the predicted state vector for frame t. Let represent the posterior state vector of the (t-1)th frame. Let represent the prediction covariance matrix of frame t. Let represent the posterior covariance matrix of the (t-1)th frame. Represents the state transition matrix. This represents the transpose of the state transition matrix F. Represent the covariance matrix; Calculate the combined cost between the centroid of the i-th point cluster and the predicted value of the j-th track, where the combined cost is: ; in, This represents the combined cost between the centroid of the i-th point cluster and the predicted value of the j-th track. This represents the Euclidean distance between the centroid of the i-th point cluster and the predicted position of the j-th track. This represents the absolute value of the radial velocity difference between the i-th point cluster and the j-th track prediction. This represents the absolute value of the acceleration difference between the i-th point cluster and the j-th track prediction. , , These represent the weighting coefficients for the Euclidean distance term, the radial velocity difference term, and the acceleration difference term, respectively. The target attribute data output by clustering tracking includes three-dimensional position information, radial velocity, radar cross section (RCS), echo power, and cluster size information.

[0010] In one embodiment of the present invention, data preprocessing is performed on the target trajectory formed by the target attribute data, including: The data preprocessing includes data cleaning and normalization; The data cleaning includes removing invalid frames, removing data that is outside the radar detection range, and removing outliers caused by multipath, mismerging, misassociation, track start data, or track disappearance data. The normalization uses max-min normalization to map each feature to the interval [-1, 1].

[0011] In one embodiment of the present invention, constructing a fixed-length time-series sample based on the trajectory data includes: A fixed-length time-series sample is constructed from the trajectory data using a sliding window approach; The category labeling uses visual data collected synchronously with millimeter-wave radar as a reference to label each sequence sample with a category label, including falling, normal walking, sitting, bending over, and abnormal postures in industrial scenarios.

[0012] In one embodiment of the present invention, the one-dimensional convolutional branch includes a first-level one-dimensional convolutional layer, a first ReLU nonlinear mapping layer, a max pooling layer, a second-level one-dimensional convolutional layer, a second ReLU nonlinear mapping layer, and a global mean pooling layer; wherein, the first-level one-dimensional convolutional layer has 32 output channels, a kernel size of 2, a stride of 1, and a padding method of "same"; the max pooling layer has a stride of 2; the second-level one-dimensional convolutional layer has 64 output channels, a kernel size of 2, a stride of 1; and the global mean pooling layer outputs a 64-dimensional local feature vector. The gated recurrent unit branch includes a cascaded first-level recurrent unit (GRU) and second-level recurrent unit (GRU); the first-level recurrent unit (GRU) has an input dimension of F, 64 hidden units, and returns a complete sequence; the second-level recurrent unit (GRU) has 64 hidden units and only returns the final state; the gated recurrent unit branch outputs a 64-dimensional global feature vector. The feature fusion and decision layer includes a feature dimension concatenation layer, a fully connected mapping layer, and a Softmax classification layer, which respectively satisfy: ; ; ; in, This represents the local feature vector output by a one-dimensional convolution branch. This represents the global feature vector output by the branch of the gated loop unit. denoted as fused feature vector, [ ] indicates concatenation according to feature dimension, W represents the weight matrix of the fully connected layer, b represents the bias vector of the fully connected layer, h represents the binary logit vector output by the fully connected layer, and ŷ represents the probability vector output by the Softmax classification layer.

[0013] In one embodiment of the present invention, training the fusion neural network using the training set includes: Cross-entropy is used as the optimization objective to measure the difference between the predicted distribution and the true label. The cross-entropy loss function is expressed as: ; in, This represents the weighted cross-entropy loss of the current batch of samples. Indicates the number of samples in the batch. Indicates the sample number. Indicates the first The true label of each sample It means to fall down. Indicates not a fall. The model predicts the first... The probability that a sample belongs to the "falling" category.

[0014] In one embodiment of the present invention, training the fusion neural network using the training set further includes: The optimizer uses adaptive moment estimation (Adam), with an initial learning rate of 0.001, a first-order moment decay coefficient of 0.9, a second-order moment decay coefficient of 0.999, and a numerical stability term ε= The batch size for each parameter iteration is set to 16, and the maximum number of iterations is set to 200. The network weights are updated iteratively through backpropagation.

[0015] In one embodiment of the present invention, model selection and saving of the optimal model using the validation set includes: After each training round, the current model's classification accuracy is verified, and the model is switched to inference mode to avoid gradient calculation and parameter updates. The validation set is input into the model for forward inference to obtain the predicted probability output; The accuracy rate is calculated to verify the model's precision, using the following formula: ; Where Acc represents the accuracy rate; TP represents the number of fall samples that were correctly detected as falls; TN represents the number of non-fall samples that were correctly identified as non-fall samples; FP represents the number of non-fall samples that were falsely detected as falls; and FN represents the number of fall samples that were missed as non-fall samples. The model saving strategy is as follows: if the current validation accuracy (Acc) exceeds the historical best accuracy... Then save the current network weights. and will Update to the current validation accuracy Acc; if the accuracy does not improve for 10 consecutive rounds, terminate training early to prevent overfitting. The evaluation of the generalization performance of the optimal model using the test set includes: Calculate accuracy (Acc), recall, precision, and F1 score. Acc is used to validate model accuracy, recall measures the ability to detect fall events, precision measures the confidence of alarms, and the F1 score is used to calculate the harmonic mean of precision and recall to comprehensively measure both metrics. The calculations are as follows: ; ; ; ; Where Acc represents accuracy; TP represents the number of fall samples correctly detected as fall samples; TN represents the number of non-fall samples correctly identified as non-fall samples; FP represents the number of non-fall samples falsely detected as fall samples; FN represents the number of fall samples missed as non-fall samples; Recall represents recall; Precision represents precision; and F1 represents the F1 score.

[0016] The present invention also provides a millimeter-wave radar fall detection system based on a fused neural network, comprising: The raw point cloud data acquisition module is used to acquire raw point cloud data detected by millimeter-wave radar. The raw point cloud data includes at least one of the following: distance, radial velocity, horizontal angle, elevation angle, radar cross section, and echo power of the point target. The moving point acquisition module is used to perform motion and static discrimination on the original point cloud data based on the absolute value of the radial velocity of the point target, and retain the moving points that meet the motion discrimination conditions; The clustering tracking module is used to perform clustering tracking on the moving point, including: introducing an anisotropic weighting strategy in the distance metric of density clustering, assigning a weight factor to the vertical direction that is lower than that to the horizontal direction, so that the clustering decision focuses more on the spatial continuity of the target in the horizontal plane relative to the vertical direction, and predicting the state vector of the existing track for the current frame based on the Kalman filter. The output module is used to output target attribute data based on clustering tracking. The trajectory data acquisition module is used to preprocess the target trajectory formed by the target attribute data to obtain trajectory data. A time-series sample construction module is used to construct time-series samples of a fixed length based on the trajectory data; The data partitioning module is used to classify the time-series samples and divide them into training set, validation set and test set; A fusion neural network construction module is used to construct a dual-branch parallel fusion neural network, including a one-dimensional convolutional branch, a gated recurrent unit branch, and a feature fusion and decision layer. The one-dimensional convolutional branch extracts the temporal samples and outputs local transient change features in the fall action, while the gated recurrent unit branch extracts and outputs a global feature vector showing the transition of the human body from a normal posture to a fall state. The feature fusion and decision layer concatenates the local transient change features and the global feature vector by their feature dimensions to obtain a fused feature vector. After performing Softmax normalization on the fused feature vector, it outputs the classification probabilities of the fall category and the non-fall category. A training module is used to train the fusion neural network using the training set; The optimal model acquisition module is used to select a model using the validation set and save the optimal model. A generalization performance evaluation module is used to evaluate the generalization performance of the optimal model using the test set. A detection module is used to perform real-time fall detection using the optimal model after generalization performance.

[0017] The technical solution of the present invention has the following advantages compared with the prior art: (1) Customized data processing: This invention uses vertical weighted DBSCAN for point cloud clustering, with the goal of improving the accuracy of fall detection tasks. The vertical weights are determined through ablation experiments to solve the problem of feature confusion between falls and squatting / bending in the vertical dimension. The optimization target, weight determination basis, and expected technical effect are deeply adapted to the fall detection task.

[0018] (2) Target trajectory collaborative dual-branch network: Existing target trajectory-based methods mostly adopt a single network or a simple serial architecture; dual-branch fusion-based methods mostly process image features (such as distance-Doppler spectra) rather than trajectory time-series data. This invention designs the trajectory input in collaboration with a CNN+GRU dual-branch parallel network: the trajectory data is simultaneously fed into the CNN branch to extract local transient features such as velocity mutations, and into the GRU branch to extract long-range dependent features such as centroid shifts. After feature-level fusion, the classification result is output, realizing the adaptation of input data and network structure.

[0019] (3) Lightweight real-time deployment: The overall parameters of the dual-branch structure are controllable, the computational overhead is low, and it can be embedded in a small-channel radar to complete inference. It ensures millisecond-level response capability and meets the productization requirements of low power consumption and low cost in home and industrial scenarios.

[0020] (4) Cross-scenario applicability: In addition to home-based elderly care, this method is also applicable to industrial environments such as factory workshops, substations, and construction sites. Millimeter-wave radar is not affected by dust, smoke, or changes in lighting, and can effectively monitor accidental falls of workers, reducing the risk of secondary injuries. Attached Figure Description

[0021] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0022] Figure 1 This is a flowchart of the millimeter-wave radar fall detection method based on a fused neural network, as described in this invention. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0024] Example 1 Reference Figure 1 As shown, a millimeter-wave radar fall detection method based on a fusion neural network includes: S1. Acquire raw point cloud data detected by millimeter-wave radar. The raw point cloud data includes at least one of the following: distance, radial velocity, horizontal angle, elevation angle, radar cross section (RCS), and echo power of the point target.

[0025] S2. Based on the absolute value of the radial velocity of the point target, the original point cloud data is subjected to motion discrimination, and moving points that meet the motion discrimination conditions are retained. In this embodiment, the threshold is set to 0.1 m / s. Points with velocities higher than the threshold are identified as moving points and retained for subsequent steps.

[0026] S3. Cluster the moving points for tracking, and adapt to fall scenarios to address the physical characteristic that millimeter-wave radar's vertical resolution is weaker than its horizontal resolution. This includes: introducing an anisotropic weighting strategy into the distance metric of density clustering, assigning a weight factor to the vertical direction that is lower than that to the horizontal direction, making the clustering decision focus more on the spatial continuity of the target in the horizontal plane than on the vertical direction, and predicting the current frame based on the state vector of the existing track using a Kalman filter.

[0027] Specifically, the anisotropic weighting strategy employs the following weighted distance metric: ; in, This represents the weighted distance between the i-th detection point and the j-th detection point. , , Let x, y, x, and y represent the horizontal, vertical, and y coordinates of the i-th detection point in the three-dimensional coordinate system, respectively. , , Let x, y, x, and y represent the horizontal, vertical, and y coordinates of the j-th detection point in the three-dimensional coordinate system, respectively. Let α represent the vertical weighting coefficient, and 0 < α < 1. The smaller α is, the lower the contribution of the vertical distance difference to the clustering distance. Preferably, α is determined to be 0.4 through ablation experiments aimed at improving the accuracy of fall detection tasks, the clustering neighborhood radius is set to 0.4 m, and the minimum number of points is set to 2.

[0028] Specifically, the current frame prediction is performed based on the state vector of the existing track using a Kalman filter, including: The track state vector of the Kalman filter is: ; Where x, y, and z represent the position components of the target in the three-dimensional coordinate system. , , These represent the velocity components of the target in the three-dimensional coordinate system. , , Let T represent the acceleration components of the target in the three-dimensional coordinate system, and T denote the matrix transpose. The Kalman filter performs state prediction and covariance prediction on an existing track. The expression is: ; ; in, This represents the predicted state vector for frame t. Let represent the posterior state vector of the (t-1)th frame. Let represent the prediction covariance matrix of frame t. Let represent the posterior covariance matrix of the (t-1)th frame. Represents the state transition matrix. This represents the transpose of the state transition matrix F. Represent the covariance matrix; The comprehensive cost between the centroid of the i-th point cluster and the predicted value of the j-th track is calculated, consisting of a weighted sum of three indicators. This comprehensive cost is: ; in, This represents the combined cost between the centroid of the i-th point cluster and the predicted value of the j-th track. This represents the Euclidean distance between the centroid of the i-th point cluster and the predicted position of the j-th track. This represents the absolute value of the radial velocity difference between the i-th point cluster and the j-th track prediction. This represents the absolute value of the acceleration difference between the i-th point cluster and the j-th track prediction. , , These represent the weighting coefficients for the Euclidean distance term, the radial velocity difference term, and the acceleration difference term, respectively.

[0029] S4. Output target attribute data based on clustering; including three-dimensional position information, radial velocity, radar cross section (RCS), echo power, and cluster size information (cluster height and cluster width, etc.).

[0030] S5. Perform data preprocessing on the target trajectory formed by the target attribute data to obtain trajectory data. Specifically, this includes: The data preprocessing includes data cleaning and normalization; The data cleaning includes removing invalid frames, removing data that is outside the radar detection range, and removing outliers caused by multipath, mismerging, misassociation, track start data, or track disappearance data. The normalization uses max-min normalization to map each feature to the interval [-1, 1].

[0031] S6. Construct a fixed-length time-series sample based on the trajectory data. Specifically, this includes: A fixed-length time-series sample is constructed from the trajectory data using a sliding window approach; The preferred window length is 15 frames, the window step size is 5 frames, and there is partial overlap between adjacent samples to enhance data diversity.

[0032] The category labeling uses visual data collected synchronously with millimeter-wave radar as a reference to label each sequence sample with a category label, including falling, normal walking, sitting, bending over, and abnormal postures in industrial scenarios.

[0033] The time-series samples were labeled with categories and divided into training, validation, and test sets. During dataset partitioning, all labeled samples were randomly divided into training, validation, and test sets in a ratio of 7:1.5:1.5, ensuring that the three sets of data do not overlap at the sample level and guaranteeing the objectivity of model evaluation.

[0034] S7, Training Set. Used for learning and optimizing model parameters, updating network weights through backpropagation.

[0035] S8, Validation set. Used for hyperparameter tuning and model selection during training, monitoring overfitting, and saving the optimal model weights.

[0036] S9, Test set. Used for the final evaluation of the model's generalization performance. The validation set and test set are not involved in parameter updates throughout the training process.

[0037] S10. Construct a dual-branch parallel fusion neural network, including a one-dimensional convolutional branch, a gated recurrent unit branch, and a feature fusion and decision layer; wherein, the one-dimensional convolutional branch extracts the temporal samples and outputs the local transient change features in the fall action, and the gated recurrent unit branch extracts and outputs the global feature vector of the human body transitioning from a normal posture to a fall state; the feature fusion and decision layer is used to concatenate the local transient change features and the global feature vector by feature dimension to obtain a fused feature vector, and after Softmax normalization of the fused feature vector, output the classification probability of the fall category and the non-fall category.

[0038] Specifically, the one-dimensional convolutional branch includes a first-level one-dimensional convolutional layer, a first ReLU nonlinear mapping layer, a max pooling layer, a second-level one-dimensional convolutional layer, a second ReLU nonlinear mapping layer, and a global mean pooling layer; wherein, the first-level one-dimensional convolutional layer has 32 output channels, a kernel size of 2, a stride of 1, and a padding method of "same"; the max pooling layer has a stride of 2; the second-level one-dimensional convolutional layer has 64 output channels, a kernel size of 2, and a stride of 1; and the global mean pooling layer outputs a 64-dimensional local feature vector.

[0039] Specifically, the gated recurrent unit branch includes a cascaded first-level recurrent unit (GRU) and a second-level recurrent unit (GRU); the first-level recurrent unit (GRU) has an input dimension of F, 64 hidden units, and returns a complete sequence; the second-level recurrent unit (GRU) has 64 hidden units and only returns the final state; the gated recurrent unit branch outputs a 64-dimensional global feature vector.

[0040] Specifically, the feature fusion and decision layer includes a feature dimension concatenation layer, a fully connected mapping layer, and a Softmax classification layer, which respectively satisfy: ; ; ; in, This represents the local feature vector output by a one-dimensional convolution branch. This represents the global feature vector output by the branch of the gated loop unit. denoted as fused feature vector, [ ] indicates concatenation according to feature dimension, W represents the weight matrix of the fully connected layer, b represents the bias vector of the fully connected layer, h represents the binary logit vector output by the fully connected layer, and ŷ represents the probability vector output by the Softmax classification layer.

[0041] In some embodiments, in step S10, a dual-branch parallel temporal feature decoupling network is constructed. The two branches take the same target trajectory sequence as input, forming a functional division of labor and complementarity. The input is represented as: ; Where X represents the input target trajectory time series sample; R represents the real number field; T represents the window length, preferably T=15; and F represents the target trajectory feature dimension. Preferably, F includes position x, y, z; radial velocity; RCS; power; point cluster height; and point cluster width, in which case F is 6.

[0042] The transient change sensing branch is a one-dimensional convolutional branch used to capture short-term, dramatic fluctuations in a fall, such as a sudden increase in radial velocity, abrupt changes in RCS, power changes, or abrupt changes in target size. This one-dimensional convolutional branch includes a first-level transform, a second-level transform, and temporal compression. The first-level transform uses Conv1D with F input channels, 32 output channels, a kernel size of 2, a stride of 1, and same padding, followed by ReLU nonlinear mapping and max pooling with a stride of 2. The second-level transform also uses Conv1D with 32 input channels, 64 output channels, a kernel size of 2, a stride of 1, followed by ReLU nonlinear mapping. Temporal compression uses global mean pooling, outputting a 64-dimensional local feature vector. .

[0043] The state evolution modeling branch is a gated recurrent unit (GRU) branch, used to learn the global evolutionary patterns of the human body transitioning from a normal posture to a falling state, such as center of gravity shift, velocity change trends, and height change trends. This GRU branch consists of two levels of GRUs. The first-level GRU has an input dimension of F and 64 hidden units, and returns the complete sequence; the second-level GRU has an input dimension of 64 and 64 hidden units, and only returns the final state, outputting a 64-dimensional global feature vector. .

[0044] The probability of class k in the Softmax output can be expressed as: ; Where h0 represents the logit corresponding to the non-fall category, h1 represents the logit corresponding to the fall category, ŷ0 represents the probability of the non-fall category, and ŷ1 represents the probability of the fall category.

[0045] S11. Train the fusion neural network using the training set.

[0046] Specifically, training the fusion neural network using the training set includes: Cross-entropy is used as the optimization objective to measure the difference between the predicted distribution and the true label. The cross-entropy loss function is expressed as: ; in, This represents the weighted cross-entropy loss of the current batch of samples. Indicates the number of samples in the batch. Indicates the sample number. Indicates the first The true label of each sample It means to fall down. Indicates not a fall. The model predicts the first... The probability that a sample belongs to the "falling" category.

[0047] Specifically, training the fusion neural network using the training set further includes: The optimizer uses adaptive moment estimation (Adam), with an initial learning rate of 0.001, a first-order moment decay coefficient of 0.9, a second-order moment decay coefficient of 0.999, and a numerical stability term ε= The batch size for each parameter iteration is set to 16, and the maximum number of iterations is set to 200. The network weights are updated iteratively through backpropagation.

[0048] S12. Use the validation set to select a model and save the optimal model. Specifically, this includes: After each training round, the current model's classification accuracy is verified, and the model is switched to inference mode to avoid gradient calculation and parameter updates. The validation set is input into the model for forward inference to obtain the predicted probability output; The accuracy rate is calculated to verify the model's precision, using the following formula: ; Where Acc represents the accuracy rate; TP represents the number of fall samples that were correctly detected as fall samples; TN represents the number of non-fall samples that were correctly identified as non-fall samples; FP represents the number of non-fall samples that were falsely detected as fall samples; and FN represents the number of fall samples that were missed as non-fall samples.

[0049] The model saving strategy is as follows: if the current validation accuracy (Acc) exceeds the historical best accuracy... Then save the current network weights. and will Update to the current validation accuracy Acc; if the accuracy does not improve for 10 consecutive rounds, terminate training early to prevent overfitting.

[0050] S13. Evaluate the generalization performance of the optimal model using the test set. Specifically, this includes: Calculate accuracy (Acc), recall, precision, and F1 score. Acc is used to validate model accuracy, recall measures the ability to detect fall events, precision measures the confidence of alarms, and the F1 score is used to calculate the harmonic mean of precision and recall to comprehensively measure both metrics. The calculations are as follows: ; ; ; ; Where Acc represents accuracy; TP represents the number of fall samples correctly detected as fall samples; TN represents the number of non-fall samples correctly identified as non-fall samples; FP represents the number of non-fall samples falsely detected as fall samples; FN represents the number of fall samples missed as non-fall samples; Recall represents recall; Precision represents precision; and F1 represents the F1 score.

[0051] S14. Model Deployment. The network structure is converted into C language code, and the saved weight parameters are extracted and stored. Integer quantization is used when necessary to reduce the storage footprint of the embedded device. The deployed model runs directly on the embedded processor inside the millimeter-wave radar, without the need for external computing devices, achieving low-latency response for real-time fall detection.

[0052] S15, End.

[0053] In summary, this method uses a dual-branch fusion neural network for fall detection and recognition in millimeter-wave radar. Compared to methods based on micro-Doppler spectra, it does not rely on time-frequency transformation, avoiding feature attenuation when the direction of motion is perpendicular to the radar radial direction, and can simultaneously capture spatial location and temporal dynamic information. Compared to methods based on point cloud data, it directly uses the trajectory sequence output by the radar, without relying on high-density point clouds, fundamentally avoiding the performance degradation caused by sparse or missing point clouds in multi-channel radar configurations. Compared to existing target trajectory-based methods, it adopts a dual-branch parallel fusion architecture, where a one-dimensional convolutional branch and a gated recurrent unit branch extract local details and long-range dependencies respectively, and output them after feature-level fusion, overcoming the limitation that a single network or simple serial architecture cannot take into account both types of features.

[0054] Example 2 Based on the same inventive concept, this embodiment provides a millimeter-wave radar fall detection system based on a fused neural network. The principle of solving the problem is similar to that of the millimeter-wave radar fall detection method based on a fused neural network, and the repeated parts will not be described again.

[0055] This embodiment provides a millimeter-wave radar fall detection system based on a fused neural network, including: The raw point cloud data acquisition module is used to acquire raw point cloud data detected by millimeter-wave radar. The raw point cloud data includes at least one of the following: distance, radial velocity, horizontal angle, elevation angle, radar cross section, and echo power of the point target. The moving point acquisition module is used to perform motion and static discrimination on the original point cloud data based on the absolute value of the radial velocity of the point target, and retain the moving points that meet the motion discrimination conditions; The clustering tracking module is used to perform clustering tracking on the moving point, including: introducing an anisotropic weighting strategy in the distance metric of density clustering, assigning a weight factor to the vertical direction that is lower than that to the horizontal direction, so that the clustering decision focuses more on the spatial continuity of the target in the horizontal plane relative to the vertical direction, and predicting the state vector of the existing track for the current frame based on the Kalman filter. The output module is used to output target attribute data based on clustering tracking. The trajectory data acquisition module is used to preprocess the target trajectory formed by the target attribute data to obtain trajectory data. A time-series sample construction module is used to construct time-series samples of a fixed length based on the trajectory data; The data partitioning module is used to classify the time-series samples and divide them into training set, validation set and test set; A fusion neural network construction module is used to construct a dual-branch parallel fusion neural network, including a one-dimensional convolutional branch, a gated recurrent unit branch, and a feature fusion and decision layer. The one-dimensional convolutional branch extracts the temporal samples and outputs local transient change features in the fall action, while the gated recurrent unit branch extracts and outputs a global feature vector showing the transition of the human body from a normal posture to a fall state. The feature fusion and decision layer concatenates the local transient change features and the global feature vector by their feature dimensions to obtain a fused feature vector. After performing Softmax normalization on the fused feature vector, it outputs the classification probabilities of the fall category and the non-fall category. A training module is used to train the fusion neural network using the training set; The optimal model acquisition module is used to select a model using the validation set and save the optimal model. A generalization performance evaluation module is used to evaluate the generalization performance of the optimal model using the test set. A detection module is used to perform real-time fall detection using the optimal model after generalization performance.

[0056] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0057] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0058] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0059] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0060] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A fall detection method based on millimeter-wave radar using a fusion neural network, characterized in that, include: Acquire raw point cloud data detected by millimeter-wave radar, wherein the raw point cloud data includes at least one of the following: distance, radial velocity, horizontal angle, elevation angle, radar cross section, and echo power of the point target; The original point cloud data is subjected to motion discrimination based on the absolute value of the radial velocity of the point target, and the moving points that meet the motion discrimination conditions are retained. Clustering and tracking of the moving point includes: introducing an anisotropic weighting strategy in the distance metric of density clustering, assigning a weight factor to the vertical direction that is lower than that to the horizontal direction, so that the clustering decision focuses more on the spatial continuity of the target in the horizontal plane than the vertical direction, and predicting the state vector of the existing track in the current frame based on the Kalman filter. Target attribute data is output based on clustering tracking; The target trajectory formed by the target attribute data is preprocessed to obtain trajectory data; Construct a fixed-length time series sample based on the trajectory data; The time-series samples are labeled with categories and divided into training set, validation set and test set; A dual-branch parallel fusion neural network is constructed, comprising a one-dimensional convolutional branch, a gated recurrent unit branch, and a feature fusion and decision layer. The one-dimensional convolutional branch extracts the temporal samples and outputs local transient change features in the fall action. The gated recurrent unit branch extracts and outputs a global feature vector showing the transition of the human body from a normal posture to a fall state. The feature fusion and decision layer concatenates the local transient change features with the global feature vector by adjusting the feature dimension to obtain a fused feature vector. After Softmax normalization of the fused feature vector, the classification probabilities of the fall category and the non-fall category are output. The fusion neural network is trained using the training set; The validation set is used to select a model and save the optimal model. The generalization performance of the optimal model is evaluated using the test set. Real-time fall detection is performed using the optimal model after generalization performance.

2. The fall detection method based on a fusion neural network using millimeter-wave radar according to claim 1, characterized in that, The anisotropic weighting strategy uses the following weighted distance metric: ; in, This represents the weighted distance between the i-th detection point and the j-th detection point. , , Let x, y, x, and y represent the horizontal, vertical, and y coordinates of the i-th detection point in the three-dimensional coordinate system, respectively. , , Let x represent the horizontal, vertical, and y coordinates of the j-th detection point in the three-dimensional coordinate system, respectively, and let α represent the vertical weight coefficient, where 0 < α < 1.

3. The fall detection method based on a fused neural network using millimeter-wave radar according to claim 1, characterized in that, Predicting the current frame based on the state vector of an existing track using a Kalman filter includes: The track state vector of the Kalman filter is: ; Where x, y, and z represent the position components of the target in the three-dimensional coordinate system. , , These represent the velocity components of the target in the three-dimensional coordinate system. , , Let T represent the acceleration components of the target in the three-dimensional coordinate system, and T denote the matrix transpose. The Kalman filter performs state prediction and covariance prediction on an existing track. The expression is: ; ; in, This represents the predicted state vector for frame t. Let represent the posterior state vector of the (t-1)th frame. Let represent the prediction covariance matrix of frame t. Let represent the posterior covariance matrix of the (t-1)th frame. Represents the state transition matrix. This represents the transpose of the state transition matrix F. Represent the covariance matrix; Calculate the combined cost between the centroid of the i-th point cluster and the predicted value of the j-th track, where the combined cost is: ; in, This represents the combined cost between the centroid of the i-th point cluster and the predicted value of the j-th track. This represents the Euclidean distance between the centroid of the i-th point cluster and the predicted position of the j-th track. This represents the absolute value of the radial velocity difference between the i-th point cluster and the j-th track prediction. This represents the absolute value of the acceleration difference between the i-th point cluster and the j-th track prediction. , , These represent the weighting coefficients for the Euclidean distance term, the radial velocity difference term, and the acceleration difference term, respectively. The target attribute data output by clustering tracking includes three-dimensional position information, radial velocity, radar cross section (RCS), echo power, and cluster size information.

4. The fall detection method based on a fused neural network using millimeter-wave radar according to claim 1, characterized in that, Data preprocessing is performed on the target trajectory formed by the target attribute data, including: The data preprocessing includes data cleaning and normalization; The data cleaning includes removing invalid frames, removing data that is outside the radar detection range, and removing outliers caused by multipath, mismerging, misassociation, track start data, or track disappearance data. The normalization uses max-min normalization to map each feature to the interval [-1, 1].

5. The fall detection method based on a fusion neural network using millimeter-wave radar according to claim 1, characterized in that, Constructing a fixed-length time-series sample based on the trajectory data includes: A fixed-length time-series sample is constructed from the trajectory data using a sliding window approach; The category labeling uses visual data collected synchronously with millimeter-wave radar as a reference to label each sequence sample with a category label, including falling, normal walking, sitting, bending over, and abnormal postures in industrial scenarios.

6. The fall detection method based on a fused neural network using millimeter-wave radar according to claim 1, characterized in that, The one-dimensional convolutional branch includes a first-level one-dimensional convolutional layer, a first ReLU nonlinear mapping layer, a max pooling layer, a second-level one-dimensional convolutional layer, a second ReLU nonlinear mapping layer, and a global mean pooling layer; wherein, the first-level one-dimensional convolutional layer has 32 output channels, a kernel size of 2, a stride of 1, and a same padding method; the max pooling layer has a stride of 2; the second-level one-dimensional convolutional layer has 64 output channels, a kernel size of 2, a stride of 1; and the global mean pooling layer outputs a 64-dimensional local feature vector. The gated recurrent unit branch includes a cascaded first-level recurrent unit (GRU) and second-level recurrent unit (GRU); the first-level recurrent unit (GRU) has an input dimension of F, 64 hidden units, and returns a complete sequence; the second-level recurrent unit (GRU) has 64 hidden units and only returns the final state; the gated recurrent unit branch outputs a 64-dimensional global feature vector. The feature fusion and decision layer includes a feature dimension concatenation layer, a fully connected mapping layer, and a Softmax classification layer, which respectively satisfy: ; ; ; in, This represents the local feature vector output by a one-dimensional convolution branch. This represents the global feature vector output by the branch of the gated recurrent unit. denoted as fused feature vector, [ ] indicates concatenation according to feature dimension, W represents the weight matrix of the fully connected layer, b represents the bias vector of the fully connected layer, h represents the binary logit vector output by the fully connected layer, and ŷ represents the probability vector output by the Softmax classification layer.

7. The fall detection method based on a fused neural network using millimeter-wave radar according to claim 1, characterized in that, Training the fusion neural network using the training set includes: Cross-entropy is used as the optimization objective to measure the difference between the predicted distribution and the true label. The cross-entropy loss function is expressed as: ; in, This represents the weighted cross-entropy loss of the current batch of samples. Indicates the number of samples in the batch. Indicates the sample number. Indicates the first The true label of each sample It means to fall down. Indicates not a fall. The model predicts the first... The probability that a sample belongs to the "falling" category.

8. The fall detection method based on a fusion neural network for millimeter-wave radar according to claim 1, characterized in that, Training the fusion neural network using the training set further includes: The optimizer uses adaptive moment estimation (Adam), with an initial learning rate of 0.001, a first-order moment decay coefficient of 0.9, a second-order moment decay coefficient of 0.999, and a numerical stability term ε= The batch size for each parameter iteration is set to 16, and the maximum number of iterations is set to 200. The network weights are updated iteratively through backpropagation.

9. The fall detection method based on a fusion neural network for millimeter-wave radar according to claim 1, characterized in that, Using the validation set to select and save the optimal model includes: After each training round, the current model's classification accuracy is verified, and the model is switched to inference mode to avoid gradient calculation and parameter updates. The validation set is input into the model for forward inference to obtain the predicted probability output; The accuracy rate is calculated to verify the model's precision, using the following formula: ; Where Acc represents the accuracy rate; TP represents the number of fall samples that were correctly detected as falls; TN represents the number of non-fall samples that were correctly identified as non-fall samples; FP represents the number of non-fall samples that were falsely detected as falls; and FN represents the number of fall samples that were missed as non-fall samples. The model saving strategy is as follows: if the current validation accuracy (Acc) exceeds the historical best accuracy... Then save the current network weights. and will Update to the current validation accuracy Acc; if the accuracy does not improve for 10 consecutive rounds, terminate training early to prevent overfitting. The evaluation of the generalization performance of the optimal model using the test set includes: Calculate accuracy (Acc), recall, precision, and F1 score. Acc is used to validate model accuracy, recall measures the ability to detect fall events, precision measures the confidence of alarms, and the F1 score is used to calculate the harmonic mean of precision and recall to comprehensively measure both metrics. The calculations are as follows: ; ; ; ; Where Acc represents accuracy; TP represents the number of fall samples correctly detected as fall samples; TN represents the number of non-fall samples correctly identified as non-fall samples; FP represents the number of non-fall samples falsely detected as fall samples; FN represents the number of fall samples missed as non-fall samples; Recall represents recall; Precision represents precision; and F1 represents the F1 score.

10. A millimeter-wave radar fall detection system based on a fused neural network, characterized in that, include: The raw point cloud data acquisition module is used to acquire raw point cloud data detected by millimeter-wave radar. The raw point cloud data includes at least one of the following: distance, radial velocity, horizontal angle, elevation angle, radar cross section, and echo power of the point target. The moving point acquisition module is used to perform motion and static discrimination on the original point cloud data based on the absolute value of the radial velocity of the point target, and retain the moving points that meet the motion discrimination conditions; The clustering tracking module is used to perform clustering tracking on the moving point, including: introducing an anisotropic weighting strategy in the distance metric of density clustering, assigning a weight factor to the vertical direction that is lower than that to the horizontal direction, so that the clustering decision focuses more on the spatial continuity of the target in the horizontal plane relative to the vertical direction, and predicting the state vector of the existing track for the current frame based on the Kalman filter. The output module is used to output target attribute data based on clustering tracking. The trajectory data acquisition module is used to preprocess the target trajectory formed by the target attribute data to obtain trajectory data. A time-series sample construction module is used to construct time-series samples of a fixed length based on the trajectory data; The data partitioning module is used to classify the time-series samples and divide them into training set, validation set and test set; A fusion neural network construction module is used to construct a dual-branch parallel fusion neural network, including a one-dimensional convolutional branch, a gated recurrent unit branch, and a feature fusion and decision layer. The one-dimensional convolutional branch extracts the temporal samples and outputs local transient change features in the fall action, while the gated recurrent unit branch extracts and outputs a global feature vector showing the transition of the human body from a normal posture to a fall state. The feature fusion and decision layer concatenates the local transient change features and the global feature vector by their feature dimensions to obtain a fused feature vector. After performing Softmax normalization on the fused feature vector, it outputs the classification probabilities of the fall category and the non-fall category. A training module is used to train the fusion neural network using the training set; The optimal model acquisition module is used to select a model using the validation set and save the optimal model. A generalization performance evaluation module is used to evaluate the generalization performance of the optimal model using the test set. A detection module is used to perform real-time fall detection using the optimal model after generalization performance.