FMCW radar-based dual-channel deep reinforcement learning tumble prediction method and system

By using a dual-channel deep reinforcement learning method based on FMCW radar, a dual-channel fall detection network is constructed, which solves the problem of low recognition rate in existing fall detection technology, achieves higher fall prediction accuracy and stability, adapts to complex environments and provides detailed human body status information.

CN120784009APending Publication Date: 2025-10-14NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510929162.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing fall detection technologies have problems with low recognition rate and low recognition accuracy, especially methods based on deep learning. Traditional methods are also susceptible to light influence and privacy issues.

Method used

A dual-channel deep reinforcement learning method based on FMCW radar is adopted. By constructing a dual-channel fall detection network, including ConvNeXt V2 module, SwinTransformer V2 module, cross-attention mechanism and intelligent agent, the FMCW radar dataset is preprocessed, trained and tested. The cross-attention mechanism is used to fuse features to predict the fall probability and stand-up probability, and feedback optimization is performed through the intelligent agent.

Benefits of technology

It improves the accuracy and stability of fall prediction, can adapt to complex and changeable actual application scenarios, reduce the uncertainty of single detection results, and provide more comprehensive human body status information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120784009A_ABST
    Figure CN120784009A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-channel deep reinforcement learning tumble prediction method and system based on an FMCW radar, and belongs to the technical field of only monitoring, and the method comprises the steps: obtaining and preprocessing an FMCW radar data set, and obtaining a target FMCW data set; constructing a dual-channel tumble detection network, and training and testing the dual-channel tumble detection network according to the target FMCW data set; acquiring to-be-predicted data, and processing the to-be-predicted data according to the dual-channel fall detection network to obtain a fall probability and a predicted standing-up probability corresponding to the fall probability; the dual-channel tumble detection network performs self feedback optimization according to the plurality of tumble probabilities and the plurality of predicted standing-up probabilities corresponding to the plurality of tumble probabilities; through combination of the FMCW radar technology and the deep reinforcement learning algorithm, the method can adapt to complex and changeable practical application scenes, can continuously learn and adapt to different tumble modes, and improves the accuracy of tumble prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent monitoring, and particularly relates to a double-channel deep reinforcement learning fall prediction method and system based on FMCW radar. BACKGROUND

[0002] With the aggravation of global population aging, the fall events of the elderly occur frequently, which has become a serious public health problem. However, the existing fall detection technology has certain limitations. For example, the detection based on video images is easily affected by factors such as light and shielding, and has privacy problems. The detection of wearable devices requires the elderly to wear devices at all times, which is inconvenient to use.

[0003] At present, the existing fall prediction method based on deep learning has certain limitations in recognition rate, which leads to low recognition rate and low recognition accuracy of the final model. SUMMARY

[0004] The purpose of the embodiment of the application is to provide a double-channel deep reinforcement learning fall prediction method and system based on FMCW radar, which can solve the technical problem of low recognition rate of the existing fall prediction method.

[0005] In order to solve the above technical problems, the application is implemented as follows: In a first aspect, the embodiment of the application provides a double-channel deep reinforcement learning fall prediction method based on FMCW radar, which comprises the following steps: An FMCW radar data set is acquired, and the FMCW radar data set is preprocessed to obtain a target FMCW data set; A double-channel fall detection network is constructed, which comprises a ConvNeXt V2 module, a SwinTransformer V2 module, a cross attention mechanism, a fall detection head and an agent; Part of the data in the target FMCW data set is used to train the double-channel fall detection network, and another part of the data is used to test the trained double-channel fall detection network; The predicted data is acquired, and the predicted data is processed according to the tested double-channel fall detection network to obtain a fall probability and a predicted standing-up probability corresponding to the fall probability; The double-channel fall detection network is fed back and optimized according to a plurality of fall probabilities and a plurality of predicted standing-up probabilities corresponding to the fall probabilities, to obtain an optimized double-channel fall detection network.

[0006] As an optional implementation manner of the first aspect of the application, the FMCW radar data set is preprocessed to obtain a target FMCW data set, specifically: performing binary parsing on the FMCW radar dataset to obtain a parsed dataset; performing data reconstruction, Hamming window addition, and denoising processing on the parsed dataset to obtain a reconstructed dataset; performing twice fast Fourier transform processing on the reconstructed dataset to obtain a transformed FMCW dataset; performing normalization and data enhancement processing on the transformed FMCW dataset to obtain the target FMCW dataset; wherein the target FMCW dataset includes a plurality of range maps and a plurality of Doppler maps, and the range maps and the Doppler maps Figure One correspond to each other.

[0007] As an optional implementation of the first aspect of the application, the double-channel fall detection network is trained according to part of the data in the target FMCW dataset, and the trained double-channel fall detection network is tested according to another part of the data; specifically: dividing the target FMCW dataset into a target training set and a target test set according to a preset ratio; constructing a loss function, and training the double-channel fall detection network according to the training function and the target training set; testing the trained double-channel fall detection network according to the target test set to obtain the tested double-channel fall detection network.

[0008] As an optional implementation of the first aspect of the application, the to-be-predicted data includes a to-be-predicted range map and a to-be-predicted Doppler map corresponding to the to-be-predicted range map; the to-be-predicted data is processed according to the tested double-channel fall detection network to obtain target prediction data, specifically: processing the to-be-predicted range map according to the ConvNeXt V2 module to obtain range features, and performing feature enhancement processing on the range features to obtain enhanced range features; processing the to-be-predicted Doppler map according to the Swin Transformer V2 module to obtain velocity features, and performing feature enhancement processing on the velocity features to obtain enhanced velocity features; performing feature fusion on the enhanced range features and the enhanced velocity features according to the cross-attention mechanism to obtain cross-fusion features; performing activation processing on the cross-fusion features according to the fall detection head to obtain a fall probability of the to-be-predicted data; processing the fall probability according to the agent to obtain a prediction standing-up probability corresponding to the fall probability.

[0009] As an optional implementation of the first aspect of the application, the dual-channel fall detection network is fed back and optimized according to the plurality of fall probabilities and the plurality of predicted standing-up probabilities corresponding to the plurality of fall probabilities, to obtain an optimized dual-channel fall detection network; specifically: The plurality of fall probabilities and the plurality of predicted standing-up probabilities are averaged to obtain a fall probability mean and a predicted standing-up probability mean; The reward function in the agent feeds back and optimizes the dual-channel fall detection network by calling different reward strategies according to the fall probability mean and the predicted standing-up probability mean; specifically as follows: If the predicted standing-up probability mean belongs to a first range, a low-level reward strategy is called to optimize the dual-channel fall detection network; If the predicted standing-up probability mean belongs to a second range, a middle-level reward strategy is called to optimize the dual-channel fall detection network; If the predicted standing-up probability mean belongs to a third range, a high-level reward strategy is called to optimize the dual-channel fall detection network; If the predicted standing-up probability mean belongs to a fourth range, a special-level reward strategy is called to optimize the dual-channel fall detection network.

[0010] As an optional implementation of the first aspect of the application, the low-level reward strategy specifically includes: When the fall probability mean is greater than a first fall threshold, if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a first value, and if the agent incorrectly selects not to stand up, the reward degree of the agent is weakened by a third value; When the fall probability mean is less than or equal to the first fall threshold, if the agent correctly selects to stand up, the reward degree of the agent is enhanced by a third value; The middle-level reward strategy specifically includes: When the fall probability mean is greater than the first fall threshold, if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a second value, and if the agent incorrectly selects not to stand up, the reward degree of the agent is weakened by a second value; When the fall probability mean is less than or equal to the first fall threshold, if the agent correctly selects to stand up, the reward degree of the agent is enhanced by a second value; The high-level reward strategy specifically includes: When the fall probability mean is greater than a second fall threshold, if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a first value, and if the agent incorrectly selects not to stand up, the reward degree of the agent is weakened by a first value; When the mean of the fall probability is greater than the third fall threshold and less than or equal to the second fall threshold, if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a third numerical value, and if the agent incorrectly selects not to stand up, the reward degree of the agent is weakened by a third numerical value; When the mean of the fall probability is less than or equal to the third fall threshold, if the agent correctly selects to stand up, the reward degree of the agent is enhanced by a third numerical value, and if the agent incorrectly selects to stand up, the reward degree of the agent is weakened by a fourth numerical value; The special level reward strategy is specifically: If the agent incorrectly selects not to stand up, the reward degree of the agent is enhanced by a first calculated numerical value, and if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a second calculated numerical value. The first calculated numerical value is equal to twice the mean of the fall probability, and the second calculated numerical value is equal to twice the complement of the mean of the fall probability. In a second aspect, an embodiment of the present application provides a double-channel deep reinforcement learning fall prediction system based on an FMCW radar, the system comprising: An acquisition module: acquiring an FMCW radar data set and preprocessing the FMCW radar data set to obtain a target FMCW data set; A construction module: constructing a double-channel fall detection network, the double-channel fall detection network comprising a ConvNeXt V2 module, a Swin Transformer V2 module, a cross attention mechanism, a fall detection head, and an agent; A training module: training the double-channel fall detection network according to part of the data in the target FMCW data set, and testing the trained double-channel fall detection network according to another part of the data; A prediction module: acquiring to-be-predicted data, processing the to-be-predicted data according to the tested double-channel fall detection network to obtain a fall probability and a predicted standing-up probability corresponding to the fall probability; A feedback optimization module: the double-channel fall detection network performs self-feedback optimization according to a plurality of fall probabilities and a plurality of predicted standing-up probabilities corresponding to the fall probabilities, to obtain an optimized double-channel fall detection network.

[0011] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor, a memory, and a program or instructions stored on the memory and executable on the processor, the program or instructions being executed by the processor to implement the steps of the method of the first aspect.

[0012] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0013] Compared with the prior art, the embodiments of the present invention have the following technical effects: (1) By combining FMCW radar technology and deep reinforcement learning algorithms, it can adapt to complex and changeable actual application scenarios and can continuously learn and adapt to different fall patterns, thereby improving the accuracy of fall prediction.

[0014] (2) The dual-channel fall detection network performs self-feedback optimization based on multiple fall probabilities and multiple predicted stand-up probabilities. By averaging the probability means, the uncertainty of a single detection result is reduced, making the optimization process more stable and reliable. (3) The dual-channel fall detection network integrates the ConvNeXt V2 module and the Swin Transformer V2 module to extract the features of radar data from the local and global perspectives respectively, and performs feature fusion through the cross-attention mechanism, which can more comprehensively capture the motion characteristics and posture changes of the human body, thereby improving the accuracy of fall prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figures 1 is a flowchart of a dual-channel deep reinforcement learning fall prediction method based on FMCW radar provided by some embodiments of the present application; Figures 2 This is a dual-channel fall detection network structure diagram of a dual-channel deep reinforcement learning fall prediction method based on FMCW radar provided by some embodiments of the present application; Figures 3 This is a flowchart of self-feedback optimization of a dual-channel fall detection network in a dual-channel deep reinforcement learning fall prediction method based on FMCW radar provided by some embodiments of the present application. DETAILED DESCRIPTION

[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0017] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0018] Below, in conjunction with the accompanying drawings, a dual-channel deep reinforcement learning fall prediction method and system based on FMCW radar provided by the embodiment of the present application is described in detail through specific embodiments and their application scenarios.

[0019] Example A dual-channel deep reinforcement learning fall prediction method based on FMCW radar includes the following steps: S100: Acquire an FMCW radar dataset, and preprocess the FMCW radar dataset to obtain a target FMCW dataset; It is important to understand that the FMCW (Frequency Modulated Continuous Wave) radar dataset comes from human activity data collected by the University of Glasgow. It includes six behaviors: walking, sitting, standing, retrieving objects, drinking water, and falling. Each dataset has a .dat file, and the data files are named using the common method KPXXAYYRZ, where K (1, 2, 3, 4, 5, 6) represents the six behaviors; XX represents the subject (individual) with ID XX (01, 02, etc.); YY represents the ongoing activity, such as A01, A02, A03, A04, A05, and A06; and Z represents the number of repetitions of the activity, such as R1, R2, and so on.

[0020] Furthermore, in S100, the FMCW radar dataset is preprocessed to obtain a target FMCW dataset, specifically: S110: performing binary parsing on the FMCW radar data set to obtain a parsed data set; S120: Reconstruct the data, add a Hamming window, and perform denoising on the corresponding parsed data set to obtain a reconstructed data set; S130: Processing the reconstructed data set by two fast Fourier transforms to obtain a transformed FMCW data set; S140: performing normalization and data enhancement processing on the transformed FMCW dataset to obtain a target FMCW dataset; Among them, the target FMCW data set includes multiple range maps and multiple Doppler maps. Figure One One to one correspondence.

[0021] It needs to be understood that before data processing, the original.dat file is first parsed in binary and parameters are extracted, and the corresponding metadata is extracted, and the label information is extracted according to the file name. Since there are six behaviors, behaviors other than falling (1 represents) are considered as non-falling (0 represents); the parsed data is reconstructed into a rectangular shape of (nc, NTS), and a Hamming window is added before the fast Fourier transform calculation to avoid spectral leakage. The background is denoised by the CFAR algorithm, and the range map generation and Doppler map generation are obtained by twice FFT (fast Fourier transform) calculation; the range map and Doppler map and label data are normalized by Min-Max, and data augmentation such as data rotation and flipping is performed. In order to adapt to the input of the model, the size of the range map and the Doppler map is adjusted to 128x128, and the channel dimension is added (1, 128, 128).

[0022] It needs to be noted that through the steps of binary parsing, data reconstruction, Hamming window addition and denoising processing, the noise and redundant information in the original data are removed, the accuracy and reliability of the data are improved, and a high-quality data basis is provided for subsequent fall prediction. Twice fast Fourier transform processing converts the time domain signal into a range map and a Doppler map, which can effectively extract the distance and Doppler information of the target. These information is an important feature for judging whether the human body falls. The range map can reflect the distance change between the human body and the radar, and the Doppler map can reflect the movement speed and direction of the human body, providing rich feature information for fall prediction. Normalization processing makes the data have a unified dimension and numerical range, speeds up the convergence speed of the model, and improves the training efficiency and stability of the model. Data augmentation processing increases the diversity of data, so that the model can learn more features and patterns, improves the generalization ability of the model, and makes the model better adapt to different scenes and environments in practical application. The target FMCW data set contains multiple one-to-one corresponding range maps and Doppler maps. This structured data form is convenient for subsequent analysis and processing.

[0023] S200: Construct a dual-channel fall detection network, which includes a ConvNeXt V2 module, a Swin Transformer V2 module, a cross-attention mechanism, a fall detection head, and an intelligent agent. It should be understood that the ConvNeXt V2 module, as an improved version of the convolutional neural network, can extract local features in the radar data, such as the shape and motion pattern of the human body, and capture spatial features in the data through convolution operations; by combining self-supervised learning techniques and architecture improvements, the performance of the pure convolutional network (ConvNet) in various visual recognition tasks is significantly improved. The Swin Transformer V2 module is based on the Transformer architecture and is good at capturing long-distance dependencies and global features in the data, and can grasp the overall trend of human motion and posture changes; it can reduce the spatial dimension of the features while increasing the depth of the features, effectively capturing multi-scale features; the cross-attention mechanism is used to fuse the features extracted by the ConvNeXt V2 module and the Swin Transformer V2 module, so that the information of the two channels can complement and enhance each other, thereby more comprehensively understanding the information in the radar data; the fall detection head performs classification or regression operations on the fused features to determine whether a fall has occurred and outputs a fall probability; the agent plays a decision-making and optimization role in the dual-channel fall detection network, and optimizes the network according to the actual prediction results of the network.

[0024] S300: training the dual-channel fall detection network according to part of the data in the target FMCW data set, and testing the trained dual-channel fall detection network according to another part of the data; It should be understood that the target FMCW data set is divided into a training set and a test set. The trained dual-channel fall detection network is trained using the training set, and the parameters in the network are continuously adjusted through the backpropagation algorithm, so that the network can learn the features and patterns in the data and more accurately predict falls; the test set is used to test the trained network, and the performance of the network on unseen data is evaluated to ensure that the network has good generalization ability.

[0025] Further, S300 is specifically: S310: dividing the target FMCW data set into a target training set and a target test set according to a predetermined ratio; S320: constructing a loss function, and training the dual-channel fall detection network according to the training function and the target training set; S330: testing the trained dual-channel fall detection network according to the target test set to obtain a tested dual-channel fall detection network.

[0026] It needs to be understood that the target FMCW dataset is divided into a target training set and a target test set according to a preset ratio, and the purpose of this step is to ensure that the model can use different data in the training and testing stages, thereby avoiding overfitting and verifying the generalization ability of the model; generally, the division ratio of the dataset can be 70% training set and 30% test set, or other suitable ratios, depending on the size and complexity of the dataset; the training set is used for parameter updating and learning of the model, and the test set is used to evaluate the performance of the model on unseen data. The loss function is a function used to measure the difference between the model's predicted results and the true labels; constructing a suitable loss function is a key step in model training; in the fall prediction task, a classification loss function (such as cross-entropy loss) is used; the selection of the loss function needs to match the model output and the task target in order to effectively guide the learning process of the model. The target training set is used to train the dual-channel fall detection network, and during the training process, the model calculates the gradient of the loss function through the backpropagation algorithm and uses an optimizer (such as Adam, SGD, etc.) to update the parameters of the model to minimize the loss function; the training process involves multiple epochs, and each epoch traverses the entire training set, gradually adjusting the parameters of the model to make it better fit the training data. The trained dual-channel fall detection network is tested using the target test set; during the testing process, the model predicts the samples in the test set and compares the predicted results with the true labels to evaluate the performance of the model; common evaluation indicators include accuracy, precision, recall, F1 score, etc. These indicators can comprehensively reflect the performance of the model on the test set, helping to judge whether the model has good generalization ability.

[0027] It needs to be understood that by dividing the dataset into a training set and a test set, it is ensured that the model does not come into contact with test data during the training process, thereby avoiding overfitting. The test set is used to verify the performance of the model on unseen data, ensuring that the model has good generalization ability. The loss function provides a clear goal for model training, and by minimizing the loss function, the model can gradually adjust its parameters to better fit the training data. This helps to improve the adaptability and accuracy of the model for the fall prediction task. The testing stage provides an objective evaluation of the performance of the model, helping to identify the strengths and weaknesses of the model in the fall prediction task. Through the test results, the model structure or training strategy can be further optimized to improve overall performance.

[0028] S400: Obtain the to-be-predicted data, process the to-be-predicted data according to the tested dual-channel fall detection network, and obtain the fall probability and the predicted standing-up probability corresponding to the fall probability; Among them, the to-be-predicted data includes a to-be-predicted range image and a to-be-predicted Doppler image corresponding to the to-be-predicted range image; It needs to be understood that the FMCW radar data to be predicted is obtained, which is input into the tested dual-channel fall detection network; the network processes the to-be-predicted data, first extracts features, and then outputs a fall probability through a fall detection head; at the same time, the network also predicts a corresponding predicted standing-up probability according to the fall probability, because in the actual scene, there may be a standing-up behavior after falling, and predicting the standing-up probability helps to more comprehensively understand the state change of the human body.

[0029] Further, in S400, the to-be-predicted data is processed according to the tested dual-channel fall detection network to obtain a fall probability and a predicted standing-up probability corresponding to the fall probability, specifically: S410: processing the to-be-predicted distance graph according to the ConvNeXt V2 module to obtain distance features, and performing feature enhancement processing on the distance features to obtain enhanced distance features; S420: processing the to-be-predicted Doppler graph according to the Swin Transformer V2 module to obtain speed features, and performing feature enhancement processing on the speed features to obtain enhanced speed features; S430: performing feature fusion on the enhanced distance features and the enhanced speed features according to the cross-attention mechanism to obtain cross-fusion features; S440: performing activation processing on the cross-fusion features according to the fall detection head to obtain a fall probability of the to-be-predicted data; S450: processing the fall probability according to the intelligent agent to obtain a predicted standing-up probability corresponding to the fall probability.

[0030] It needs to be understood that the to-be-predicted distance graph (1, 128, 128) is input into the ConvNeXt V2 module, and after 4 times of convolution and down-sampling, distance features with a dimension of (768, 32, 32) are obtained, and then feature enhancement is performed to output enhanced distance features with a dimension of (256, 32, 32) as a query matrix (Query) in the cross-attention mechanism; the to-be-predicted Doppler graph is input into the Swin Transformer V2 module to obtain speed features with the same dimension (768, 32, 32), and then feature enhancement is performed to output enhanced speed features with a dimension of (256, 32, 32) as a key / value (Key / Value) in the cross-attention mechanism. Through the cross-attention mechanism, the distance information and the Doppler information are fused, the similarity between Query and Key is calculated (usually using dot product), to determine the weight of Value, so as to extract the most relevant feature information of the current attention spatial position. The output dimension of the cross-attention mechanism in this embodiment is , through a fully connected layer , wherein is a weight matrix, is a bias term. is the output. The Softmax activation function is used to convert it into a fall probability. The probability output by the fall detection head affects the decision-making of subsequent actions as input to the agent. The difference between the output of the fall detection head and the true label is also used to optimize the network weights. The agent plays a role in decision-making and reasoning in the dual-channel fall detection network. It reasons and predicts the situation of standing up after falling based on the fall probability, combined with its knowledge base and pre-set rules.

[0031] It should be noted that by extracting distance features and speed features respectively and enhancing them, the motion information of the human body can be more comprehensively captured. Distance features reflect the change in distance between the human body and the radar, while speed features reflect the speed and direction of the human body. The fusion of the two makes the cross-fusion features contain more information, thereby improving the accuracy of fall probability calculation. In addition to obtaining the fall probability, the agent also calculates the predicted standing up probability. This provides more comprehensive information about the state of the human body, which helps to better understand the fall and recovery of the human body in practical applications, such as in the elderly care scene, appropriate measures can be taken in a timely manner based on this information. The application of feature enhancement processing and cross-attention mechanism enables the model to better cope with noise and interference in the data, enhancing the robustness of the model. Even in complex environments, the model can accurately extract and process features, improving the reliability of fall prediction.

[0032] The introduction of the agent enables the model to make intelligent decisions and reasoning, providing more valuable suggestions and action plans based on the fall probability and predicted standing up probability.

[0033] S500: The dual-channel fall detection network performs self-feedback optimization based on the multiple fall probabilities and the multiple predicted standing up probabilities corresponding to the multiple fall probabilities, to obtain an optimized dual-channel fall detection network.

[0034] It should be understood that the dual-channel fall detection network performs self-feedback optimization based on the multiple fall probabilities and the multiple corresponding predicted standing up probabilities; the agent analyzes these output results and compares them with the pre-set target to find problems and deficiencies in the prediction process of the network. Based on the analysis results, the agent adjusts the parameters of the network to optimize the structure and performance of the network, and obtains an optimized dual-channel fall detection network, which can be more accurate and reliable in subsequent prediction.

[0035] Further, S500 specifically comprises: S510: Average the multiple fall probabilities and the multiple predicted standing up probabilities to obtain the mean fall probability and the mean predicted standing up probability; S520: The reward function in the agent performs self-feedback optimization on the dual-channel fall detection network according to the mean of the fall probability and the mean of the predicted standing-up probability; the specific process is as follows: S521: If the mean of the predicted standing-up probability belongs to the first range, a low-level reward strategy is called to optimize the dual-channel fall detection network; S522: If the mean of the predicted standing-up probability belongs to the second range, a middle-level reward strategy is called to optimize the dual-channel fall detection network; S523: If the mean of the predicted standing-up probability belongs to the third range, a high-level reward strategy is called to optimize the dual-channel fall detection network; S524: If the mean of the predicted standing-up probability belongs to the fourth range, a special-level reward strategy is called to optimize the dual-channel fall detection network.

[0036] It should be understood that in actual application, multiple fall detection is performed on the same scene or the same object, thereby obtaining multiple fall probabilities and multiple predicted standing-up probabilities. In order to obtain more stable and more representative results, the probabilities are averaged; the mean of the fall probability is obtained by calculating the average of all fall probabilities, and the mean of the predicted standing-up probability is obtained by calculating the average of all predicted standing-up probabilities; this step can reduce the error and fluctuation of single detection result, so that the subsequent decision and optimization are based on more reliable data. The reward function in the agent is a key part for evaluating the performance of the dual-channel fall detection network and guiding its optimization; the reward function judges the performance of the current network according to the calculated mean of the fall probability and the mean of the predicted standing-up probability, and calls different reward strategies to feedback and optimize the network.

[0037] Specifically, the mean of the predicted standing-up probability is divided into different ranges (first range, second range, third range, and fourth range), and each range corresponds to a different reward strategy. The division of different ranges can be determined according to actual needs and application scenarios, for example, the first range represents a situation where the predicted standing-up probability is low and the possibility of standing up after falling is small; the fourth range represents a situation where the predicted standing-up probability is high and the possibility of standing up after falling is large. When the mean of the predicted standing-up probability belongs to a certain range, the agent calls the corresponding reward strategy to optimize the dual-channel fall detection network. Different levels of reward strategies have different adjustment intensity and direction on the network in the optimization process. The low-level reward strategy only fine-tunes the network parameters to adapt to the current situation; while the high-level or special-level reward strategy adjusts the network structure or parameters more significantly to promote the network to adapt to complex scenarios faster and improve the accuracy of fall detection.

[0038] It should be noted that averaging multiple probabilities reduces the uncertainty of individual detection results, making the model's output mean fall probability and predicted mean standing probability more stable. This helps provide more reliable results in practical applications and avoids misjudgments caused by single-detection errors. By invoking different reward strategies based on the range of the predicted mean standing probability, adaptive optimization of the dual-channel fall detection network is achieved. The agent can dynamically adjust the optimization strategy based on the current network performance, allowing the network to better adapt to different scenarios and data distributions. This adaptive optimization approach improves the model's generalization and performance, enabling it to maintain strong performance in complex and changing real-world application scenarios. Different levels of reward strategies allow for targeted optimization of the network for different situations. Low-level reward strategies can be used to handle common situations, ensuring basic network performance; high-level and special reward strategies can be used to handle complex or extreme situations, further improving network performance. This hierarchical optimization approach fully utilizes the network's potential and improves the accuracy and reliability of fall detection. The process by which the agent optimizes the network based on the mean probability by invoking reward strategies is actually a process of intelligent decision-making and control. It can automatically adjust the network parameters and structure based on the current detection results to achieve better performance. This intelligent decision-making and control capability makes the entire fall detection system more intelligent and automated, and can better meet the needs of practical applications.

[0039] Furthermore, the specific low-level reward strategies in S521 are: When the mean probability of falling is greater than the first falling threshold, if the agent correctly chooses not to stand up, the reward level of the agent is increased by a first value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by a third value; When the mean fall probability is less than or equal to the first fall threshold, if the agent correctly chooses to stand up, the reward level of the agent is enhanced by a third value; The specific S522 intermediate reward strategy is: When the mean probability of falling is greater than the first falling threshold, if the agent correctly chooses not to stand up, the reward level of the agent is enhanced by a second value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by a second value; When the mean falling probability is less than or equal to the first falling threshold, if the agent correctly chooses to stand up, the reward level of the agent is enhanced by a second value; The specific S523 intermediate and advanced reward strategies are as follows: When the mean probability of falling is greater than the second falling threshold, if the agent correctly chooses not to stand up, the reward level of the agent is enhanced by the first value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by the first value; When the mean of the falling probability is greater than the third falling threshold and less than or equal to the second falling threshold, if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a third numerical value, and if the agent incorrectly selects not to stand up, the reward degree of the agent is weakened by a third numerical value; When the mean of the falling probability is less than or equal to the third falling threshold, if the agent correctly selects to stand up, the reward degree of the agent is enhanced by a third numerical value, and if the agent incorrectly selects to stand up, the reward degree of the agent is weakened by a fourth numerical value; The special level reward strategy in S524 is specifically: If the agent incorrectly selects not to stand up, the reward degree of the agent is enhanced by a first calculated numerical value, and if the agent correctly selects not to stand up, the reward degree of the agent is enhanced by a second calculated numerical value. The first calculated numerical value is equal to twice the mean of the falling probability, and the second calculated numerical value is equal to twice the complement of the mean of the falling probability.

[0040] It needs to be understood that in the low-level reward strategy, when the mean of the fall probability is greater than the first fall threshold (0.7, high fall risk): if the agent correctly chooses not to stand up (i.e., actually not falling, and the agent gives the action of not standing up), it means that the network has made a reasonable decision in the current high fall risk scenario, so the reward degree is enhanced by the first value (2.0) to encourage the agent to continue to make such correct judgments in similar scenarios. If the agent incorrectly chooses not to stand up (i.e., actually not falling, but the agent does not give the action of not standing up), it means that the agent has made a wrong decision, so the reward degree is weakened by the third value (1.0) to prompt the agent to improve its judgment when encountering similar situations in the future. When the mean of the fall probability is less than or equal to the first fall threshold: if the agent correctly chooses to stand up (i.e., actually falls, and the agent gives the action of standing up), it means that the agent has made a correct judgment for the low-risk scenario, so the reward degree is enhanced by the third value to strengthen the agent's correct decision-making ability in such scenarios. In the intermediate-level reward strategy, when the mean of the fall probability is greater than the first fall threshold: if the correct choice is not to stand up, the reward degree is enhanced by the second value (1.5), and if the incorrect choice is not to stand up, the reward degree is weakened by the second value. When the mean of the fall probability is less than or equal to the first fall threshold: if the correct choice is to stand up, the reward degree is enhanced by the second value, further strengthening the agent's correct decision-making in low-risk scenarios. In the advanced-level reward strategy, when the mean of the fall probability is greater than the second fall threshold (0.9, extremely high fall risk): if the correct choice is not to stand up, the reward degree is enhanced by the first value to encourage the agent to make correct decisions in this risk range; if the incorrect choice is not to stand up, the reward degree is weakened by the first value to prompt the agent to improve the wrong judgment. When the mean of the fall probability is greater than the third fall threshold (0.6) and less than or equal to the second fall threshold: if the correct choice is not to stand up, the reward degree is enhanced by the third value to encourage the agent to make correct decisions in this risk range; if the incorrect choice is not to stand up, the reward degree is weakened by the third value to prompt the agent to improve the wrong judgment. When the mean of the fall probability is less than or equal to the third fall threshold: if the correct choice is to stand up, the reward degree is enhanced by the third value to strengthen the agent's correct decision-making in this risk range. If the incorrect choice is to stand up (i.e., actually falls, but the agent gives the action of not standing up), the reward degree is weakened by the fourth value (0.5) to prompt the agent to more accurately identify the fall situation in this risk range. In the special-level reward strategy, when the incorrect choice is not to stand up: the reward degree is enhanced by the first calculation value, which is equal to twice the mean of the fall probability. This means that when the agent makes a wrong decision not to stand up, the reward enhancement degree is proportional to the mean of the fall probability, the higher the mean of the fall probability, the greater the reward enhancement degree, to more strongly prompt the agent to improve the wrong decision. When the correct choice is not to stand up: the reward degree is enhanced by the second calculation value, which is equal to twice the complement of the mean of the fall probability.The complement of the mean of the fall probability is 1 minus the mean of the fall probability. When the mean of the fall probability is low, the complement is high, and the reward enhancement degree is greater at this time, encouraging the agent to correctly judge not to stand up in a low fall risk scenario.

[0041] It should be noted that the reward strategies of different levels enhance or weaken the degree of reward to different extents according to different ranges of the mean of the fall probability and the correctness of the agent's decision, which can accurately guide the dual-channel fall detection network to optimize. Make the network better adapt to different fall risk scenarios and improve the accuracy of fall detection. The high-level and special-level reward strategies make more detailed divisions and different processing methods for the mean of the fall probability, which can better adapt to complex and variable actual application scenarios. For example, different reward strategies are used when the fall risk is in different intervals, allowing the network to more accurately determine the fall and standing up of the human body. By enhancing the degree of reward when the agent makes a correct decision, the agent's behavior pattern in these situations is strengthened, making the network more inclined to make correct judgments when encountering similar scenarios later, improving the stability and reliability of the network. When the agent makes a wrong decision, the degree of reward is weakened, prompting the network to correct errors more quickly, reducing the frequency of incorrect decisions, and improving the overall performance of the network.

[0042] According to the dual-channel deep reinforcement learning fall prediction method based on FMCW radar, the ConvNeXt V2 module and the Swin Transformer V2 module in the dual-channel fall detection network process the range map and the Doppler map respectively, and extract the range features and the velocity features respectively; these two kinds of features describe the motion state of the human body from different angles, the range features reflect the distance change between the human body and the radar, and the velocity features reflect the motion speed and direction of the human body. The enhanced range features and the enhanced velocity features are fused through the cross-attention mechanism to obtain the cross-fusion features; the features fuse the information of the range and the velocity dimensions, which can more comprehensively and accurately describe the motion state of the human body, and provide more rich feature representation for the subsequent calculation of the fall probability and the predicted standing up probability. The dual-channel fall detection network processes the to-be-predicted data to obtain the accurate fall probability and the predicted standing up probability, the fall probability reflects the possibility of the human body falling, and the predicted standing up probability reflects the possibility of standing up after falling; these two probability indicators provide more comprehensive human state information for the user, which helps to better understand the fall and recovery of the human body in actual application; the self-feedback optimization is performed according to multiple fall probabilities and multiple predicted standing up probabilities; the uncertainty of the single detection result is reduced through the average processing of the probability mean, making the optimization process more stable and reliable.

[0043] It should be noted that the execution subject of the FMCW radar-based dual-channel deep reinforcement learning fall prediction method provided in the embodiments of the present application can be an FMCW radar-based dual-channel deep reinforcement learning fall prediction system, or a control module in the FMCW radar-based dual-channel deep reinforcement learning fall prediction system for executing the FMCW radar-based dual-channel deep reinforcement learning fall prediction method. In the embodiments of the present application, the FMCW radar-based dual-channel deep reinforcement learning fall prediction system is taken as an example to execute the FMCW radar-based dual-channel deep reinforcement learning fall prediction method, and the FMCW radar-based dual-channel deep reinforcement learning fall prediction method provided in the embodiments of the present application is described.

[0044] An FMCW radar-based dual-channel deep reinforcement learning fall prediction system comprises: An acquisition module: acquiring an FMCW radar dataset and preprocessing the FMCW radar dataset to obtain a target FMCW dataset; A construction module: constructing a dual-channel fall detection network, wherein the dual-channel fall detection network comprises a ConvNeXt V2 module, a Swin Transformer V2 module, a cross attention mechanism, a fall detection head and an agent; A training module: training the dual-channel fall detection network according to part of the data in the target FMCW dataset, and testing the trained dual-channel fall detection network according to another part of the data; A prediction module: acquiring to-be-predicted data, processing the to-be-predicted data according to the tested dual-channel fall detection network to obtain a fall probability and a predicted standing-up probability corresponding to the fall probability; A feedback optimization module: the dual-channel fall detection network performs self feedback optimization according to multiple fall probabilities and multiple predicted standing-up probabilities corresponding to the multiple fall probabilities, to obtain an optimized dual-channel fall detection network.

[0045] The FMCW radar-based dual-channel deep reinforcement learning fall prediction system in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an ios operating system or other possible operating systems, and the embodiments of the present application do not make specific limitations.

[0046] The FMCW radar-based dual-channel deep reinforcement learning fall prediction system provided in the embodiments of the present application can implement Figures 1 to 3 The processes implemented by the FMCW radar-based dual-channel deep reinforcement learning fall prediction method in the method embodiment are not repeated here to avoid repetition.

[0047] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the various processes of the above-mentioned embodiment of the dual-channel deep reinforcement learning fall prediction method based on FMCW radar are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here.

[0048] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned embodiment of the dual-channel deep reinforcement learning fall prediction method based on FMCW radar are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0049] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0050] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0051] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, also can be through hardware, but many cases the former is the better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part of the contribution to the prior art can be embodied in the form of software products, the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including a number of instructions to make a terminal (may be a mobile phone, computer, server, air conditioner, or network equipment, etc.) executes the method described in various embodiments of the present application.

[0052] The embodiments of the present application are described above in conjunction with the drawings, but the present application is not limited to the above-mentioned specific embodiments, the above-mentioned specific embodiments are only illustrative, but not limited, those skilled in the art can make many forms without departing from the purpose of the present application and the scope of the claims under the inspiration of the present application, all belong to the protection of the present application.

Claims

1. A dual-channel deep reinforcement learning fall prediction method based on FMCW radar, characterized in that: The method comprises: Acquire an FMCW radar dataset, and preprocess the FMCW radar dataset to obtain a target FMCW dataset; Constructing a dual-channel fall detection network, comprising a ConvNeXt V2 module, a SwinTransformer V2 module, a cross-attention mechanism, a fall detection head, and an agent; training the dual-channel fall detection network based on a portion of the data in the target FMCW dataset, and testing the trained dual-channel fall detection network based on another portion of the data; Acquiring data to be predicted, and processing the data to be predicted according to the tested dual-channel fall detection network to obtain a fall probability and a predicted standing probability corresponding to the fall probability; The dual-channel fall detection network performs self-feedback optimization according to the multiple fall probabilities and the multiple predicted standing-up probabilities corresponding to the multiple fall probabilities to obtain an optimized dual-channel fall detection network.

2. The dual-channel deep reinforcement learning fall prediction method based on FMCW radar according to claim 1 is characterized in that: The FMCW radar data set is preprocessed to obtain a target FMCW data set, specifically: performing binary parsing on the FMCW radar data set to obtain a parsed data set; Perform data reconstruction, Hamming window addition and denoising on the parsed data set to obtain a reconstructed data set; Processing the reconstructed data set by two fast Fourier transforms to obtain a transformed FMCW data set; performing normalization and data enhancement processing on the transformed FMCW dataset to obtain the target FMCW dataset; The target FMCW data set includes a plurality of range maps and a plurality of Doppler maps, and the range maps and the Doppler maps are in one-to-one correspondence.

3. The dual-channel deep reinforcement learning fall prediction method based on FMCW radar according to claim 1 is characterized in that: The dual-channel fall detection network is trained according to a portion of the data in the target FMCW data set, and the trained dual-channel fall detection network is tested according to another portion of the data; specifically: Dividing the target FMCW dataset into a target training set and a target test set according to a preset ratio; Constructing a loss function and training the dual-channel fall detection network according to the training function and the target training set; The trained dual-channel fall detection network is tested according to the target test set to obtain the tested dual-channel fall detection network.

4. The dual-channel deep reinforcement learning fall prediction method based on FMCW radar according to claim 1 is characterized in that: The data to be predicted includes a distance map to be predicted and a Doppler map to be predicted corresponding to the distance map to be predicted; the data to be predicted is processed according to the tested dual-channel fall detection network to obtain target prediction data, specifically: Processing the distance map to be predicted according to the ConvNeXt V2 module to obtain a distance feature, and performing feature enhancement processing on the distance feature to obtain an enhanced distance feature; Processing the Doppler image to be predicted according to the Swin Transformer V2 module to obtain a velocity feature, and performing feature enhancement processing on the velocity feature to obtain an enhanced velocity feature; Performing feature fusion on the enhanced distance feature and the enhanced speed feature according to the cross attention mechanism to obtain a cross fusion feature; activating the cross-fusion features according to the fall detection head to obtain the fall probability of the data to be predicted; The falling probability is processed according to the intelligent agent to obtain the predicted standing up probability corresponding to the falling probability.

5. The dual-channel deep reinforcement learning fall prediction method based on FMCW radar according to claim 4 is characterized in that: The dual-channel fall detection network performs self-feedback optimization based on the multiple fall probabilities and the multiple predicted standing-up probabilities corresponding to the multiple fall probabilities to obtain an optimized dual-channel fall detection network; specifically: Averaging the plurality of fall probabilities and the plurality of predicted stand-up probabilities to obtain a mean fall probability and a mean predicted stand-up probability; The reward function in the agent performs self-feedback optimization by invoking different reward strategies on the dual-channel fall detection network based on the mean fall probability and the predicted mean stand-up probability; specifically, as follows: If the predicted mean standing probability falls within a first range, invoking a low-level reward strategy to optimize the dual-channel fall detection network; If the predicted mean standing probability falls within the second range, invoking the intermediate reward strategy to optimize the dual-channel fall detection network; If the predicted mean standing probability falls within a third range, invoking an advanced reward strategy to optimize the dual-channel fall detection network; If the predicted mean standing probability falls within the fourth range, a special reward strategy is invoked to optimize the dual-channel fall detection network.

6. The dual-channel deep reinforcement learning fall prediction method based on FMCW radar according to claim 5 is characterized in that: The low-level reward strategy is specifically: When the mean fall probability is greater than a first fall threshold, if the agent correctly chooses not to stand up, the reward level of the agent is increased by a first value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by a third value; When the mean of the fall probability is less than or equal to the first fall threshold, if the agent correctly chooses to stand up, the reward level of the agent is enhanced by a third value; The intermediate reward strategy is specifically as follows: When the mean fall probability is greater than the first fall threshold, if the agent correctly chooses not to stand up, the reward level of the agent is enhanced by a second value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by a second value; When the mean of the fall probability is less than or equal to the first fall threshold, if the agent correctly chooses to stand up, the reward level of the agent is enhanced by a second value; The advanced reward strategy is specifically as follows: When the mean fall probability is greater than a second fall threshold, if the agent correctly chooses not to stand up, the reward level of the agent is enhanced by a first value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by the first value; When the mean fall probability is greater than the third fall threshold and less than or equal to the second fall threshold, if the agent correctly chooses not to stand up, the reward level of the agent is enhanced by a third value; if the agent incorrectly chooses not to stand up, the reward level of the agent is weakened by a third value; When the mean fall probability is less than or equal to a third fall threshold, if the agent correctly chooses to stand up, the reward level of the agent is increased by a third value; if the agent incorrectly chooses to stand up, the reward level of the agent is decreased by a fourth value; The special reward strategy is specifically as follows: If the agent incorrectly chooses not to stand up, the agent's reward level is enhanced by a first calculated value; If the agent correctly chooses not to stand up, the agent's reward level is enhanced by a second calculated value; The first calculated value is equal to twice the mean of the fall probability, and the second calculated value is equal to twice the complement of the mean of the fall probability.

7. A dual-channel deep reinforcement learning fall prediction system based on FMCW radar, capable of implementing the dual-channel deep reinforcement learning fall prediction method based on FMCW radar according to any one of claims 1 to 6, characterized in that: The system comprises: Acquisition module: acquires an FMCW radar dataset and preprocesses the FMCW radar dataset to obtain a target FMCW dataset; Building modules: Build a dual-channel fall detection network, which includes a ConvNeXt V2 module, a Swin Transformer V2 module, a cross-attention mechanism, a fall detection head, and an agent; Training module: training the dual-channel fall detection network according to a portion of the target FMCW dataset, and testing the trained dual-channel fall detection network according to another portion of the data; Prediction module: obtains data to be predicted, processes the data to be predicted according to the tested dual-channel fall detection network, and obtains a fall probability and a predicted standing probability corresponding to the fall probability; Feedback optimization module: The dual-channel fall detection network performs self-feedback optimization based on the multiple fall probabilities and the multiple predicted standing-up probabilities corresponding to the multiple fall probabilities to obtain an optimized dual-channel fall detection network.

8. An electronic device, characterized in that: The invention comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the dual-channel deep reinforcement learning fall prediction method based on FMCW radar are implemented as described in claims 1-6.

9. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the dual-channel deep reinforcement learning fall prediction method based on FMCW radar are implemented as described in claims 1-6.