Method and device for behavior detection based on multi-domain data fusion of millimeter wave radar
By employing a multi-domain data fusion method based on millimeter-wave radar, utilizing range-velocity spectrum and point cloud data features, and combining them with a frame feature fusion network, the problem of high false alarm rate in behavior detection in privacy-sensitive locations was solved, achieving efficient behavior recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-14
- Publication Date
- 2026-03-31
AI Technical Summary
Existing behavior detection technologies based on vision and wearable sensors suffer from high false alarm rates in privacy-sensitive scenarios, and millimeter-wave radar often experiences system misidentification during identification, making it difficult to deploy effectively in privacy-sensitive locations.
A multi-domain data fusion method based on millimeter-wave radar is adopted. By determining the range-velocity spectrum, micro-Doppler signal and point cloud data, and combining micro-Doppler features and point cloud features, a frame feature fusion network is used for behavior detection to reduce the false alarm rate.
By extracting human motion speed and spatial location features and combining them with a frame feature fusion network, the false alarm rate of behavior detection is reduced and the detection effect is improved.
Smart Images

Figure CN116400313B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of millimeter-wave radar technology, and more specifically, to a behavior detection method and apparatus based on multi-domain data fusion of millimeter-wave radar. Background Technology
[0002] With my country's aging population trend intensifying, the demand for care among the elderly is growing. Commonly used behavior detection technologies include vision-based and wearable sensor-based behavior monitoring systems. However, both of these mainstream technologies require consideration of the specific application context of indoor home monitoring during practical deployment. For example, to protect privacy, vision-based behavior detection solutions cannot be deployed in private locations such as bedrooms and bathrooms; wearable devices are inconvenient to carry in certain scenarios and require frequent charging. Millimeter-wave radar has unique advantages in these privacy-sensitive scenarios, and thanks to the penetrating power of millimeter waves, it can perform normal identification even under obstructed conditions. However, millimeter-wave radar is prone to false alarms. Therefore, reducing the false alarm rate of behavior detection and ensuring its detection effectiveness has become an urgent technical problem to be solved. Summary of the Invention
[0003] The embodiments of this application provide a behavior detection method and apparatus based on multi-domain data fusion of millimeter-wave radar, which can at least reduce the false alarm rate of behavior detection and ensure its detection effect to a certain extent.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0005] According to one aspect of the embodiments of this application, a behavior detection method based on millimeter-wave radar multi-domain data fusion is provided, the method comprising:
[0006] Based on the echo signal obtained by the millimeter-wave radar from the space to be detected, determine its corresponding range-velocity spectrum;
[0007] Based on the distance-velocity spectrum, each identical velocity grid point at each distance is summed to obtain a single frame of micro-Doppler signal;
[0008] Based on the range-velocity spectrum corresponding to multiple antennas, the point cloud data corresponding to the space to be detected is determined, and the point cloud data frame corresponds to the micro-Doppler signal in time;
[0009] Based on the micro-Doppler signal, it is determined whether an action has occurred. If an action is determined to have occurred, the detection data corresponding to the action is acquired. The detection data includes several frames of micro-Doppler signals and their corresponding point cloud data.
[0010] Feature extraction is performed on several frames of micro-Doppler signals and their corresponding point cloud data contained in the data to be detected, and the corresponding micro-Doppler features and point cloud features are determined.
[0011] The micro-Doppler features and the point cloud features are input into a pre-trained frame feature fusion network so that the frame feature fusion network outputs the corresponding action classification result.
[0012] According to one aspect of the embodiments of this application, a behavior detection device based on millimeter-wave radar multi-domain data fusion is provided, the device comprising:
[0013] The first determining module is used to determine the corresponding range-velocity spectrum based on the echo signal obtained by the millimeter-wave radar from the space to be detected.
[0014] The second determining module is used to add up each identical velocity grid point at each distance according to the distance-velocity spectrum to obtain a single frame of micro-Doppler signal;
[0015] The third determining module is used to determine the point cloud data corresponding to the space to be detected based on the range-velocity spectrum corresponding to multiple antennas, wherein the point cloud data corresponds to the micro-Doppler signal in time;
[0016] The acquisition module is used to determine whether an action has occurred based on the micro-Doppler signal. If the action is determined to have occurred, the module acquires the detection data corresponding to the action. The detection data includes several frames of micro-Doppler signals and their corresponding point cloud data.
[0017] The extraction module is used to extract features from several frames of micro-Doppler signals and their corresponding point cloud data contained in the data to be detected, and to determine their corresponding micro-Doppler features and point cloud features.
[0018] The processing module is used to input the micro-Doppler features and the point cloud features into a pre-trained frame feature fusion network, so that the frame feature fusion network outputs the corresponding action classification result.
[0019] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the behavior detection method based on millimeter-wave radar multi-domain data fusion as described in the above embodiments.
[0020] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the behavior detection method based on millimeter-wave radar multi-domain data fusion as described in the above embodiments.
[0021] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the behavior detection method based on millimeter-wave radar multi-domain data fusion provided in the above embodiments.
[0022] In some embodiments of this application, the technical solutions are as follows: The range-velocity spectrum is determined based on the echo signal obtained by millimeter-wave radar probing the space to be detected. Then, based on this range-velocity spectrum, each velocity grid point at each distance is summed to obtain a single frame of micro-Doppler signal. Point cloud data corresponding to the space to be detected is determined based on the range-velocity spectrum corresponding to multiple antennas. This point cloud data corresponds to the micro-Doppler signal in time. The micro-Doppler signal is then used to determine whether an action has occurred. If an action is determined, the corresponding detection data is acquired. This detection data includes several frames of micro-Doppler signals and their corresponding point cloud data. Feature extraction is performed on each frame of micro-Doppler signals and their corresponding point cloud data contained in the detection data to determine their corresponding micro-Doppler features and point cloud features. Finally, the micro-Doppler features and point cloud features are input into a pre-trained frame feature fusion network to output the corresponding action classification result. Therefore, by extracting the motion speed and spatial position features of the human body in the space to be detected, and combining the two features to output the corresponding behavior detection results, the false alarm rate of behavior detection is reduced and the detection effect is guaranteed.
[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0025] Figure 1 A flowchart illustrating an embodiment of a behavior detection method based on millimeter-wave radar multi-domain data fusion according to this application is shown.
[0026] Figure 2 A schematic flowchart illustrating the generation of a micro-Doppler signal according to an embodiment of this application is shown;
[0027] Figure 3 A schematic flowchart of the action detection process in a behavior detection method based on millimeter-wave radar multi-domain data fusion according to an embodiment of this application is shown;
[0028] Figure 4 A schematic diagram of the architecture of a time point cloud feature extraction network according to an embodiment of this application is shown;
[0029] Figure 5 A schematic diagram of the architecture of a frame feature fusion network according to an embodiment of this application is shown;
[0030] Figure 6 The diagram illustrates an application scenario where the technical solutions of the embodiments of this application can be applied.
[0031] Figure 7 A block diagram of a behavior detection device based on millimeter-wave radar multi-domain data fusion according to an embodiment of this application is shown;
[0032] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation
[0033] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0034] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0035] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0036] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0037] Figure 1 A flowchart illustrating an embodiment of a behavior detection method based on millimeter-wave radar multi-domain data fusion according to this application is shown. (Refer to...) Figure 1 As shown, the method includes at least steps S110 to S160, which are described in detail below:
[0038] In step S110, the range-velocity spectrum is determined based on the echo signal obtained by the millimeter-wave radar from the space to be detected.
[0039] In this embodiment, based on the care needs of the elderly, the millimeter-wave radar can send electromagnetic waves to the space to be detected (such as one or more of the bedroom, living room, and bathroom) in real time to sense the surrounding environment, and the radar can receive the corresponding echo signals.
[0040] In one example, the millimeter-wave radar sensor is a multi-antenna array sensor that emits electromagnetic waves with a length of 1-10 mm. It is a sensor capable of acquiring information about the external environment through calculation. This includes dedicated millimeter-wave radar sensors, or communication devices operating within the millimeter-wave range, such as millimeter-wave wireless network transmitters, and so on.
[0041] like Figure 2 As shown, upon receiving the corresponding echo signal, its corresponding range-velocity spectrum can be determined. Specifically, the echo signal is first mixed with the transmitted signal in a mixer to obtain an intermediate frequency (IF) signal, which is then sampled by an ADC. A range-dimensional FFT is then performed on the sampled signal to obtain the range-intensity spectrum. Furthermore, in indoor environments, electromagnetic waves reflected from static objects severely interfere with human body monitoring; therefore, static noise removal can be performed on the signal to eliminate static wall surface signals. A velocity-dimensional FFT is then performed to obtain the range-velocity spectrum.
[0042] In step S120, based on the distance-velocity spectrum, each identical velocity grid point at each distance is added together to obtain a single frame of micro-Doppler signal.
[0043] In this embodiment, such as Figure 2 As shown, summing the values of each identical velocity grid point at each distance (i.e., summing the rangeBins) yields the micro-Doppler signal for a single frame. Connecting the signals from consecutive frames together provides the signal shown in the image. Figure 2 The last matrix shown contains information about the speed and time of the action.
[0044] In step S130, point cloud data corresponding to the space to be detected is determined based on the range-velocity spectrum corresponding to multiple antennas, and the point cloud data corresponds to the micro-Doppler signal in time.
[0045] In one example, an FFT algorithm can be performed based on the range-velocity spectrum corresponding to multiple antennas to determine the point cloud data corresponding to the space to be detected. It should be understood that the point cloud data and the micro-Doppler signal have a temporal correspondence, that is, according to the time sequence, one frame of micro-Doppler signal corresponds to one frame of point cloud data.
[0046] In another example, the MUSIC estimation algorithm can be applied to the range-intensity spectrum corresponding to multiple antennas to obtain angle information, thereby determining the corresponding point cloud data. This method has super-resolution characteristics, but the computational cost is large.
[0047] It should be noted that those skilled in the art can determine the corresponding point cloud data determination method according to actual implementation needs, and no special limitations are made in this regard.
[0048] In step S140, it is determined whether an action has occurred based on the micro-Doppler signal. If an action is determined to have occurred, the detection data corresponding to the action is acquired. The detection data includes several frames of micro-Doppler signals and their corresponding point cloud data.
[0049] In this embodiment, due to the characteristics of micro-Doppler signals, it is possible to determine whether an action has occurred based on the acquired micro-Doppler signals. If an action is determined to have occurred, the detection data corresponding to the action can be acquired. This detection data includes several frames of micro-Doppler signals and their corresponding point cloud data. This allows for accurate identification of the detection data in subsequent steps, ensuring the accuracy of the subsequent identification.
[0050] In step S150, feature extraction is performed on several frames of micro-Doppler signals and their corresponding point cloud data contained in the data to be detected, and the corresponding micro-Doppler features and point cloud features are determined.
[0051] In this embodiment, feature extraction is performed on several frames of micro-Doppler signals and their corresponding point cloud data contained in the data to be detected, thereby obtaining the corresponding micro-Doppler features and point cloud features. Thus, the motion velocity features and spatial position features of the human body in the space to be detected can be obtained.
[0052] In step S160, the micro-Doppler features and the point cloud features are input into a pre-trained frame feature fusion network so that the frame feature fusion network outputs the corresponding action classification result.
[0053] In this embodiment, those skilled in the art can pre-build and train a frame feature fusion network. In actual use, the extracted micro-Doppler features and corresponding point cloud features can be input into the frame feature fusion network so that the frame feature fusion network can output the corresponding action classification results, such as walking, falling, sitting down, jumping, standing up, running, etc.
[0054] Therefore, in Figure 1 In the illustrated embodiment, the range-velocity spectrum corresponding to the echo signal obtained by the millimeter-wave radar probing the space to be detected is determined. Based on this range-velocity spectrum, each velocity grid point at each distance is added together to obtain a single frame of micro-Doppler signal. Based on the range-velocity spectrum corresponding to multiple antennas, the point cloud data corresponding to the space to be detected is determined. This point cloud data corresponds to the micro-Doppler signal in time. Then, based on the micro-Doppler signal, it is determined whether an action has occurred. If an action is determined to have occurred, the detection data corresponding to the action is acquired. The detection data includes several frames of micro-Doppler signals and their corresponding point cloud data. Features are extracted from the several frames of micro-Doppler signals and their corresponding point cloud data contained in the detection data to determine their corresponding micro-Doppler features and point cloud features. Then, the micro-Doppler features and point cloud features are input into a pre-trained frame feature fusion network to output the corresponding action classification result. Therefore, by extracting the motion speed and spatial position features of the human body in the space to be detected, and combining the two features to output the corresponding behavior detection results, the false alarm rate of behavior detection is reduced and the detection effect is guaranteed.
[0055] based on Figure 1 In one embodiment of this application, as shown in the example, it is determined whether an action has occurred based on the micro-Doppler signal. If an action is determined to have occurred, the detection data corresponding to the action is acquired, including:
[0056] When the predetermined conditions are met, the multi-frame micro-Doppler signal corresponding to the predetermined time length is acquired. The position with the highest motion correlation in the multi-frame micro-Doppler signal is determined by the mask convolution algorithm, and the signal is divided into velocity energy grid points and non-velocity energy grid points.
[0057] An action detection algorithm is used to calculate the action correlation of the segmented microDoppler signal to determine whether any action has occurred.
[0058] If an action is determined to have occurred, the micro-Doppler signal and point cloud data within a predetermined range before and after the correlation peak point of the action are copied and acquired as the data to be detected.
[0059] In this embodiment, it should be understood that the action is a continuous process. Therefore, the predetermined conditions can be detection triggering conditions pre-determined by those skilled in the art based on prior experience. For example, the predetermined conditions can be a fixed time interval, energy threshold judgment, or a mask convolution segmentation algorithm (see below). When the predetermined conditions are met, such as reaching a fixed time interval, multiple frames of micro-Doppler signals corresponding to a predetermined time length (e.g., 2s, 5s, etc.) can be acquired. The position with the highest action correlation in the multiple frames of micro-Doppler signals is determined by the mask convolution algorithm, and an activation function is used to segment the micro-Doppler signals into those with velocity energy grids and those without. It should be understood that when there is reflective motion, the energy amplitude at that position in the micro-Doppler signal will increase, and the corresponding velocity grid value will increase. Thus, through segmentation, the segment in which the action actually occurs can be accurately located, the interference of preceding and following actions can be shielded, and subsequent segments without action can be ignored, improving the accuracy and efficiency of the system's recognition.
[0060] Then, an action detection algorithm is used to calculate the action correlation of the segmented micro-Doppler signal to determine whether action has occurred. For example, action detection can be performed using energy peak detection or fixed threshold judgment. Once action is determined to have occurred, the micro-Doppler signal and point cloud data within a predetermined range before and after the correlation peak point of the action can be copied as the detection data. For example, the micro-Doppler signal and point cloud data of 20 frames before and after the correlation peak point of the action can be copied as the detection data.
[0061] In one embodiment of this application, a masked convolutional-peak detection method is used for action detection. For details, please refer to... Figure 3The mask segment is a special segment of length m containing the velocity characteristics of human motion. In some instances, it can be obtained by weighting typical micro-Doppler segments of the human motion to be measured. The mask segment performs convolution operations on multiple frames of micro-Doppler signals corresponding to a predetermined time length at one or more steps. The mask segment pools the n frames of micro-Doppler signals within the segment window and outputs the result, i.e., W frames of mask convolution results (where W = nm / detha). This output result contains the correlation between the motion data within this segment window and the typical motion data. After completing the mask operation on one frame, the operation is repeated by moving forward one frame until n frames (i.e., the number of frames corresponding to the predetermined time length) of micro-Doppler data are obtained. The length of the correlation data obtained by convolution is W.
[0062] At this point, a peak detection algorithm is used to detect whether there are correlation maxima in the W-frame mask convolution results. If so, a motion confidence score is calculated for each maxima. This confidence score can be used to represent the probability that motion actually occurred at that location. It is understandable that slight human movements, such as moving a leg or even writing, can generate peaks. Therefore, an n-order Markov model can be used to jointly calculate the magnitude of the extreme value with the motion correlation data before and after the extreme value (as a confidence score) and make a judgment. If the confidence score exceeds a threshold, the location with the highest confidence score during the motion is determined to identify the motion location. After determining the motion location, redundant micro-Doppler signals before and after the motion are removed, i.e., only the micro-Doppler signals within a predetermined range before and after the correlation peak point are retained. This allows for the determination of motion occurrence with a smaller time delay, thereby improving time sensitivity, enabling more timely identification of user actions, and reducing system false alarms.
[0063] It should be understood that those skilled in the art can also use other signal processing methods to segment user behavior action segments. In this case, if the user action occurrence determination condition is not met, it is identified as no action has occurred, and the system returns to continue receiving and listening to data. When the action occurrence determination condition is met, the micro-Doppler signal and point cloud data within a predetermined range before and after the correlation peak point of the action that occurred are copied as the data to be detected.
[0064] In one embodiment of this application, feature extraction is performed on several frames of micro-Doppler signals and their corresponding point cloud data contained in the data to be detected, to determine their corresponding micro-Doppler features and point cloud features, including:
[0065] For the several frames of micro-Doppler signals contained in the data to be detected, a pre-trained CNN network or RNN network is used to extract features from them;
[0066] Several frames of point cloud data and their corresponding time dimension information contained in the data to be detected are input into a pre-trained point cloud feature extraction network so that the point cloud feature extraction network outputs the corresponding point cloud features.
[0067] In this embodiment, the horizontal axis of the micro-Doppler signal represents time (s), and the vertical axis represents velocity V (m / s). The micro-Doppler signal is acquired at a rate of X frames per second, with T frames as a recognition segment. Each frame has C features, and the corresponding numerical values of the features reflect the intensity of the reflection of human motion speed. Therefore, the corresponding micro-Doppler image exhibits both distinct image and temporal features. Based on these two features, feature extraction from the micro-Doppler image can employ either a CNN network or an RNN network. In other examples, other neural network processing methods can also be used for feature extraction, without any particular limitation.
[0068] Furthermore, for point cloud data, each point has four defined parameters: x, y, z, and t. x, y, and z represent spatial coordinates, and t represents time. Since the number of point clouds in each frame is not a fixed value, the number of point clouds in the acquired recognition segment is also not a fixed value. This makes it impossible to directly use CNN networks for feature extraction of point cloud data like image data. This problem is commonly referred to as the irregular point cloud data feature extraction problem. Because human behavior data has significant characteristics in the time dimension, it is also necessary to preserve the temporal characteristics of point cloud information while performing regularization.
[0069] To consider the impact of temporal variations on features while achieving feature extraction from irregular point cloud data, this application also provides a PointTimeNet (PTN) temporal point cloud feature extraction network. For example... Figure 4 As shown, inputtransform is the spatiotemporal point cloud rotation module. T-Net is used to calculate the rotation matrix, and the calculation result is multiplied with the spatiotemporal point cloud to obtain the point cloud after spatiotemporal rotation. The purpose of this step is to transform the point cloud in the temporal and spatial dimensions, so as to rotate different point clouds to a uniform angle as much as possible. The function of MLP is to transform the dimension, transforming the point cloud features from 4 to C. Frame count is used to calculate the number of point clouds in each frame. Frame pooling uses the calculation result of frame count to pool the point cloud features of each frame to obtain frame features. The frame features are transformed by MLP to obtain the final k-classification result.
[0070] Based on the aforementioned time point cloud feature extraction network, its time extraction has been optimized as follows:
[0071] Optimization 1: Introduce temporal dimension information, change the input format of the point cloud from n*(x,y,z) to n*(x,y,z,t), and calculate the number of point clouds in each frame in the framecount layer of the PTN network using the introduced temporal information and retain this information.
[0072] Optimization 2: Point cloud rotation matrix optimization. PointNet's point cloud rotation matrix is a rotation performed on the spatial point cloud with a matrix parameter of 3*3. After introducing time, the PTN network performs spatiotemporal rotation with a spatiotemporal rotation matrix parameter of 4*4.
[0073] Optimization 3: Based on the point cloud count information obtained in Optimization 1, the global pooling operation of PointNet is changed to the frame pooling operation in PTN, and finally the features of each frame are obtained. The data format is T*C, which prepares for subsequent frame feature fusion.
[0074] Optimization 4: Since point cloud data is collected by millimeter-wave radar and static point cloud is removed using a static error elimination algorithm, it cannot be guaranteed that point cloud exists in every frame. This means that when a frame of data does not contain point cloud, the PTN network cannot extract features from the point cloud data. To solve this problem, a global zero-padding operation can be used, that is, a non-existent (0,0,0,t) point is added to each frame of point cloud, thereby ensuring the operation of the PTN network.
[0075] Optimization 5: Combining the point cloud data acquisition method of this application, the rotation matrix can be further optimized. Most of the acquired point cloud data is based on the horizontal ground, so there is no need to rotate around the x and y axes. Therefore, the rotation matrix is degenerated to rotate around the z axis. The size of the parameter matrix that needs to be determined for the rotation matrix is reduced from 3*3 to 2*2. Combined with optimization 2, only 5 unknown parameters need to be determined, including a 2*2 matrix and time scaling parameters.
[0076] In one embodiment of this application, the frame feature extraction network includes a connected recurrent neural network and a multilayer perceptron.
[0077] In this embodiment, to maximize the prediction accuracy, micro-Doppler features and point cloud features are fused to extract features from the two different types of data. The advantages of micro-Doppler in velocity detection are complemented by the advantages of point cloud data in time and space, thereby improving both the robustness and accuracy of the algorithm. Figure 5The diagram below illustrates the architecture of a frame feature fusion network according to an embodiment of this application. A CNN network is used to extract features from micro-Doppler data, and a PTN network is used to extract features from point cloud data. The feature data of both are extracted, and then the features are fused through feature concatenation. Since the data has temporal characteristics, an RNN (i.e., recurrent neural network) is used as a self-attention module to further extract features. Finally, the classification result is obtained through an MLP (i.e., multilayer perceptron).
[0078] Based on the technical solutions of the above embodiments, the following describes a specific application scenario of an embodiment of this application:
[0079] Figure 6 The diagram illustrates an application scenario where the technical solutions of the embodiments of this application can be applied, such as... Figure 6 As shown, in a bathroom setting, the millimeter-wave radar intelligent care terminal 202 can be installed in a location such as... Figure 6 The ceiling light fixture at location 201, near the center of the room at location 203, can also be installed in the center of the room, or around the room, or near areas of interest where behavioral monitoring is required, depending on the installation environment.
[0080] The millimeter-wave radar in the care system senses the environment by transmitting linear frequency-modulated electromagnetic waves with wavelengths between 1 and 10 mm. Upon encountering a transmissive obstacle, it eventually illuminates the target human body and transmits the electromagnetic wave. The reflected electromagnetic wave is received and mixed with the transmitted echo. Passing through a low-pass filter, an intermediate frequency (IF) signal containing reflections from the target, the human body, furniture, and walls is obtained. The radar data is then further processed by a data processing chip. Range-reflection intensity information is extracted from the IF signal using a range-dimensional FFT.
[0081] Due to the limited indoor space, most of the signals detected by radar are echoes reflected from static objects such as the ground or walls. After extracting the range information, static clutter removal is required to filter out static reflection noise caused by walls and the ground. Useful human motion information is then extracted. Next, a velocity-dimensional FFT is performed on the signal to obtain range-velocity information. Finally, the signals are summed along the range dimension to obtain the micro-Doppler signal. An angle-of-arrival (AOA) estimation algorithm (including MUSIC, CAPON, etc.) is applied to the signal from each antenna obtained by the range-dimensional FFT, or the CFAR constant false alarm rate algorithm is applied to the velocity-dimensional FFT result. An angle-dimensional FFT is then performed on the grid points where target velocity exists to obtain target point cloud data.
[0082] The millimeter-wave radar sends each frame of data (including micro-Doppler signals and target point cloud data) processed as described above back to the embedded device. After the millimeter-wave radar front-end completes its data transmission, it enters the next frame of electromagnetic wave transmission mode to repeat the above process.
[0083] Embedded devices can cache data and use micro-Doppler signals to identify human behavior. For human behavior, the action is a continuous process. Convolutional neural networks are used to identify user behavior through micro-Doppler. Therefore, this system needs to accurately locate the starting point of the action. If the user's action is not accurately located, it is easy to confuse the previous action with the next action, resulting in the network recognizing the action incorrectly.
[0084] In one embodiment, the action detection thread is activated when certain conditions are met. The action detection thread uses an action detection algorithm to detect the occurrence of actions and make accurate segmentations from n frames of micro-Doppler data within a certain time period in the receive buffer. In particular, in some instances, action detection and segmentation can use a masked convolution-peak detection algorithm, which will not be described in detail here.
[0085] Once an action is detected, the action recognition thread extracts the micro-Doppler data and point cloud data containing the user's action from the buffer and simultaneously normalizes the micro-Doppler data.
[0086] Specifically, the accurately cropped micro-Doppler spectrum is fed into a multi-domain feature fusion neural network for identification to determine the type of movement. If the movement is normal, it is classified as normal activity, and statistics are collected on the ward's lifestyle and health status. The intelligent care system reports to the system in a timely manner through the Internet of Things module and simultaneously returns to the monitoring step, allowing the system to continue operating normally. At the same time, the background server collects statistics on the user's lifestyle and provides remote health services.
[0087] In some embodiments, if there is a follow-up action and a subsequent response to a voice inquiry, it is determined to be a fall injury. If there is a follow-up action and the fall energy is too great, it is determined to be a serious fall injury. In some cases, if there is no response for a long time, it is determined that the monitored person has suffered a serious fall injury leading to coma.
[0088] Once the system confirms a fall, it will issue an alarm and notify a remote server and computer system via the network. The remote server system will then calculate the available medical and logistical resources to be deployed near the alarm location.
[0089] Specifically, the computing device described in this application is a device with a storage medium capable of processing data and performing operations on neural networks. It can also be implemented as an embedded system.
[0090] A network can be any communication network through which data can be transmitted and shared. For example, a network can be a local area network (LAN) or a wide area network (WAN), such as the Internet. Another example is the Internet of Things (IoT) or a cellular communication network. Networks can be implemented using various network interfaces, such as wireless network interfaces (e.g., Wi-Fi, Bluetooth, or infrared) or wired network interfaces (e.g., Ethernet or serial connections). A network can also comprise a combination of more than one network and can be implemented using one or more network interfaces.
[0091] The following describes an embodiment of the apparatus described in this application, which can be used to execute the behavior detection method based on millimeter-wave radar multi-domain data fusion described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the behavior detection method based on millimeter-wave radar multi-domain data fusion described above in this application.
[0092] Figure 7 A block diagram of a behavior detection device based on millimeter-wave radar multi-domain data fusion according to an embodiment of this application is shown.
[0093] Reference Figure 7 As shown, a behavior detection device based on millimeter-wave radar multi-domain data fusion according to an embodiment of this application includes:
[0094] The first determining module 710 is used to determine the corresponding range-velocity spectrum based on the echo signal obtained by the millimeter-wave radar from the space to be detected.
[0095] The second determining module 720 is used to add up each identical velocity grid point at each distance according to the distance-velocity spectrum to obtain a single frame of micro-Doppler signal;
[0096] The third determining module 730 is used to determine the point cloud data corresponding to the space to be detected based on the range-velocity spectrum corresponding to multiple antennas, wherein the point cloud data corresponds to the micro-Doppler signal in time;
[0097] The acquisition module 740 is used to determine whether an action has occurred based on the micro-Doppler signal. If the action is determined to have occurred, the module acquires the detection data corresponding to the action. The detection data includes several frames of micro-Doppler signals and their corresponding point cloud data.
[0098] The extraction module 750 is used to extract features from several frames of micro-Doppler signals and their corresponding point cloud data contained in the data to be detected, and to determine their corresponding micro-Doppler features and point cloud features.
[0099] The processing module 760 is used to input the micro-Doppler features and the point cloud features into a pre-trained frame feature fusion network so that the frame feature fusion network outputs the corresponding action classification result.
[0100] In one embodiment of this application, the acquisition module 740 is configured to: when a predetermined condition is met, acquire multiple frames of micro-Doppler signals corresponding to a predetermined time length; determine the position with the highest motion correlation in the multiple frames of micro-Doppler signals using a mask convolution algorithm; and segment the micro-Doppler signals into those with velocity energy grids and those without velocity energy grids; use a motion detection algorithm to calculate the motion correlation of the segmented micro-Doppler signals to determine whether any motion has occurred; if motion is determined to have occurred, copy and acquire the micro-Doppler signals and point cloud data within a predetermined range before and after the motion correlation peak point as the data to be detected.
[0101] In one embodiment of this application, the extraction module 750 is used to: extract features from several frames of micro-Doppler signals contained in the data to be detected using a pre-trained CNN network or RNN network; and input several frames of point cloud data and their corresponding time dimension information contained in the data to be detected into a pre-trained point cloud feature extraction network so that the point cloud feature extraction network outputs the corresponding point cloud features.
[0102] In one embodiment of this application, the frame feature extraction network includes a connected recurrent neural network and a multilayer perceptron.
[0103] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0104] It should be noted that, Figure 8 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0105] like Figure 8 As shown, the computer system includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 802 or programs loaded from storage portion 808 into Random Access Memory (RAM) 803, such as performing the methods described in the above embodiments. The RAM 803 also stores various programs and data required for system operation. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.
[0106] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.
[0107] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit (CPU) 801, it performs various functions defined in the system of this application.
[0108] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0110] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0111] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.
[0112] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0113] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.
[0114] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0115] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A behavior detection method based on multi-domain data fusion of millimeter wave radar, characterized in that, The method comprises the following steps: determining a distance-velocity spectrum corresponding to the echo signal obtained by the millimeter wave radar detecting the space to be detected; adding each same velocity grid point on each distance according to the distance-velocity spectrum to obtain a single frame of micro-Doppler signal; determining point cloud data corresponding to the space to be detected based on the distance-velocity spectrum corresponding to multiple antennas, wherein the point cloud data corresponds to the micro-Doppler signal in time; determining whether an action occurs according to the micro-Doppler signal, and if it is determined that an action occurs, obtaining detection data corresponding to the action, wherein the detection data comprises a plurality of frames of micro-Doppler signal and corresponding point cloud data; extracting features of the plurality of frames of micro-Doppler signal and corresponding point cloud data in the detection data respectively to determine corresponding micro-Doppler features and point cloud features; inputting the micro-Doppler features and the point cloud features into a pre-trained frame feature fusion network to enable the frame feature fusion network to output corresponding action classification results; wherein the step of determining whether an action occurs according to the micro-Doppler signal, and if it is determined that an action occurs, obtaining detection data corresponding to the action comprises: when a predetermined condition is met, obtaining a plurality of frames of micro-Doppler signal corresponding to a predetermined time length, determining a position with the highest action correlation in the plurality of frames of micro-Doppler signal by a mask convolution algorithm, and cutting the position into a position with velocity energy grid points and a position without velocity energy grid points; using an action detection algorithm to calculate the action correlation of the cut micro-Doppler signal to determine whether an action occurs; if it is determined that an action occurs, copying the micro-Doppler signal and the point cloud data in a predetermined range before and after the correlation peak point of the action as detection data.
2. The method of claim 1, wherein, The step of extracting features of the plurality of frames of micro-Doppler signal and corresponding point cloud data in the detection data respectively to determine corresponding micro-Doppler features and point cloud features comprises: using a pre-trained CNN network or RNN network to extract features of the plurality of frames of micro-Doppler signal in the detection data; inputting the plurality of frames of point cloud data and corresponding time dimension information in the detection data into a pre-trained point cloud feature extraction network to enable the point cloud feature extraction network to output corresponding point cloud features.
3. The method according to any one of claims 1-2, characterized in that, The frame feature fusion network comprises a recurrent neural network and a multilayer perceptron connected in series.
4. A behavior detection apparatus based on multi-domain data fusion of millimeter wave radar, characterized by, The method comprises the following steps: a first determining module configured to determine a distance-velocity spectrum corresponding to the echo signal obtained by the millimeter wave radar detecting the space to be detected; a second determining module configured to add each same velocity grid point on each distance according to the distance-velocity spectrum to obtain a single frame of micro-Doppler signal; a third determining module configured to determine point cloud data corresponding to the space to be detected based on the distance-velocity spectrum corresponding to multiple antennas, wherein the point cloud data corresponds to the micro-Doppler signal in time; The acquisition module is configured to determine whether an action occurs according to the micro-Doppler signal, and if it is determined that the action occurs, acquire detection data corresponding to the action, wherein the detection data includes a plurality of frames of micro-Doppler signals and corresponding point cloud data; The extraction module is configured to respectively extract features from the plurality of frames of micro-Doppler signals and the corresponding point cloud data included in the detection data, and determine corresponding micro-Doppler features and point cloud features; The processing module is configured to input the micro-Doppler features and the point cloud features into a pre-trained frame feature fusion network, so that the frame feature fusion network outputs a corresponding action classification result. The acquisition module is configured to: When a predetermined condition is met, acquire a plurality of frames of micro-Doppler signals corresponding to a predetermined time length, determine a position with the highest action correlation in the plurality of frames of micro-Doppler signals by using a mask convolution algorithm, and cut the position into a position with velocity energy grid points and a position without velocity energy grid points; An action detection algorithm is used to detect the cut micro-Doppler signals to determine whether an action occurs; If it is determined that the action occurs, the micro-Doppler signals and the point cloud data in a predetermined range before and after a correlation peak point of the action are copied and acquired as detection data.
5. The apparatus of claim 4, wherein, The extraction module is configured to: For the plurality of frames of micro-Doppler signals included in the detection data, a pre-trained CNN network or RNN network is used to extract features from the plurality of frames of micro-Doppler signals; The plurality of frames of point cloud data and corresponding time dimension information included in the detection data are input into a pre-trained point cloud feature extraction network, so that the point cloud feature extraction network outputs corresponding point cloud features.
6. The apparatus of any one of claims 4-5, wherein, The frame feature fusion network includes a recurrent neural network and a multilayer perceptron connected in series.
7. A computer readable medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the behavior detection method based on millimeter wave radar multi-domain data fusion according to any one of claims 1 to 3.
8. An electronic device, comprising: The computer program, when executed by a processor, implements the behavior detection method based on millimeter wave radar multi-domain data fusion according to any one of claims 1 to 3. The computer program, when executed by a processor, implements the behavior detection method based on millimeter wave radar multi-domain data fusion according to any one of claims 1 to 3. The computer program, when executed by a processor, implements the behavior detection method based on millimeter wave radar multi-domain data fusion according to any one of claims 1 to 3.
Citation Information
Patent Citations
SYSTEM AND METHOD FOR HUMAN BEHAVIOR MODELLING AND POWER CONTROL and storage medium
CN110068815A
Millimeter wave radar human body action real-time detection method based on neural network
CN114429672A