Space-time fusion human body posture recognition system and method based on millimeter wave radar
Through the space-time fusion human posture recognition system based on millimeter wave radar, the problem of human posture recognition accuracy and privacy protection in occlusion and lighting changes is solved, and high-precision and real-time recognition effects are achieved.
Patent Information
- Application Number
- CN202510671758.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
AI Technical Summary
The existing human posture recognition technology has insufficient accuracy and privacy protection in occlusion and lighting changes, and the calculation volume is large, making it difficult to achieve real-time and high-precision recognition.
The space-time fusion human posture recognition system based on millimeter wave radar is adopted. The human body distance and angle information are extracted from the frequency-modulated continuous wave signal through the feature extraction module, and the human body posture point cloud is constructed, and the space-time feature fusion module is used for identification. Combining the time domain and spatial domain features, multipath interference is reduced and recognition accuracy is improved.
It realizes the accuracy and robustness of human posture recognition in complex environments, reduces the impact of multipath interference, and improves the security and real-time recognition.
Smart Images

Figure CN120491055A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless intelligent perception of the Internet of Things, and specifically relates to a time-space fusion human posture recognition system and method based on millimeter-wave radar. Background Art
[0002] With the rapid development of AI-assisted sensing systems, intelligent perception has been widely applied in various scenarios, addressing numerous application needs. In particular, human posture recognition in various scenarios is being applied in applications ranging from improving the efficiency of intelligent driving in cars and enabling smart home interactions to intelligent medical monitoring. For example, in intelligent driving, human posture can provide traffic systems and autonomous vehicles with information about the status of people, predict their upcoming actions, and enable preventive decision-making. Furthermore, in whole-home intelligent human-computer interaction, human posture can enable devices to start, stop, and interact without touching a switch. In intelligent medical monitoring, human posture can effectively monitor patient conditions and perform gait analysis, saving medical personnel. However, existing human posture estimation methods are mostly based on visual modalities. In many scenarios, they are prone to privacy risks, are significantly affected by lighting, and are easily obscured by obstacles, resulting in loss of target data. Furthermore, the computational effort and processing speed required by visual models also hinder the real-time performance of posture estimation. Therefore, other sensing modalities are needed to collect human posture information, such as human posture recognition in homes and protected areas.
[0003] The sensing system has multiple modes to solve the problem of human posture recognition in non-open and occluded scenes:
[0004] 1. Wi-Fi-based human posture recognition: Multiple Wi-Fi transmitting and receiving antennas are used to collect human posture information. Given the wide coverage of Wi-Fi, using commercial Wi-Fi devices for human posture estimation can effectively overcome interference caused by occlusion without collecting private data. However, perception using Wi-Fi modules often suffers from poor accuracy, is highly sensitive to the surrounding environment, and cannot detect previously unseen free movements.
[0005] 2. Human pose recognition based on laser modalities: Using multiple sets of laser beam transmitting and receiving antennas, objects in the scene are depicted using point clouds. Using LiDAR point clouds to estimate human pose allows for accurate target signal perception while protecting privacy. However, LiDAR has difficulty detecting objects obscured by obstacles, and because the resulting point clouds are so numerous, additional computation is required to merge and segment the corresponding objects. Furthermore, the sensors are very expensive, making them difficult to scale.
[0006] 3. Human gesture recognition based on millimeter-wave modality: Using multiple transmitting and receiving antennas of millimeter-wave radar, target information is acquired through point clouds or intermediate-frequency signals. Detection algorithms are then used to filter and fit the target information, completing human gesture recognition. Because millimeter-wave radar has a shorter wavelength than Wi-Fi, it can obtain more accurate location information than Wi-Fi. Because its frequency is lower than that of laser modality, it experiences less energy loss in the medium, allowing it to penetrate clouds, rain, and fog to acquire location information regardless of weather conditions. However, millimeter waves are also subject to the influence of wavelength and frequency. The short wavelength increases the likelihood of specular reflections from the medium's surface, while the low frequency results in less energy loss in the medium. These two factors together result in a more severe multipath effect for millimeter waves, severely impacting accurate target recognition.
[0007] Therefore, based on the above considerations, it is necessary to propose a human posture recognition system that can solve the problems of privacy leakage, susceptibility to light and occlusion. It utilizes the signal characteristics of both time domain and space domain, takes into account the continuity in time domain and consistency in space domain, reduces the recognition error caused by multipath interference, and improves the accuracy, robustness and security of human posture recognition. Summary of the Invention
[0008] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a millimeter-wave radar-based spatiotemporal fusion human posture recognition system and method to solve the problem of multipath interference that is prone to exist in wireless sensing modes in the prior art.
[0009] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0010] The present invention provides a millimeter wave radar-based time-space fusion human posture recognition system, comprising: a feature extraction module and a time-space feature fusion and posture recognition module;
[0011] The feature extraction module is used to extract the distance and angle information of the human body from the frequency modulated continuous wave signal after scanning the target human body, and construct a human body posture point cloud;
[0012] The spatiotemporal feature fusion and posture recognition module is used to extract the time domain and spatial domain features of the human body posture point cloud signal and perform human body posture recognition through feature fusion.
[0013] Furthermore, the feature extraction module uses the millimeter wave radar to transmit frequency modulated continuous wave as the transmission signal. The signal characteristics are: each group consists of M frames of Chirp signal, the corresponding signal period is T, and the corresponding initial frequency is f c ; Receive the reflected signal after scanning the target human body, the interval between the transmitted signal and the received signal is t dAnd the corresponding reflected signal of each group of Chirp signals is mixed with the transmitted signal to obtain the demodulated intermediate frequency signal. The function expression of the transmitted signal TX(t) and the received signal RX(t) with respect to time t is:
[0014]
[0015] Where e is the natural logarithm, j is the imaginary unit, and B represents the signal bandwidth corresponding to each group of Chirp signals. The corresponding intermediate frequency signal IF(t) is expressed as a function of time t:
[0016]
[0017] Furthermore, the expression of the intermediate frequency signal phase φ composed of the irrelevant terms at time t in the intermediate frequency signal IF(t) is:
[0018]
[0019] By measuring the phase φ of the mixed intermediate frequency signal, the interval time t between the transmitted signal and the received signal can be determined. d .
[0020] Furthermore, the extracting of the distance information of the human body specifically includes:
[0021] A single-frame Chirp signal is extracted from the intermediate frequency signal, and a discrete point Fourier transform algorithm is performed on each Chirp signal. The Chirp signal corresponding to the target person is determined by detecting the peak position, and then the interval time t between the transmitted signal and the received signal is determined based on the intermediate frequency signal phase φ after demodulation of the transmitted signal and the received signal. d ;
[0022] The time interval between the transmitted signal and the received signal is t d Extract the distance d, t between the millimeter wave antenna and the target human body d The ratio of the distance r from the signal transmission to the reception to the electromagnetic wave propagation speed c is used to calculate the distance d=ct between the millimeter wave antenna and the target human body. d / 2.
[0023] Furthermore, the extracting of angle information of the human body specifically includes:
[0024] The phase difference Δφ of the intermediate frequency signal under multiple sets of receiving antennas on the millimeter wave radar describes the distance difference Δd between the millimeter wave antenna and the target human body. The corresponding expression is:
[0025]
[0026] Where λ is the wavelength of the millimeter-wave radar signal;
[0027] The distance difference Δd between the millimeter wave antenna and the target human body represents the product of the distance l between each millimeter wave receiving antenna and the sine value of the angle θ between the millimeter wave antenna and the target human body, that is, Δd = lsin(θ). The expression for the angle between the millimeter wave antenna and the target human body is:
[0028]
[0029] Furthermore, constructing the human body posture point cloud specifically includes:
[0030] The constant false alarm rate algorithm is used to filter the reflection intensity of the intermediate frequency signal. By setting the adjustable detection threshold Th, the local peak point with large signal intensity is selected to control the target point detection sensitivity. The expression of the detection threshold Th is:
[0031] Th=αP n
[0032] Among them, α is the threshold coefficient, P n is the noise intensity estimate, and the expression is:
[0033]
[0034] Among them, P fa is the false alarm rate, N is the number of samples, x m is the sample value; the local peak point of the acquired millimeter wave signal intensity is converted, and the distance and angle of the peak point in the polar coordinate system are converted into point cloud points of the horizontal, vertical and depth coordinates of the physical space, thereby constructing the human body posture point cloud.
[0035] Furthermore, the spatiotemporal feature fusion and posture recognition module uses a coarse-grained human posture recognition algorithm to convert the constructed human posture point cloud into coarse-grained human posture skeleton points; uses the consistency of human posture changes in consecutive frames in the time domain, and uses a merging algorithm of continuous frame skeleton point sets to make consistency judgments on the human postures of consecutive frames, thereby reducing the large deviation of posture skeleton points caused by multipath and ghost points, and merging them into the correct posture range of the current frame; uses the continuity of posture position in the spatial domain, and uses a discrete point heat map generation algorithm to convert the discrete posture skeleton points obtained in the time domain into a continuous spatial domain confidence mask, thereby fusing the confidence with the human posture point cloud; uses a fine-grained human posture recognition algorithm to fuse the confidence mask with the human posture point cloud to obtain fine-grained human posture information.
[0036] Furthermore, the coarse-grained human posture recognition algorithm is used to obtain coarse-grained human posture skeleton points, and the specific steps are as follows:
[0037] (21) Vectorized human body posture point cloud: The physical space position of the target human body is divided into positions, and the minimum spatial resolution of the millimeter wave is used as a reference. The measurable space where the target human body is located is used as the range boundary. A spatial voxel dictionary of the horizontal, vertical and depth dimensions is created, and the corresponding point cloud is mapped into the spatial voxel dictionary, thereby obtaining n groups of voxel vectors consisting of n consecutive frames of point cloud;
[0038] (22) Extract voxel vector features: The encoder structure of the temporal Transformer is used to extract features from the n sets of voxel vectors output in step (21). Specifically, the feature dimension is expanded through the embedding layer (·), and the context relevance of the features is extracted through the encoder (·). Each encoder (·) contains two sets of tensors: the intermediate unit state and the final unit state. The state tensor transmission is controlled by the memory gate structure.
[0039] (23) Calculating feature correlation: Using the decoder structure of the time-series Transformer, the features output in step (22) are correlated. Specifically, the correlation matrix of each set of features is calculated through the attention layer Attention(·), and the features are fused with the correlation matrix through the decoder Decoder(·) to merge and extract features, and the features are passed back.
[0040] (24) Predicting coarse-grained human posture skeleton points: The features output in step (23) are mapped to the voxel vectors using the fully connected layer structure of the temporal Transformer to obtain the human posture skeleton points. Specifically, the vector sequence is integrated and dimensionally transformed through the fully connected layer FC(·), and the highest possible voxel classification is found by the normalization function Softmax(·) activation function. Finally, the voxels are restored to physical spatial positions through the spatial voxel dictionary to obtain the position prediction results of the coarse-grained human posture skeleton points.
[0041] Furthermore, the correct posture range of the current frame is obtained by the merging algorithm of the skeleton point sets of consecutive frames, and the specific steps are as follows:
[0042] (31) Extracting continuous frame posture skeleton points: Select the current frame and its previous m-1 frames to form m groups of continuous human posture estimation vectors, and extract the coarse-grained human posture skeleton points with the same label in the corresponding time series in chronological order. The coarse-grained human posture skeleton points are the 25 skeleton points that constitute the human posture frame, and form a continuous frame posture skeleton point vector of [m×25];
[0043] (32) Calculate the distance between the continuous frame posture skeleton points: Calculate the distance of the 25 skeleton points corresponding to the continuous frame posture skeleton point vector output in step (31), and obtain the Manhattan distance matrix M between the skeleton point i and the skeleton point j by calculating the Manhattan distance Manh() of the corresponding label skeleton points in the continuous frame. ij ;
[0044] (33) Find the reference point of the skeleton point in the continuous frame posture: half the Manhattan distance of the skeleton points in different frames with the same label output in step (32) is summed, that is, the sum of the (m-1) / 2 items with the smallest Manhattan distance is calculated; the sum of the (m-1) / 2 items with the smallest Manhattan distance corresponding to the skeleton point i and the skeleton points of the continuous frames with the same label is recorded as D i ;
[0045] The reference point of the continuous frame pose skeleton point is to find the sum of (m-1) / 2 items with the smallest Manhattan distance from each group of skeleton points with the same label. i , the smallest D i The corresponding bone point P i As the reference point of the label skeleton point, the reference point and the skeleton point corresponding to the smallest (m-1) / 2 Manhattan distance are recorded as the reference point Rf j , merged and added to the reference point set Refer{P i ,Rf1,…,Rf (m-1) / 2}middle;
[0046] (34) Find the center point of the continuous frame posture skeleton point: The reference point set output in step (33) is calculated by calculating the arithmetic average of the corresponding coordinates of the reference points under the corresponding labels to obtain the position coordinates P of the geometric center point. center , the expression is:
[0047]
[0048] By finding the center point P of 25 skeleton points center , merge m groups of consecutive frames, each with 25 skeleton points, to obtain the final center point set matrix Pose of the human body posture skeleton points center .
[0049] Furthermore, the discrete point heat map generation algorithm is used to obtain a continuous spatial domain confidence mask, thereby obtaining a human body posture point cloud and a confidence fusion vector. The specific steps are as follows:
[0050] (41) Draw a discrete point heat map: The center point matrix Pose of the human body posture skeleton points obtained in step (34) centerConvert it into a vector with four dimensions of H, W, D, and C, which is a discrete point heat map. H, W, and D represent the vertical height, horizontal width, and depth range of the discrete point heat map, respectively. C represents the confidence weight. The confidence weight of the point on the discrete point heat map corresponding to the center point is set to 1, and the rest are set to 0.
[0051] (42) Expanding the confidence heat map: The center point obtained in step (34) is quantitatively expanded by the Gaussian expansion algorithm (·) to form a confidence heat map. The heat map of concentric circles is expanded outward with the center point as the center. The corresponding confidence value C is expressed as:
[0052]
[0053] Among them, h i , w i and d i are the three-dimensional (H, W, D) coordinates corresponding to the 25 center points, and N is expressed as:
[0054]
[0055] Among them, σ h , σ w , σ d They represent the distribution parameters of the heat map along H, W, and D respectively;
[0056] (43) Human body posture point cloud and confidence fusion: The confidence heat map obtained in step (42) is used as a continuous spatial domain confidence mask and fused with the human body posture point cloud; the skeleton point P in the human body posture point cloud i , whose corresponding vertical, horizontal and depth coordinates are h i 、w i and d i , bring the corresponding coordinates into the confidence heat map in step (42) to find the skeleton point P i The corresponding confidence value C(P i ), and then combine the confidence value with the three-dimensional coordinates of the skeleton point to obtain the point cloud fusion vector P i (h,w,d,c).
[0057] Furthermore, the method of obtaining fine-grained human posture information by using a fine-grained human posture fitting algorithm specifically includes:
[0058] Multi-dimensional feature fusion: The coordinate coord and confidence conf in the point cloud fusion vector are fused as multi-dimensional inputs, and the point cloud position feature encoder is used to encode coord into Encoder(coord), and the point cloud confidence feature encoder is used to encode conf into Encoder(conf); then the encoded point cloud position feature Encoder(coord) and point cloud confidence feature Encoder(conf) are fused using the TransFuser structure to obtain a fine-grained human posture vector F Pose , expressed as:
[0059] F Pose =TransFuser[Encoder(coord),Encoder(conf)]
[0060] Among them, the feature fusion structure TransFuser consists of an encoder, a decoder, and a multi-dimensional input feature attention module Attention(·), which can be expressed as:
[0061]
[0062] Among them, Q represents the query matrix, K represents the key-value matrix, and V represents the value matrix. is the scale factor; the human body posture is fitted through a set of linear layers Linear(·), as follows:
[0063] (51) Input the point cloud fusion vector in step (43) into the position feature encoder and the confidence feature encoder respectively, including the point cloud position feature coord and the point cloud confidence feature conf;
[0064] (52) Point cloud position feature encoding: The coordinate feature of the point cloud fusion vector is The three-dimensional coordinates are converted into a single-dimensional morpheme vector through the spatial coordinate mapping voxel (·) Divide into two-dimensional plane slices according to the time dimension The corresponding attention features are Q coord , K coord and V coord ;
[0065] (53) Point cloud confidence feature encoding: The confidence feature of the point cloud fusion vector is conf = (conf1, conf2, ..., conf n ),in The corresponding attention features are Q conf , K conf and V conf ;
[0066] (54) Multi-head attention extraction: position attention feature Q of point cloud fusion vector coord , K coord , V coord And the confidence attention feature Q of the point cloud fusion vector conf , K conf , V conf , the corresponding features are exchanged using the multi-head attention mechanism, and the two sets of features after the exchange are A coord =att((K coord ,V coord ),Q conf ) and A conf =att((K conf ,V conf ),Q coord );
[0067] (55) Joint attention mechanism: The features extracted by multi-head attention are used as the input of the joint feature encoder, and the joint attention mechanism is used to fuse the position and confidence of the human posture point cloud into high-dimensional features. The corresponding feature representation is: A coord =Concat(A coord ,A conf );
[0068] Decoder(·) is used to decode the fine-grained human posture vector F Pose Decode to get the corresponding position encoding feature F coord , denoted as F coord =Decoder(F Pose ), and use the fully connected layer to map the features into the physical space to obtain the coordinates of the corresponding human posture points.
[0069] The present invention provides a method for human posture recognition based on time-space fusion of millimeter-wave radar, based on the above system, and the steps are as follows:
[0070] 1) Collect the frequency modulated continuous wave signal emitted by the millimeter wave radar after scanning the target human body;
[0071] 2) Demodulate the collected signal to obtain an intermediate frequency signal;
[0072] 3) Extract the distance and angle information between the millimeter wave antenna and the target human body from the intermediate frequency signal;
[0073] 4) Using the constant false alarm rate (CFAR) algorithm to filter the intermediate frequency signal, and converting the polar coordinates into spatial coordinates to construct the human body posture point cloud;
[0074] 5) Perform coarse-grained human posture skeleton point recognition;
[0075] 6) Using the continuous frame pose skeleton points, calculate the sum of the (m-1) / 2 items with the smallest Manhattan distance D i , D i The smallest bone point is used as the reference point, and a reference point set of 25 bone points is constructed, and the center point matrix of the human body posture bone points is merged;
[0076] 7) Using the discrete point heat map and Gaussian expansion algorithm, the continuous spatial domain confidence mask is calculated and fused with the human body posture point cloud to obtain the point cloud fusion vector;
[0077] 8) Construct a point cloud position feature and confidence feature encoder to process the point cloud fusion vector, use the joint attention mechanism to fuse the position and confidence features, and use the decoder to identify the coordinates of the human body posture points.
[0078] Beneficial effects of the present invention:
[0079] 1. The present invention uses a continuous frame reference point extraction algorithm to extract more reliable and stable millimeter wave point cloud signals, thereby reducing the impact of ghost points caused by multipath effects;
[0080] 2. The present invention realizes the confidence calculation of discrete points by adopting the confidence heat map and Gaussian expansion algorithm, so that confidence mapping can be performed on subsequent frame point clouds;
[0081] 3. By fusing two dimensional features, point cloud coordinates and confidence, the present invention achieves the complementarity and enhancement of multi-dimensional features, thereby further improving the accuracy of human posture recognition;
[0082] 4. The present invention uses a single sensor to simultaneously acquire two-dimensional features and fuse them, thereby achieving effective recognition of human posture under multipath interference in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 Schematic diagram of the system of the present invention.
[0084] Figure 2 Flowchart of coarse-grained pose recognition.
[0085] Figure 3 Schematic diagram for merging skeleton points in consecutive frames.
[0086] Figure 4 Schematic diagram of point cloud confidence fusion.
[0087] Figure 5 Schematic diagram of multi-dimensional feature fusion and recognition.
[0088] Figure 6 This is a demonstration diagram of the multipath effect. DETAILED DESCRIPTION
[0089] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and drawings. The contents mentioned in the embodiments are not intended to limit the present invention.
[0090] Reference Figures 1 to 5 As shown, a time-space fusion human posture recognition system based on millimeter wave radar of the present invention includes: a feature extraction module and a time-space feature fusion and posture recognition module;
[0091] The feature extraction module is used to extract the distance and angle information of the human body from the frequency modulated continuous wave (FMCW) signal after scanning the target human body, and construct a human body posture point cloud;
[0092] The feature extraction module uses the millimeter wave radar to transmit frequency modulated continuous wave as the transmission signal. The signal characteristics are: each group consists of M frames of Chirp signal, the corresponding signal period is T, and the corresponding initial frequency is f c ; Receive the reflected signal after scanning the target human body, the interval between the transmitted signal and the received signal is t d And the corresponding reflected signal of each group of Chirp signals is mixed with the transmitted signal to obtain the demodulated intermediate frequency signal. The function expression of the transmitted signal TX(t) and the received signal RX(t) with respect to time t is:
[0093]
[0094] Where e is the natural logarithm, j is the imaginary unit, and B represents the signal bandwidth corresponding to each group of Chirp signals. The corresponding intermediate frequency signal IF(t) is expressed as a function of time t:
[0095]
[0096] The expression of the intermediate frequency signal phase φ composed of the irrelevant terms at time t in the intermediate frequency signal IF(t) is:
[0097]
[0098] By measuring the phase φ of the mixed intermediate frequency signal, the interval time t between the transmitted signal and the received signal can be determined. d .
[0099] The step of extracting the distance information of the human body specifically includes:
[0100] A single-frame Chirp signal is extracted from the intermediate frequency signal, and a discrete point Fourier transform algorithm is performed on each Chirp signal. The Chirp signal corresponding to the target person is determined by detecting the peak position, and then the interval time t between the transmitted signal and the received signal is determined based on the intermediate frequency signal phase φ after demodulation of the transmitted signal and the received signal. d ;
[0101] The time interval between the transmitted signal and the received signal is t d Extract the distance d, t between the millimeter wave antenna and the target human body d The ratio of the distance r from the signal transmission to the reception to the electromagnetic wave propagation speed c is used to calculate the distance d=ct between the millimeter wave antenna and the target human body. d / 2.
[0102] The step of extracting the angle information of the human body specifically includes:
[0103] The phase difference Δφ of the intermediate frequency signal under multiple sets of receiving antennas on the millimeter wave radar describes the distance difference Δd between the millimeter wave antenna and the target human body. The corresponding expression is:
[0104]
[0105] Where λ is the wavelength of the millimeter-wave radar signal;
[0106] The distance difference Δd between the millimeter wave antenna and the target human body represents the product of the distance l between each millimeter wave receiving antenna and the sine value of the angle θ between the millimeter wave antenna and the target human body, that is, Δd = lsin(θ). The expression for the angle between the millimeter wave antenna and the target human body is:
[0107]
[0108] The step of constructing a human body posture point cloud specifically includes:
[0109] The constant false alarm rate (CFAR) algorithm is used to filter the reflection intensity of the intermediate frequency signal. By setting the adjustable detection threshold Th, the local peak point with large signal intensity is selected to control the target point detection sensitivity. The expression of the detection threshold Th is:
[0110] Th=αP n
[0111] Among them, α is the threshold coefficient, P n is the noise intensity estimate, and the expression is:
[0112]
[0113] Among them, P fa is the false alarm rate, N is the number of samples, x m is the sample value; the local peak point of the acquired millimeter wave signal intensity is converted, and the distance and angle of the peak point in the polar coordinate system are converted into point cloud points of the horizontal, vertical and depth coordinates of the physical space, thereby constructing the human body posture point cloud.
[0114] The spatiotemporal feature fusion and posture recognition module is used to extract the time domain and spatial domain features of the human body posture point cloud signal and perform human body posture recognition through feature fusion.
[0115] Among them, the spatiotemporal feature fusion and posture recognition module uses a coarse-grained human posture recognition algorithm to convert the constructed human posture point cloud into coarse-grained human posture skeleton points; uses the consistency of human posture changes in consecutive frames in the time domain, and uses a merging algorithm of continuous frame skeleton point sets to make consistency judgments on the human postures of consecutive frames, thereby reducing the large deviation of posture skeleton points caused by multipath and ghost points, and merging them into the correct posture range of the current frame; uses the continuity of posture position in the spatial domain, and uses a discrete point heat map generation algorithm to convert the discrete posture skeleton points obtained in the time domain into a continuous spatial domain confidence mask, thereby fusing the confidence with the human posture point cloud; uses a fine-grained human posture recognition algorithm to fuse the confidence mask with the human posture point cloud to obtain fine-grained human posture information.
[0116] In this context, multipath refers to the fact that wireless signals may travel along multiple paths, not just straight ones, from the transmitter to the receiver. Because signals can propagate through reflection, refraction, and scattering, the time, phase, and intensity of signals from different paths may vary, resulting in multiple sets of reflected signals from the same target. When these reflected signals are processed into a point cloud, points originating from direct paths are called true target points, while points originating from indirect paths are called false target points, also known as ghost points.
[0117] Among them, in order to verify the interference of multipath effect and ghost points on the construction of the above human posture point cloud, the direct path (LOS) and multiple reflection path (NLOS) signals were constructed, and the distance and azimuth of the ghost points were calculated. Figure 6 As shown, the details are as follows:
[0118] (11) Construction of direct path and multiple reflection path: The direct path corresponds to the millimeter wave radar antenna directly reaching the target point of the human posture through the optical path, and the multiple reflection path corresponds to reaching the target point of the human posture after being reflected from the ground or the wall; set the height of the millimeter wave radar from the ground h1, the height of the human posture target point from the ground h2, the elevation angle α of the transmitted signal through the multiple reflection path, the path r1 from the millimeter wave antenna to the ground reflection point, the path r2 and reflection angle β from the ground reflection point to the human posture target point, the path r3 from the human posture target point to the millimeter wave antenna, and the angle between the reflection surface and the vertical plane corresponding to the multiple reflection path at the human posture target point is γ;
[0119] (12) Calculate the ghost point distance r g and azimuth angle θ g: The ghost point g is on the reverse extension line of the path from the human body posture target point to the millimeter wave antenna, and has the same azimuth angle as the real point r, that is, θ g =θ r =2β-α; using the relevant angle and distance information in step (11), eliminating the linearly related variables, and obtaining the three segments of the multiple reflection path r1, r2, r3 and r g They are:
[0120]
[0121] The distance difference Δd between the corresponding human posture target point and the ghost point is:
[0122]
[0123] The coarse-grained human posture recognition algorithm is used to obtain coarse-grained human posture skeleton points. The specific steps are as follows:
[0124] (21) Vectorized human body posture point cloud: The physical space position of the target human body is divided into positions, with the minimum spatial resolution of millimeter wave (5cm) as the reference and the measurable space (5m×5m×5m) where the target human body is located as the range boundary. A spatial voxel dictionary of the horizontal, vertical and depth dimensions is created, and the corresponding point cloud is mapped into the spatial voxel dictionary, thereby obtaining n groups of voxel vectors consisting of n consecutive frames of point cloud;
[0125] (22) Extract voxel vector features: The encoder structure of the temporal Transformer is used to extract features from the n sets of voxel vectors output in step (21). Specifically, the feature dimension is expanded through the embedding layer (·), and the context relevance of the features is extracted through the encoder (·). Each encoder (·) contains two sets of tensors: the intermediate unit state and the final unit state. The state tensor transmission is controlled by the memory gate structure.
[0126] (23) Calculating feature correlation: Using the decoder structure of the time-series Transformer, the features output in step (22) are correlated. Specifically, the correlation matrix of each set of features is calculated through the attention layer Attention(·), and the features are fused with the correlation matrix through the decoder Decoder(·) to merge and extract features, and the features are passed back.
[0127] (24) Predicting coarse-grained human posture skeleton points: The features output in step (23) are mapped to the voxel vectors using the fully connected layer structure of the temporal Transformer to obtain the human posture skeleton points. Specifically, the vector sequence is integrated and dimensionally transformed through the fully connected layer FC(·), and the highest possible voxel classification is found by the normalization function Softmax(·) activation function. Finally, the voxels are restored to physical spatial positions through the spatial voxel dictionary to obtain the position prediction results of the coarse-grained human posture skeleton points.
[0128] The correct posture range of the current frame is obtained by merging the skeleton point sets of consecutive frames. The specific steps are as follows:
[0129] (31) Extracting continuous frame posture skeleton points: Select the current frame and its previous m-1 frames to form m groups of continuous human posture estimation vectors, and extract the coarse-grained human posture skeleton points with the same label in the corresponding time series in chronological order. The coarse-grained human posture skeleton points are the 25 skeleton points that constitute the human posture frame, and form a continuous frame posture skeleton point vector of [m×25];
[0130] (32) Calculate the distance between the continuous frame posture skeleton points: Calculate the distance of the 25 skeleton points corresponding to the continuous frame posture skeleton point vector output in step (31), and obtain the Manhattan distance matrix M between the skeleton point i and the skeleton point j by calculating the Manhattan distance Manh() of the corresponding label skeleton points in the continuous frame. ij ;
[0131] (33) Find the reference point of the skeleton point in the continuous frame posture: half the Manhattan distance of the skeleton points in different frames with the same label output in step (32) is summed, that is, the sum of the (m-1) / 2 items with the smallest Manhattan distance is calculated; the sum of the (m-1) / 2 items with the smallest Manhattan distance corresponding to the skeleton point i and the skeleton points of the continuous frames with the same label is recorded as D i ;
[0132] The reference point of the continuous frame pose skeleton point is to find the sum of (m-1) / 2 items with the smallest Manhattan distance from each group of skeleton points with the same label. i , the smallest D i The corresponding bone point P i As the reference point of the label skeleton point, the reference point and the skeleton point corresponding to the smallest (m-1) / 2 Manhattan distance are recorded as the reference point Rf j , merged and added to the reference point set Refer{P i ,Rf1,…,Rf (m-1) / 2}middle;
[0133] (34) Find the center point of the continuous frame posture skeleton point: The reference point set output in step (33) is calculated by calculating the arithmetic average of the corresponding coordinates of the reference points under the corresponding labels to obtain the position coordinates P of the geometric center point. center , the expression is:
[0134]
[0135] By finding the center point P of 25 skeleton points center , merge m groups of consecutive frames, each with 25 skeleton points, to obtain the final center point set matrix Pose of the human body posture skeleton points center .
[0136] The discrete point heat map generation algorithm is used to obtain a continuous spatial domain confidence mask, thereby obtaining a human body posture point cloud and a confidence fusion vector. The specific steps are as follows:
[0137] (41) Draw a discrete point heat map: The center point matrix Pose of the human body posture skeleton points obtained in step (34) center Converted into a vector with four dimensions of H, W, D, and C, which is a discrete point heat map (DiscreteHeatMap), where H, W, and D represent the vertical height, horizontal width, and depth range of the discrete point heat map, respectively. C represents the confidence weight. The confidence weight of the point on the discrete point heat map corresponding to the center point is set to 1, and the rest are set to 0.
[0138] (42) Expanding the confidence heat map: The center point obtained in step (34) is quantitatively expanded by the Gaussian expansion algorithm (·) to form a confidence heat map. The heat map of concentric circles is expanded outward with the center point as the center. The corresponding confidence value C is expressed as:
[0139]
[0140] Among them, h i , w i and d i are the three-dimensional (H, W, D) coordinates corresponding to the 25 center points, and N is expressed as:
[0141]
[0142] Among them, σ h , σ w , σ d Represent the distribution parameters of the heat map along H, W, and D, respectively, controlling the size and value of the fixation area and the boundary area;
[0143] (43) Human body posture point cloud and confidence fusion: The confidence heat map obtained in step (42) is used as a continuous spatial domain confidence mask and fused with the human body posture point cloud; the skeleton point P in the human body posture point cloud i , whose corresponding vertical, horizontal and depth coordinates are h i 、w i and d i , bring the corresponding coordinates into the confidence heat map in step (42) to find the skeleton point P i The corresponding confidence value C(P i ), and then combine the confidence value with the three-dimensional coordinates of the skeleton point to obtain the point cloud fusion vector P i (h,w,d,c).
[0144] The method of obtaining fine-grained human posture information by using a fine-grained human posture fitting algorithm specifically includes:
[0145] Multi-dimensional feature fusion: The coordinate coord and confidence conf in the point cloud fusion vector are fused as multi-dimensional inputs, and the point cloud position feature encoder is used to encode coord into Encoder(coord), and the point cloud confidence feature encoder is used to encode conf into Encoder(conf); then the encoded point cloud position feature Encoder(coord) and point cloud confidence feature Encoder(conf) are fused using the TransFuser structure to obtain a fine-grained human posture vector F Pose , expressed as:
[0146] F Pose =TransFuser[Encoder(coord),Encoder(conf)]
[0147] Among them, the feature fusion structure TransFuser consists of an encoder, a decoder, and a multi-dimensional input feature attention module Attention(·), which can be expressed as:
[0148]
[0149] Among them, Q represents the query matrix, K represents the key-value matrix, and V represents the value matrix. is the scale factor; the human body posture is fitted through a set of linear layers Linear(·), as follows:
[0150] (51) Input the point cloud fusion vector in step (43) into the position feature encoder and the confidence feature encoder respectively, including the point cloud position feature coord and the point cloud confidence feature conf;
[0151] (52) Point cloud position feature encoding: The coordinate feature of the point cloud fusion vector is The three-dimensional coordinates are converted into a single-dimensional morpheme vector through the spatial coordinate mapping voxel (·) Divide into two-dimensional plane slices according to the time dimension The corresponding attention features are Q coord , K coord and V coord ;
[0152] (53) Point cloud confidence feature encoding: The confidence feature of the point cloud fusion vector is conf = (conf1, conf2, ..., conf n ),in The corresponding attention features are Q conf , K conf and V conf ;
[0153] (54) Multi-head attention extraction: position attention feature Q of point cloud fusion vector coord , K coord , V coord And the confidence attention feature Q of the point cloud fusion vector conf , K conf , V conf , the corresponding features are exchanged using the multi-head attention mechanism, and the two sets of features after the exchange are A coord =att((K coord ,V coord ),Q conf ) and A conf =att((K conf ,V conf ),Q coord );
[0154] (55) Joint attention mechanism: The features extracted by multi-head attention are used as the input of the joint feature encoder, and the joint attention mechanism is used to fuse the position and confidence of the human posture point cloud into high-dimensional features. The corresponding feature representation is: V coord =Concat(A coord ,A conf );
[0155] Decoder(·) is used to decode the fine-grained human posture vector F Pose Decode to get the corresponding position encoding feature F coord , denoted as F coord =Decoder(F Pose ), and use the fully connected layer to map the features into the physical space to obtain the coordinates of the corresponding human posture points.
[0156] The present invention provides a method for human posture recognition based on time-space fusion of millimeter-wave radar, based on the above system, and the steps are as follows:
[0157] 1) Collect the frequency modulated continuous wave (FMCW) signal emitted by the millimeter wave radar after scanning the target human body;
[0158] 2) Demodulate the collected signal to obtain an intermediate frequency signal;
[0159] 3) Extract the distance and angle information between the millimeter wave antenna and the target human body from the intermediate frequency signal;
[0160] 4) Using the constant false alarm rate (CFAR) algorithm to filter the intermediate frequency signal, and converting the polar coordinates into spatial coordinates to construct the human body posture point cloud;
[0161] 5) Perform coarse-grained human posture skeleton point recognition;
[0162] 6) Using the continuous frame pose skeleton points, calculate the sum of the (m-1) / 2 items with the smallest Manhattan distance D i , D i The smallest bone point is used as the reference point, and a reference point set of 25 bone points is constructed, and the center point matrix of the human body posture bone points is merged;
[0163] 7) Using the discrete point heat map and Gaussian expansion algorithm, the continuous spatial domain confidence mask is calculated and fused with the human body posture point cloud to obtain the point cloud fusion vector;
[0164] 8) Construct a point cloud position feature and confidence feature encoder to process the point cloud fusion vector, use the joint attention mechanism to fuse the position and confidence features, and use the decoder to identify the coordinates of the human body posture points.
[0165] The present invention has many specific application paths. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements can be made without departing from the principles of the present invention. These improvements should also be considered as the scope of protection of the present invention.
Claims
1. A time-space fusion human posture recognition system based on millimeter wave radar, characterized in that: include: Feature extraction module and spatiotemporal feature fusion and posture recognition module; The feature extraction module is used to extract the distance and angle information of the human body from the frequency modulated continuous wave signal after scanning the target human body, and construct a human body posture point cloud; The spatiotemporal feature fusion and posture recognition module is used to extract the time domain and spatial domain features of the human body posture point cloud signal and perform human body posture recognition through feature fusion.
2. The time-space fusion human posture recognition system based on millimeter wave radar according to claim 1 is characterized in that: The feature extraction module uses the millimeter wave radar to transmit frequency modulated continuous wave as the transmission signal. The signal characteristics are: each group consists of M frames of Chirp signal, the corresponding signal period is T, and the corresponding initial frequency is f c ; Receive the reflected signal after scanning the target human body, the interval between the transmitted signal and the received signal is t d And the corresponding reflected signal of each group of Chirp signals is mixed with the transmitted signal to obtain the demodulated intermediate frequency signal. The function expression of the transmitted signal TX(t) and the received signal RX(t) with respect to time t is: Where e is the natural logarithm, j is the imaginary unit, and B represents the signal bandwidth corresponding to each group of Chirp signals. The corresponding intermediate frequency signal IF(t) is expressed as a function of time t: The expression of the intermediate frequency signal phase φ composed of the irrelevant terms at time t in the intermediate frequency signal IF(t) is: By measuring the phase φ of the intermediate frequency signal after mixing, the interval time t between the transmitted signal and the received signal can be determined. d .
3. The time-space fusion human posture recognition system based on millimeter wave radar according to claim 2 is characterized in that: The step of extracting the distance information of the human body specifically includes: A single-frame Chirp signal is extracted from the intermediate frequency signal, and a discrete point Fourier transform algorithm is performed on each Chirp signal. The Chirp signal corresponding to the target person is determined by detecting the peak position, and then the interval time t between the transmitted signal and the received signal is determined based on the intermediate frequency signal phase φ after demodulation of the transmitted signal and the received signal. d ; The time interval between the transmitted signal and the received signal is t d Extract the distance d, t between the millimeter wave antenna and the target human body d The ratio of the distance r from the signal transmission to the reception to the electromagnetic wave propagation speed c is used to calculate the distance d=ct between the millimeter wave antenna and the target human body. d / 2.
4. The millimeter-wave radar-based spatiotemporal fusion human posture recognition system according to claim 2, characterized in that: The extracting of angle information of the human body specifically includes: The phase difference Δφ of the intermediate frequency signal under multiple sets of receiving antennas on the millimeter wave radar describes the distance difference Δd between the millimeter wave antenna and the target human body. The corresponding expression is: Where λ is the wavelength of the millimeter-wave radar signal; The distance difference Δd between the millimeter wave antenna and the target human body represents the product of the distance l between each millimeter wave receiving antenna and the sine value of the angle θ between the millimeter wave antenna and the target human body, that is, Δd = lsin(θ). The expression for the angle between the millimeter wave antenna and the target human body is:
5. The time-space fusion human posture recognition system based on millimeter wave radar according to claim 1 is characterized in that: The constructing of the human body posture point cloud specifically includes: The constant false alarm rate algorithm is used to filter the reflection intensity of the intermediate frequency signal. By setting the adjustable detection threshold Th, the local peak point with large signal intensity is selected to control the target point detection sensitivity. The expression of the detection threshold Th is: Th=αP n Among them, α is the threshold coefficient, P n is the noise intensity estimate, and the expression is: Among them, P fa is the false alarm rate, N is the number of samples, x m is the sample value; the local peak point of the acquired millimeter wave signal intensity is converted, and the distance and angle of the peak point in the polar coordinate system are converted into point cloud points of the horizontal, vertical and depth coordinates of the physical space, thereby constructing the human body posture point cloud.
6. The millimeter wave radar-based spatiotemporal fusion human posture recognition system according to claim 1, characterized in that: The spatiotemporal feature fusion and posture recognition module uses a coarse-grained human posture recognition algorithm to convert the constructed human posture point cloud into coarse-grained human posture skeleton points; uses the consistency of human posture changes in consecutive frames in the time domain, and uses a merging algorithm of consecutive frame skeleton point sets to make consistency judgments on the human postures of consecutive frames, thereby reducing the large deviation of posture skeleton points caused by multipath and ghost points, and merging them into the correct posture range of the current frame; uses the continuity of posture position in the spatial domain, and uses a discrete point heat map generation algorithm to convert the discrete posture skeleton points obtained in the time domain into a continuous spatial domain confidence mask, thereby fusing the confidence with the human posture point cloud; uses a fine-grained human posture recognition algorithm to fuse the confidence mask with the human posture point cloud to obtain fine-grained human posture information.
7. The millimeter wave radar-based spatiotemporal fusion human posture recognition system according to claim 6, characterized in that: The coarse-grained human posture recognition algorithm is used to obtain coarse-grained human posture skeleton points, and the specific steps are as follows: (21) Vectorized human body posture point cloud: The physical space position of the target human body is divided into positions, and the minimum spatial resolution of the millimeter wave is used as a reference. The measurable space where the target human body is located is used as the range boundary. A spatial voxel dictionary of the horizontal, vertical and depth dimensions is created, and the corresponding point cloud is mapped into the spatial voxel dictionary, thereby obtaining n groups of voxel vectors consisting of n consecutive frames of point cloud; (22) Extract voxel vector features: The encoder structure of the temporal Transformer is used to extract features from the n sets of voxel vectors output in step (21). Specifically, the feature dimension is expanded through the embedding layer (·), and the context relevance of the features is extracted through the encoder (·). Each encoder (·) contains two sets of tensors: the intermediate unit state and the final unit state. The state tensor transmission is controlled by the memory gate structure. (23) Calculating feature correlation: Using the decoder structure of the time-series Transformer, the features output in step (22) are correlated. Specifically, the correlation matrix of each set of features is calculated through the attention layer Attention(·), and the features are fused with the correlation matrix through the decoder Decoder(·) to merge and extract features, and the features are passed back. (24) Predicting coarse-grained human posture skeleton points: The features output in step (23) are mapped to the voxel vectors using the fully connected layer structure of the temporal Transformer to obtain the human posture skeleton points. Specifically, the vector sequence is integrated and dimensionally transformed through the fully connected layer FC(·), and the highest possible voxel classification is found by the normalization function Softmax(·) activation function. Finally, the voxels are restored to physical spatial positions through the spatial voxel dictionary to obtain the position prediction results of the coarse-grained human posture skeleton points.
8. The millimeter wave radar-based spatiotemporal fusion human posture recognition system according to claim 6, characterized in that: The correct posture range of the current frame is obtained by the merging algorithm of the skeleton point sets of consecutive frames. The specific steps are as follows: (31) Extracting continuous frame posture skeleton points: Select the current frame and its previous m-1 frames to form m groups of continuous human posture estimation vectors, and extract the coarse-grained human posture skeleton points with the same label in the corresponding time series in chronological order. The coarse-grained human posture skeleton points are the 25 skeleton points that constitute the human posture frame, and form a continuous frame posture skeleton point vector of [m×25]; (32) Calculate the distance between the continuous frame posture skeleton points: Calculate the distance of the 25 skeleton points corresponding to the continuous frame posture skeleton point vector output in step (31), and obtain the Manhattan distance matrix M between the skeleton point i and the skeleton point j by calculating the Manhattan distance Manh() of the corresponding label skeleton points in the continuous frame. ij ; (33) Find the reference point of the skeleton points of the continuous frame posture: half sum the Manhattan distances of the skeleton points in different frames with the same label output in step (32), that is, calculate the sum of the (m-1) / 2 terms with the smallest Manhattan distance; The sum of the (m-1) / 2 terms with the smallest Manhattan distance between skeleton point i and its consecutive frame skeleton points with the same label is recorded as D i ; The reference point of the continuous frame pose skeleton point is to find the sum of (m-1) / 2 items with the smallest Manhattan distance from each group of skeleton points with the same label. i , the smallest D i The corresponding bone point P i As the reference point of the label skeleton point, the reference point and the skeleton point corresponding to the smallest (m-1) / 2 Manhattan distance are recorded as the reference point Rf j , merged and added to the reference point set Refer{P i ,Rf1,…,Rf (m-1) / 2 }middle; (34) Find the center point of the continuous frame posture skeleton point: The reference point set output in step (33) is calculated by calculating the arithmetic average of the corresponding coordinates of the reference points under the corresponding labels to obtain the position coordinates P of the geometric center point. center , the expression is: By finding the center point P of 25 skeleton points center , merge m groups of consecutive frames, each with 25 skeleton points, to obtain the final center point set matrix Pose of the human body posture skeleton points center .
9. The millimeter wave radar-based spatiotemporal fusion human posture recognition system according to claim 8, characterized in that: The discrete point heat map generation algorithm is used to obtain a continuous spatial domain confidence mask, thereby obtaining a human body posture point cloud and a confidence fusion vector. The specific steps are as follows: (41) Draw a discrete point heat map: The center point matrix Pose of the human body posture skeleton points obtained in step (34) center Convert it into a vector with four dimensions of H, W, D, and C, which is a discrete point heat map. H, W, and D represent the vertical height, horizontal width, and depth range of the discrete point heat map, respectively. C represents the confidence weight. The confidence weight of the point on the discrete point heat map corresponding to the center point is set to 1, and the rest are set to 0. (42) Expanding the confidence heat map: The center points obtained in step (34) are quantitatively expanded by the Gaussian expansion algorithm (·) to form a confidence heat map; The corresponding confidence value C of the heat map with the center point as the center and concentric circles expanding outward is expressed as: Among them, h i , w i and d i are the three-dimensional (H, W, D) coordinates corresponding to the 25 center points, and N is expressed as: Among them, σ h , σ w , σ d They represent the distribution parameters of the heat map along H, W, and D respectively; (43) Human body posture point cloud and confidence fusion: The confidence heat map obtained in step (42) is used as a continuous spatial domain confidence mask and fused with the human body posture point cloud; the skeleton point P in the human body posture point cloud i , whose corresponding vertical, horizontal and depth coordinates are h i 、w i and d i , bring the corresponding coordinates into the confidence heat map in step (42) to find the skeleton point P i The corresponding confidence value C(P i ), and then combine the confidence value with the three-dimensional coordinates of the skeleton point to obtain the point cloud fusion vector P i (h,w,d,c).
10. A method for human posture recognition based on millimeter wave radar and spatiotemporal fusion, based on the system according to any one of claims 1 to 9, characterized in that: The steps are as follows: 1) Collect the frequency modulated continuous wave signal emitted by the millimeter wave radar after scanning the target human body; 2) Demodulate the collected signal to obtain an intermediate frequency signal; 3) Extract the distance and angle information between the millimeter wave antenna and the target human body from the intermediate frequency signal; 4) Using the constant false alarm rate algorithm to filter the intermediate frequency signal, and converting the polar coordinates into spatial coordinates to construct the human body posture point cloud; 5) Perform coarse-grained human posture skeleton point recognition; 6) Using the continuous frame pose skeleton points, calculate the sum of the (m-1) / 2 items with the smallest Manhattan distance D i , D i The smallest bone point is used as the reference point, and a reference point set of 25 bone points is constructed, and the center point matrix of the human body posture bone points is merged; 7) Using the discrete point heat map and Gaussian expansion algorithm, the continuous spatial domain confidence mask is calculated and fused with the human body posture point cloud to obtain the point cloud fusion vector; 8) Construct a point cloud position feature and confidence feature encoder to process the point cloud fusion vector, use the joint attention mechanism to fuse the position and confidence features, and use the decoder to identify the coordinates of the human body posture points.