Millimeter wave signal reconstruction method and device for indoor monitoring, equipment and medium
By combining video frame data with the spatiotemporal alignment and enhancement processing of millimeter-wave radar signals, the problems of data sparsity and insufficient privacy protection in indoor monitoring by millimeter-wave radar are solved, achieving high-precision target detection and tracking, and adapting to different environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SCI & TECH PATENT OFFICE
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-15
AI Technical Summary
Existing millimeter-wave radar suffers from data sparsity and low signal-to-noise ratio in indoor monitoring, making it difficult to accurately identify subtle posture changes. Furthermore, its privacy protection capabilities are insufficient, limiting its application in fall detection and posture recognition for the elderly.
By combining video frame data with millimeter-wave radar signals, spatiotemporal alignment, preprocessing, and prior-driven mapping of compressed sensing or deep network fusion models are performed to enhance radar echo data. Target tracking is then performed using constant false alarm rate detection and Kalman filtering to generate high-quality motion trajectory information.
It significantly improves the accuracy and robustness of target detection and tracking, providing high-precision monitoring and tracking data in complex indoor environments, ensuring privacy protection, and adapting to different deployment conditions.
Smart Images

Figure CN122043399A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of millimeter-wave radar technology, and in particular to a millimeter-wave signal reconstruction method, system and storage medium for indoor monitoring. Background Technology
[0002] With the accelerating aging of the population, home-based elderly care and smart nursing have become key areas of public health. Real-time fall monitoring, posture recognition, and abnormal behavior alerts for elderly people indoors can significantly reduce secondary injuries and medical costs caused by falls, and also improve their quality of life and sense of security. Currently, monitoring solutions based on visible light cameras can provide clear images and rich posture information, thus performing excellently in tasks such as fall detection and motion recognition. However, cameras directly capture images, faces, and details of daily life, which can easily leak the personal privacy of those being monitored, especially in home environments. Privacy has become a key bottleneck hindering the large-scale deployment of this technology.
[0003] Millimeter-wave radar, due to its "invisible and audible" characteristics, can acquire the distance, speed, and micro-motion information of a target without recording visible images, naturally possessing strong privacy protection capabilities. At the same time, millimeter-wave signals can penetrate obstructions such as clothing and furniture, and are insensitive to changes in lighting, making them particularly suitable for fall detection scenarios in dim lighting or at night. However, the application of existing millimeter-wave radar in crowd or single-person posture monitoring still has two major limitations: sparse data and low signal-to-noise ratio. The weak echoes generated when elderly people fall are often masked by environmental noise, resulting in discrete time-frequency characteristics, making it difficult to directly use for accurate fall detection; insufficient spatial resolution. Traditional point cloud or spectrum processing methods struggle to recover high-resolution features sufficient to distinguish subtle posture changes such as standing, sitting, and falling, limiting the reliability of detection in multi-target, dense scenes.
[0004] Therefore, there is an urgent need for a millimeter-wave signal reconstruction method for indoor monitoring, which can achieve high-precision indoor personnel monitoring through millimeter-wave radar. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a millimeter-wave signal reconstruction method, system, and storage medium for indoor monitoring, which improves the accuracy of target object detection in indoor areas.
[0006] In a first aspect, this application provides a millimeter-wave signal reconstruction method for indoor monitoring, the method comprising: Step S1: Acquire video frame data of the target object in the indoor area and raw I / Q data collected by millimeter-wave radar. Extract key feature data of the target object from the video frame data. Perform spatiotemporal alignment on the video frame data and the raw I / Q data. Use a rigid body transformation matrix to map them to the radar coordinate system to form prior information about the human body. Step S2: Preprocess the raw I / Q data to obtain the amplitude matrix and phase matrix. Calculate the average energy based on the non-target region of the amplitude matrix to obtain the noise baseline. Normalize the amplitude matrix based on the noise baseline and combine it with the phase matrix to generate radar echo data. Step S3: Establish a priori-driven mapping model based on the weighted reconstruction method of compressed sensing or the deep network fusion model, and perform enhancement processing on the radar echo data based on the priori-driven mapping model to obtain enhanced radar echo data; Step S4: Apply constant false alarm rate (CFAR) detection to the enhanced radar echo data amplitude matrix to generate a binary detection map, and use Kalman filtering to track the target object and obtain its motion trajectory information.
[0007] In conjunction with the first aspect, in a first implementation of the first aspect of this application, spatiotemporal alignment of the video frame data and the raw I / Q data includes: Set a preset time error, wherein the time difference between the sampling time of the kth frame of the video frame data and the acquisition time of the original I / Q data in the kth frame acquisition period is less than the preset time error; Set the rigid body transformation matrix , ,in, This represents the rotation matrix of the camera coordinate system relative to the radar coordinate system. The translation vector from the camera origin to the radar origin is represented by the rigid body transformation matrix. Key features in the camera coordinate system are mapped to the radar coordinate system based on the rigid body transformation matrix, thereby achieving spatial alignment between the video frame data and the original I / Q data.
[0008] In conjunction with the first aspect, in the second implementation of the first aspect of this application, the raw I / Q data is preprocessed to obtain an amplitude matrix and a phase matrix, including: The original I / Q data is subjected to a Fast Fourier Transform to obtain a complex matrix containing multi-dimensional information of the target object. The complex matrix is then separated into amplitude and phase to obtain an amplitude matrix and a phase matrix.
[0009] In conjunction with the first aspect, in a third implementation of the first aspect of this application, calculating the average energy to obtain the noise baseline based on the non-target region of the amplitude matrix includes: Semantic segmentation is performed on the simultaneously acquired video frame data to obtain the segmentation results of human and non-human regions. Based on the rigid body transformation matrix, the segmentation results are mapped to the radar coordinate system to determine the radar grid set corresponding to the non-human region. The average amplitude energy of the radar grid set of the non-human region is calculated to obtain the noise baseline.
[0010] In conjunction with the first aspect, in the fourth implementation of the first aspect of this application, a priori-driven mapping model is established based on a weighted reconstruction method of compressed sensing, including: The magnitude matrix is vectorized, and the vectorized magnitude matrix is expressed as sparse coefficients under a sparse basis based on the first formula, which is: ,in, It is a sparse transformation basis. It is a sparse coefficient vector. This is the vectorized magnitude matrix; Calculate the minimum difference between each radar beam angle in the amplitude matrix and the human body angle in the prior human body data, and define the weighting coefficient of the k-th radar beam angle based on the second formula. The second formula is: = ,in, The minimum difference, These are the angle bandwidth control parameters; A diagonal weight matrix is constructed based on the weighting coefficients for each radar beam angle; A weighted observation equation is constructed based on the third formula, which is: ,in, For weighted observation signals, For radar observation matrix, For noise terms, W is the diagonal weight matrix; The weighted observation equation is solved using an iterative threshold shrinkage algorithm or a fast iterative algorithm to solve the sparse optimization problem and obtain a sparse solution. ; The sparse solution is reconstructed through the inverse transformation of the sparse basis to obtain the enhanced amplitude vector. ; For the amplitude vector An inverse vectorization operation is performed to obtain an enhanced amplitude matrix. The enhanced amplitude matrix is then combined with the original phase matrix to obtain enhanced radar echo data.
[0011] In conjunction with the first aspect, in the fifth implementation of the first aspect of this application, a priori-driven mapping model is established based on a deep network fusion model, including: A priori-driven mapping model is established based on a deep network fusion model. The amplitude matrix and phase matrix are input as radar branches into the radar encoder of the priori-driven mapping model to obtain radar features. The human body prior data is input as an auxiliary branch into the prior encoder of the priori-driven mapping model to obtain prior features. Feature fusion is performed on the radar features and the prior features to obtain joint features. The joint features are input into the decoder to output the enhanced amplitude, and combined with the original phase to obtain the enhanced radar echo data.
[0012] In conjunction with the first aspect, in the sixth implementation of the first aspect of this application, constant false alarm rate detection is applied to the enhanced radar echo data amplitude matrix to generate a binary detection map, including: A constant false alarm rate (CFAR) detection algorithm is applied to the amplitude matrix of the enhanced radar echo data. The amplitude values of each radar grid point are compared with the background noise threshold to mark the regions where target objects reflect, and a binary detection map is generated.
[0013] In conjunction with the first aspect, in the seventh implementation of the first aspect of this application, Kalman filtering is used to track the target object and obtain the motion trajectory information of the target object, including: Based on the binary detection map, adjacent detection response points are clustered to obtain candidate target clusters. The candidate target clusters are used as observation inputs to extract the center position, amplitude energy, and motion trend features of the target object, construct the observation vector of the target object, and use the Kalman filter algorithm to recursively estimate the state parameters of the target object. The state parameters include distance, radial velocity, and angle information. The state parameters are converted into position information in a plane coordinate system and combined with the velocity vector to form the motion trajectory information of the target object.
[0014] Secondly, this application provides a millimeter-wave signal reconstruction system for indoor monitoring, the system comprising: The acquisition module acquires video frame data of target objects in the indoor area and raw I / Q data collected by millimeter-wave radar. It extracts key feature data of the target objects from the video frame data, performs spatiotemporal alignment on the video frame data and raw I / Q data, and maps them to the radar coordinate system using a rigid body transformation matrix to form prior human information. The preprocessing module preprocesses the raw I / Q data to obtain an amplitude matrix and a phase matrix. Based on the non-target region of the amplitude matrix, it calculates the average energy to obtain a noise baseline. Based on the noise baseline, it normalizes the amplitude matrix and combines it with the phase matrix to generate radar echo data. The enhancement module establishes a priori-driven mapping model based on a weighted reconstruction method of compressed sensing or a deep network fusion model, and performs enhancement processing on the radar echo data based on the priori-driven mapping model to obtain enhanced radar echo data. The generation module applies constant false alarm rate detection to the enhanced radar echo data amplitude matrix to generate a binary detection map, and uses Kalman filtering to track the target object and obtain the target object's motion trajectory information.
[0015] A third aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the millimeter-wave signal reconstruction method for indoor monitoring described above.
[0016] Compared with the prior art, the beneficial effects of the present invention are at least as follows: This application provides a millimeter-wave signal reconstruction method for indoor monitoring, which can effectively improve the accuracy and robustness of target detection and tracking by combining video data with millimeter-wave radar signals. First, by synchronizing video frame data and radar data through spatiotemporal alignment, key features in the video can be accurately mapped to the radar coordinate system, thus forming prior information about the human body. Next, by preprocessing the raw I / Q data, calculating the noise baseline and normalizing the amplitude, the influence of environmental noise and static clutter is significantly suppressed, ensuring the stability of subsequent radar signal enhancement. Furthermore, a priori-driven mapping model is established using a weighted reconstruction method based on compressed sensing or a deep network fusion model, effectively enhancing the target signal in the radar echo data, making the target's reflection information clearer. In addition, by tracking the target object through constant false alarm rate detection and Kalman filtering, the target's motion trajectory information can be accurately extracted, ensuring efficient target tracking and abnormal behavior monitoring in complex indoor environments. Finally, this application can provide high-quality target monitoring and tracking data under different environmental and deployment conditions, providing stable and reliable support for indoor monitoring systems and significantly improving the system's accuracy, robustness, and adaptability in practical applications. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of an embodiment of a millimeter-wave signal reconstruction method for indoor monitoring in this application. Figure 2 This is a schematic diagram of one embodiment of a millimeter-wave signal reconstruction system for indoor monitoring in this application. Detailed Implementation
[0019] This application provides a millimeter-wave signal reconstruction method, system, and storage medium for indoor monitoring. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0020] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of a millimeter-wave signal reconstruction method for indoor monitoring in this application includes: Step S1: Acquire video frame data of the target object in the indoor area and the raw I / Q data collected by the millimeter-wave radar. Extract key feature data of the target object from the video frame data. Perform spatiotemporal alignment on the video frame data and the raw I / Q data. Use a rigid body transformation matrix to map them to the radar coordinate system to form prior information about the human body.
[0021] Specifically, in indoor monitoring applications, video acquisition devices (e.g., high-definition cameras, RGB-D cameras) are used to acquire video frame data of the target object, such as the elderly, within a preset time period. The video frame data is a color image or an image containing depth information. At the same time, millimeter-wave radar is used to acquire raw I / Q data. The raw I / Q data is the in-phase (I) and quadrature (Q) baseband sampling signal obtained by down-conversion after the millimeter-wave radar receives the echo. It is usually represented in complex form and is used to monitor the location information, motion characteristics and dynamic state of the elderly. To ensure a one-to-one correspondence between video frames and radar frames, a timestamp is recorded for each video frame and radar frame, and a preset time error threshold is set. Based on this threshold, the video frame data and the original I / Q data are time-aligned. A rigid transformation matrix is established between the video camera coordinate system and the radar coordinate system. This matrix, composed of a rotation matrix and a translation vector, describes the spatial relationship between the camera coordinate system and the radar coordinate system. Based on the rigid transformation matrix, the key feature points of the target extracted from the video frames are mapped to the radar coordinate system, achieving spatial alignment between the video frame data and the original I / Q data, thus forming prior information about the human body. The detailed spatial alignment process is as follows: First, semantic segmentation is performed on the video frame data to obtain mask images of the human and non-human regions, denoted as: Its value takes the range {0, 1}, where 1 represents the human body region and 0 represents the background region; human keypoint detection is performed within the human body region to obtain the pixel coordinates of each human joint point in the image plane. Visibility With confidence information , denoted as: Using the camera's intrinsic parameter matrix K and depth information By projecting the two-dimensional coordinates of key points into three-dimensional space, the corresponding three-dimensional coordinates are obtained. This process ensures that video key points correspond to target locations in physical space; it also performs differential calculations on the key point positions in consecutive frames to compute the three-dimensional motion vector for each key point. ; Semantic segmentation results Key point set The combination of three-dimensional coordinates and motion vectors at the same time stamp constitutes the prior information of the human body. This provides input for subsequent fusion with radar data. By establishing a spatiotemporal alignment relationship between video and radar data in indoor monitoring scenarios, key human features from the video side can be accurately mapped to the radar coordinate system, thereby forming prior human information that can be directly used for signal reconstruction. This effectively reduces mismatches between sensors, improves the accuracy of subsequent noise baseline estimation and enhancement processing, and ultimately enhances the robustness and reliability of the overall monitoring.
[0022] Step S2: Preprocess the raw I / Q data to obtain the amplitude matrix and phase matrix. Calculate the average energy based on the non-target area of the amplitude matrix to obtain the noise baseline. Normalize the amplitude matrix based on the noise baseline and combine it with the phase matrix to generate radar echo data.
[0023] Specifically, the raw I / Q data acquired by millimeter-wave radar includes hardware bias, static clutter, and environmental noise. The amplitude magnitude varies greatly in different scenarios, and direct use for detection can lead to unstable thresholds and false alarms. Therefore, the raw I / Q data needs to be preprocessed, including range-directed transform (FFT), and combined with slow time / antenna dimension to obtain a complex matrix of angle-range or range-velocity-angle. From the complex matrix, the amplitude matrix (echo intensity at each grid point) and phase matrix (phase information at each grid point) are extracted. The radar grid point set of non-target areas is obtained from the prior human data. The average energy of the amplitude values in the non-target areas is used as the noise baseline. The specific noise baseline generation process will be explained in detail later. The amplitude matrix is normalized grid by grid according to the noise baseline, and reasonable pruning or logarithmic compression is performed to obtain the normalized amplitude matrix. The original phase matrix remains unchanged. The normalized amplitude matrix and phase matrix are combined to obtain the radar echo data. Through the prior-guided noise baseline and normalization processing, the influence of scene differences and static clutter can be significantly reduced, stabilizing the subsequent threshold setting and detection performance.
[0024] Step S3: Establish a priori-driven mapping model based on the weighted reconstruction method of compressed sensing or the deep network fusion model, and perform enhancement processing on the radar echo data based on the priori-driven mapping model to obtain enhanced radar echo data.
[0025] Specifically, in indoor monitoring scenarios, millimeter-wave radar echoes are significantly affected by multipath propagation, non-Gaussian noise, equipment differences, and variations in the deployment environment. A single algorithm struggles to simultaneously address multiple metrics, including interpretability, dependence on data scale, computational power and latency constraints, and cross-scenario generalization capabilities. To ensure the method consistently achieves enhancement effects under different deployment conditions, this application provides two methods for constructing a priori-driven mapping models. These models are used to enhance radar echo data, yielding enhanced radar echo data. The first method is a weighted reconstruction method based on compressed sensing, while the second method is based on a deep network fusion model. The first method's general process is as follows: the amplitude matrix is represented as a sparse coefficient vector under a sparse basis, and a weighted observation equation is established by combining human priors with a weight matrix. The enhanced amplitude is obtained through iterative sparse optimization, and then combined with the original phase to generate enhanced echo data. Detailed explanations follow. The second construction method is roughly as follows: radar amplitude and phase are input into the radar branch, and human prior is input into the prior branch. After extracting and fusing features, the enhanced amplitude is output by the decoder and then combined with the original phase matrix to generate enhanced echo data. The detailed process will be explained later. Compressed sensing provides physically interpretable, lightweight and robust baseline capabilities, while deep network methods provide nonlinear expression and cross-domain adaptation performance limits. The two complement each other in terms of data scale, computing power budget, and interpretability requirements. The strategy selection of the two construction methods is based on the available training data scale, online computing power and latency budget, and the degree of change of the target scene: for example, compressed sensing is preferred when training data is insufficient or edge computing power is limited, while deep network implementation is preferred when training conditions are available and the upper limit needs to be improved in complex environments. Optionally, environmental drift and performance indicators are monitored during runtime to trigger strategy switching or joint activation (with deep network output as the main method and compressed sensing output used for confidence calibration / consistency constraints). By setting two alternative implementation methods, this application can obtain stable enhancement effects under different resource and data conditions.
[0026] Step S4: Apply constant false alarm rate detection to the enhanced radar echo data amplitude matrix to generate a binary detection map, and use Kalman filtering to track the target object and obtain the target object's motion trajectory information.
[0027] Specifically, a constant false alarm rate (CFAR) detection algorithm is applied to the enhanced radar echo amplitude matrix. The CFAR detection algorithm is a radar target detection algorithm that ensures a constant false alarm rate by adaptively setting a threshold. It is used to identify target objects in noisy backgrounds. During the detection of target objects, the amplitude value of each radar grid point of the target object is compared with the background noise estimate of the surrounding reference window. Here, the radar grid point is a single sampling point on the angle-range discrete grid, and the amplitude value represents the echo intensity of the point. When the amplitude value exceeds the adaptive threshold, the grid point is marked as a detection response point. The detection response point is a grid point that may have the reflection of the target object. Finally, all detection response points are combined to form a binary detection map, which is used to distinguish the candidate region of the target object from the background region. Adjacent response points in the binary detection map are clustered and merged into candidate target clusters. The center position, energy level, and spatial distribution characteristics of each candidate cluster are calculated and used as observation inputs for subsequent tracking. The observation information of the candidate target clusters is input into a Kalman filter to recursively estimate the target object's state parameters (including distance, radial velocity, and angle information). Through prediction and update cycles, the Kalman filter effectively filters out false alarms and occasional lost observations, achieving a smooth estimation of the target object's continuous state. The filtered state parameters are converted into position in a planar coordinate system and combined with the velocity vector to form the target object's continuous motion trajectory. The output trajectory information is used for subsequent behavior recognition or security monitoring applications. Through the above detection and tracking process, target objects can be effectively identified and continuously tracked in complex indoor environments. This not only ensures a high detection rate but also significantly reduces false alarms and trajectory jitter, achieving stable and reliable motion trajectory output, thus providing higher accuracy and robustness for indoor monitoring.
[0028] In one specific embodiment, spatiotemporal alignment of video frame data and raw I / Q data specifically includes the following steps: Set a preset time error, and set the time difference between the sampling time of the kth frame of the video frame data and the acquisition time of the original I / Q data in the kth frame acquisition cycle to be less than the preset time error; Set the rigid body transformation matrix , ,in, This represents the rotation matrix of the camera coordinate system relative to the radar coordinate system. This represents the translation vector from the camera origin to the radar origin. Based on the rigid body transformation matrix, key features in the camera coordinate system are mapped to the radar coordinate system, achieving spatial alignment between video frame data and raw I / Q data.
[0029] Specifically, a preset time error threshold is set. Assuming Compare the acquisition time of the k-th frame of the video frame with the acquisition time of the original I / Q data in the acquisition period of the k-th frame. If the time difference between the two is less than the preset time error... If the threshold is met, the video frame is considered to correspond successfully with the radar frame; otherwise, it is discarded or compensated using interpolation to ensure data matching in the time dimension. A rigid transformation matrix is established between the video camera coordinate system and the radar coordinate system. Composed of a rotation matrix R and a displacement vector T, this rigid body transformation matrix is used for spatial mapping between different sensor coordinate systems. Based on this matrix, key features extracted from the camera coordinate system are mapped to the radar coordinate system, achieving spatial unification. Through this spatiotemporal alignment process, video frame data and radar I / Q data are fused under the same time reference and unified coordinate system, effectively avoiding mismatches in target position and time. This allows key features extracted from the video to accurately guide radar signal processing, thereby improving the reliability of subsequent noise estimation and signal enhancement.
[0030] In one specific embodiment, preprocessing the raw I / Q data to obtain the amplitude matrix and phase matrix specifically includes the following steps: The original I / Q data is subjected to a fast Fourier transform to obtain a complex matrix containing multi-dimensional information of the target object. The complex matrix is then separated into amplitude and phase to obtain the amplitude matrix and phase matrix.
[0031] Specifically, a Fast Fourier Transform (FFT) is performed on the raw I / Q data. FFT is a commonly used and efficient frequency domain transformation algorithm used to convert time-domain signals into frequency-domain representations, thereby extracting the target's range or velocity features and obtaining a complex matrix containing multi-dimensional information about the target object. This complex matrix is a two-dimensional or multi-dimensional data array whose elements are complex numbers, containing amplitude and phase information, representing the target object's distribution characteristics in dimensions such as range, angle, and velocity. The complex matrix is then separated into amplitude and phase matrices. The amplitude matrix, obtained by taking the modulus of the complex matrix, is used to represent the echo energy magnitude at each radar grid point. The phase matrix, obtained by taking the phase angle of the complex matrix, reflects the phase information of the echo and describes the target object's motion and micro-motion characteristics.
[0032] In one specific embodiment, calculating the average energy to obtain the noise baseline based on the non-target region of the amplitude matrix specifically includes the following steps: Semantic segmentation is performed on the simultaneously acquired video frame data to obtain the segmentation results of human and non-human regions. The segmentation results are mapped to the radar coordinate system based on the rigid body transformation matrix to determine the radar grid set corresponding to the non-human region. The average amplitude energy of the radar grid set of the non-human region is calculated to obtain the noise baseline.
[0033] Specifically, semantic segmentation is performed on synchronously acquired video frame data to obtain segmentation results for human and non-human regions. This process can employ deep learning methods to perform pixel-level classification of the image, distinguishing between human and background regions, and ultimately generating a binary mask (i.e., human regions are 1, and non-human regions are 0). Based on the rigid body transformation matrix, the segmentation results of human and non-human regions in the video frame are mapped from the camera coordinate system to a two-dimensional grid structure in the radar coordinate system. In the radar coordinate system, the set of radar grid points corresponding to the non-human regions is determined. These grid points represent the background regions (i.e., non-target regions) in the radar data. Based on the amplitude matrix, the amplitude values in the radar grid point set corresponding to the non-human regions are extracted. By calculating the average value of the amplitude energy of these grid points, the noise baseline is obtained. The noise baseline reflects the intensity of environmental background noise.
[0034] In one specific embodiment, the prior-driven mapping model established by the weighted reconstruction method based on compressed sensing includes the following steps: The magnitude matrix is vectorized, and then expressed as sparse coefficients under a sparse basis based on the first formula: ,in, It is a sparse transformation basis. It is a sparse coefficient vector. This is the vectorized magnitude matrix; Calculate the minimum difference between each radar beam angle in the amplitude matrix and the human body angle in the prior human body data, and define the weight coefficient of the k-th radar beam angle based on the second formula. The second formula is: = ,in, To be the minimum difference, These are the angle bandwidth control parameters; A diagonal weight matrix is constructed based on the weighting coefficients for each radar beam angle; The weighted observation equation is constructed based on the third formula, which is: ,in, For weighted observation signals, For radar observation matrix, For noise terms, W is the diagonal weight matrix; The weighted observation equation is solved, and the sparse optimization problem is solved using an iterative threshold shrinkage algorithm or a fast iterative algorithm to obtain a sparse solution. ; The sparse solution is reconstructed through the inverse transformation of the sparse basis to obtain the enhanced magnitude vector. ; For the magnitude vector The enhanced amplitude matrix is obtained by performing inverse vectorization, and then combined with the original phase matrix to obtain the enhanced radar echo data.
[0035] Specifically, the process of enhancing radar echo data by establishing a prior-driven mapping model based on the weighted reconstruction method of compressed sensing is as follows: First, the amplitude matrix is vectorized, converting the two-dimensional amplitude matrix into a one-dimensional vector. The vectorized amplitude matrix is x = vec(A), where A is the amplitude matrix of the radar echo, and x is the vectorized amplitude vector. Then, based on the vectorization result of the amplitude matrix, a sparse transformation basis is used... This can be represented as a sparse coefficient vector s, which is achieved by the following formula: ,in It is the inverse transformation matrix of the sparse transformation basis. The vectorized amplitude matrix is represented under a sparse basis. Through sparse representation, the effective information of the target signal is concentrated on a few coefficients, while the remaining coefficients are close to zero, achieving a sparsity effect. Furthermore, based on the prior human data obtained in step S1, the radar beam angle of each radar beam in the amplitude matrix is calculated. With the set of prior angles of the human body The minimum difference is used to determine the proximity of the beam angle to the direction relevant to the human body; the weighting coefficient for each radar beam angle is calculated based on this minimum difference. A weight matrix is generated, where a larger weight coefficient indicates that the angle is closer to the human target and therefore more important. Based on the definition of the second formula, the angle bandwidth control parameter is used to adjust the attenuation rate of the weights. Furthermore, based on the weighted observation signal, the radar observation matrix, and the noise term, a weighted observation equation is constructed. This equation establishes the relationship between the radar echo data and the sparse coefficients. Specifically, it is assumed that the signal y observed from the radar can be represented as the sparse coefficients s of the target and... The observation matrix is combined, plus a noise term n. However, since radar signals are greatly affected by the environment, directly using the original signal will be subject to a lot of background noise interference. Therefore, this application introduces a weighted observation matrix, which reflects the importance of the target signal in the radar echo. Different weights are assigned to each observation point through weighting coefficients, so that grid points related to the target signal (e.g., the direction where the radar beam angle is close to the human body) receive higher weights, while background noise is suppressed. To solve for the sparse coefficients s, the objective function is set as follows: ,in, Represents the reconstruction error, measuring the difference between the signal generated by the model and the actual observed signal. This represents a sparse regularization term that encourages the solution s to be sparse, meaning that most coefficients are close to zero. The sparse regularization coefficients are used to balance the reconstruction error and sparsity constraints. Sparse optimization problems are typically solved using the Iterative Threshold Shrinking (ISTA) algorithm or the Fast Iterative Strategy (FISTA) algorithm. ISTA rapidly approximates the sparsest solution by gradually shrinking less important coefficients, while FISTA is an accelerated version of ISTA that utilizes gradient information and historical iteration results to speed up the optimization process. Through these iterative algorithms, the enhanced sparse coefficients are finally obtained. The obtained sparse solution The enhanced magnitude vector is obtained by reconstructing the vector through the inverse transformation of the sparse basis. The enhancement amplitude vector is restored to the enhancement amplitude matrix. The enhanced radar echo data is then combined with the original phase matrix to obtain the final enhanced radar echo data. A priori-driven mapping model is established using a weighted reconstruction method based on compressed sensing. Through sparse reconstruction optimization and weighted processing, the target signal detection capability is significantly improved and the influence of background noise is reduced. The enhanced radar echo data provides higher quality input for subsequent target detection and tracking, ensuring the accuracy and stability of indoor monitoring.
[0036] In one specific embodiment, establishing a priori-driven mapping model based on a deep network fusion model includes the following steps: A priori-driven mapping model is established based on a deep network fusion model. The amplitude matrix and phase matrix are used as radar branches input to the radar encoder of the priori-driven mapping model to obtain radar features. Human body prior data is used as an auxiliary branch input to the prior encoder of the priori-driven mapping model to obtain prior features. Feature fusion is performed on the radar features and prior features to obtain joint features. The joint features are input to the decoder to output the enhanced amplitude, and combined with the original phase to obtain the enhanced radar echo data.
[0037] Specifically, the process of enhancing radar echo data by establishing a prior-driven mapping model based on a deep network fusion model is as follows: First, the amplitude matrix and phase matrix are input into the radar branch (radar encoder) of the deep network. Human body prior data, including human body position and angle information in the video, is input into the prior encoder as an auxiliary branch. The radar encoder outputs radar features, representing the spatial and temporal patterns of the radar echo data. The prior encoder extracts spatial information, target position, angle distribution, and other features from the human body prior through a deep neural network (such as a convolutional network or a fully connected network). The prior features represent a high-level abstraction of human body information, providing geometric constraints on the target and helping to guide the enhancement of radar data. The extracted radar features and prior features are then fused. Feature fusion can be performed using techniques such as concatenation, weighted summation, or other deep fusion techniques. The fused joint features contain the low-level signal information of the radar data and the high-level spatial information of the human body prior. The joint features are input into the decoder, which is usually composed of multiple convolutional or fully connected layers. It is used to generate enhanced amplitude information from the joint features. The output of the decoder is the enhanced amplitude, that is, the signal components related to the target are enhanced on the basis of the original radar amplitude matrix. Finally, the augmentation amplitude and the original phase matrix are combined to obtain the enhanced radar echo data. By multiplying the augmentation amplitude by the original phase, the dynamic characteristics of the target (such as velocity and micromotions) are preserved, while the augmentation amplitude helps to suppress background noise and clutter.
[0038] In one specific embodiment, applying constant false alarm rate (CFAR) detection to the enhanced radar echo data amplitude matrix to generate a binary detection map specifically includes the following steps: A constant false alarm rate (CFAR) detection algorithm is applied to the amplitude matrix of the enhanced radar echo data. The amplitude values of each radar grid point are compared with the background noise threshold to mark the areas where target objects reflect, and a binary detection map is generated.
[0039] Specifically, a constant false alarm rate (CFAR) detection algorithm is applied to the enhanced amplitude matrix. The core of this algorithm is to calculate the amplitude value of each radar grid point and compare it with a dynamically adjusted noise threshold to determine whether a target signal exists at that grid point. The specific steps are as follows: The CFAR algorithm selects a reference window (usually a sliding window) around each radar grid point and estimates the background noise level of that grid point using the amplitude values of other grid points within that window; based on the background noise estimation, an adaptive threshold is set, which is related to the environmental noise level. This threshold is generally calculated using statistical methods (e.g., median, mean, etc.) to ensure a constant false alarm rate; the amplitude value of each radar grid point is compared with this threshold. If the amplitude value exceeds the threshold, it is considered that a target reflection may exist at that location; if the amplitude value is below the threshold, it is considered background noise. After constant false alarm rate (CFAR) detection, each radar grid point is marked as either a "target area" or a "non-target area." Grid points in the target area have amplitude values exceeding the noise threshold; these areas are considered target reflections detected by the radar. Finally, the CFAR algorithm outputs a binary detection map, where grid points with a value of 1 represent target reflection areas, and grid points with a value of 0 represent background noise areas. The binary detection map provides reliable input for subsequent target tracking and behavior analysis, significantly improving the accuracy and robustness of the indoor monitoring system.
[0040] In one specific embodiment, Kalman filtering is used to track the target object, and obtaining the target object's motion trajectory information specifically includes the following steps: Clustering of adjacent detection response points based on binary detection maps yields candidate target clusters. These clusters are then used as observation inputs to extract the center position, amplitude energy, and motion trend features of the target objects, constructing observation vectors for the target objects. A Kalman filter algorithm is then used to recursively estimate the state parameters of the target objects, including distance, radial velocity, and angle information. These state parameters are converted into position information in a planar coordinate system and combined with the velocity vector to form the motion trajectory information of the target objects.
[0041] Specifically, the binary detection image is first passed as input to the target clustering module. In this module, adjacent detection response points are merged into candidate target clusters using a clustering algorithm (such as DBSCAN or other clustering methods). Each cluster represents a possible target region. The clustered candidate target clusters are further processed to extract the target object's center position, amplitude energy, and motion trend features. Specifically, the geometric center of the target cluster is obtained by calculating the mean position of all grid points in the candidate target cluster, which serves as the target's spatial position. The amplitude value of each radar grid point in the candidate target cluster is calculated, and its mean or weighted average is obtained as the target's reflection intensity (energy) feature. Based on the positional changes between consecutive frames, the target's motion trend in the time dimension is calculated. The extracted center position, amplitude energy, and motion trend features are combined into an observation vector. The observation vector of the target object is input into a Kalman filter for recursive estimation. The Kalman filter is used to recursively estimate the target's state parameters, which typically include: the distance between the target object and the radar, the radial velocity, and the target's angle information relative to the radar, usually including azimuth and elevation angles. Based on the current observation and prediction information, the Kalman filter performs a smooth estimation of the target's state parameters. Through the recursive step, the Kalman filter continuously updates the target state, maintaining an accurate estimate of the target's position and velocity. The target state parameters output by the Kalman filter are converted into position information in a planar coordinate system, usually by converting polar coordinates to a Cartesian coordinate system (X, Y, Z). Combined with the target object's radial velocity and angle information, the target's trajectory in the planar coordinate system can be calculated. The trajectory information includes the target's position, velocity, direction of motion, and possible abnormal behaviors or events (e.g., fall detection, standing, walking, etc.). By clustering targets in the binary detection map and recursively estimating using Kalman filtering, the motion trajectory of the target object can be accurately and stably tracked. Especially in complex environments (such as indoor environments), by combining the target's center position, amplitude energy, and motion trend characteristics, the accuracy and robustness of target tracking are further improved, providing reliable data support for subsequent behavior analysis, anomaly monitoring, and safety early warning.
[0042] The above describes a millimeter-wave signal reconstruction method for indoor monitoring in an embodiment of this application. The following describes a millimeter-wave signal reconstruction system for indoor monitoring in an embodiment of this application. Please refer to [link to relevant documentation]. Figure 2 One embodiment of a millimeter-wave signal reconstruction system for indoor monitoring in this application includes: The acquisition module acquires video frame data of target objects in the indoor area and raw I / Q data collected by millimeter-wave radar. It extracts key feature data of target objects from video frame data, performs spatiotemporal alignment on video frame data and raw I / Q data, and uses rigid body transformation matrix to map them to radar coordinate system to form human body prior information. The preprocessing module preprocesses the raw I / Q data to obtain the amplitude matrix and phase matrix. Based on the non-target area of the amplitude matrix, the average energy is calculated to obtain the noise baseline. Based on the noise baseline, the amplitude matrix is normalized and combined with the phase matrix to generate radar echo data. The enhancement module establishes a priori-driven mapping model based on the weighted reconstruction method of compressed sensing or the deep network fusion model, and performs enhancement processing on the radar echo data based on the priori-driven mapping model to obtain enhanced radar echo data. The generation module applies constant false alarm rate detection to the enhanced radar echo data amplitude matrix to generate a binary detection map, and uses Kalman filtering to track the target object and obtain the target object's motion trajectory information.
[0043] This application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, storing instructions that, when executed on a computer, cause the computer to perform the steps of a millimeter-wave signal reconstruction method for indoor monitoring.
[0044] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0045] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0046] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A millimeter-wave signal reconstruction method for indoor monitoring, characterized in that, The method includes: Step S1: Acquire video frame data of the target object in the indoor area and raw I / Q data collected by millimeter-wave radar. Extract key feature data of the target object from the video frame data. Perform spatiotemporal alignment on the video frame data and the raw I / Q data. Use a rigid body transformation matrix to map them to the radar coordinate system to form prior information about the human body. Step S2: Preprocess the raw I / Q data to obtain the amplitude matrix and phase matrix. Calculate the average energy based on the non-target region of the amplitude matrix to obtain the noise baseline. Normalize the amplitude matrix based on the noise baseline and combine it with the phase matrix to generate radar echo data. Step S3: Establish a priori-driven mapping model based on the weighted reconstruction method of compressed sensing or the deep network fusion model, and perform enhancement processing on the radar echo data based on the priori-driven mapping model to obtain enhanced radar echo data; Step S4: Apply constant false alarm rate detection to the enhanced radar echo data amplitude matrix to generate a binary detection map, and use Kalman filtering to track the target object and obtain the target object's motion trajectory information.
2. The method according to claim 1, characterized in that, Spatiotemporal alignment of the video frame data and the raw I / Q data includes: Set a preset time error, wherein the time difference between the sampling time of the kth frame of the video frame data and the acquisition time of the original I / Q data in the kth frame acquisition period is less than the preset time error; Set the rigid body transformation matrix , ,in, This represents the rotation matrix of the camera coordinate system relative to the radar coordinate system. The translation vector from the camera origin to the radar origin is represented by the rigid body transformation matrix. Key features in the camera coordinate system are mapped to the radar coordinate system based on the rigid body transformation matrix, thereby achieving spatial alignment between the video frame data and the original I / Q data.
3. The method according to claim 1, characterized in that, The raw I / Q data is preprocessed to obtain the amplitude matrix and phase matrix, including: The original I / Q data is subjected to a Fast Fourier Transform to obtain a complex matrix containing multi-dimensional information of the target object. The complex matrix is then separated into amplitude and phase to obtain an amplitude matrix and a phase matrix.
4. The method according to claim 1, characterized in that, Based on the non-target region of the amplitude matrix, the average energy is calculated to obtain the noise baseline, including: Semantic segmentation is performed on the simultaneously acquired video frame data to obtain the segmentation results of human and non-human regions. Based on the rigid body transformation matrix, the segmentation results are mapped to the radar coordinate system to determine the radar grid set corresponding to the non-human region. The average amplitude energy of the radar grid set of the non-human region is calculated to obtain the noise baseline.
5. The method according to claim 1, characterized in that, The prior-driven mapping model established by the weighted reconstruction method based on compressed sensing includes: The magnitude matrix is vectorized, and the vectorized magnitude matrix is expressed as sparse coefficients under a sparse basis based on the first formula, which is: ,in, It is a sparse transformation basis. It is a sparse coefficient vector. This is the vectorized magnitude matrix; Calculate the minimum difference between each radar beam angle in the amplitude matrix and the human body angle in the prior human body data, and define the weighting coefficient of the k-th radar beam angle based on the second formula. The second formula is: = ,in, The minimum difference, These are the angle bandwidth control parameters; A diagonal weight matrix is constructed based on the weighting coefficients for each radar beam angle; A weighted observation equation is constructed based on the third formula, which is: ,in, For weighted observation signals, For radar observation matrix, For noise terms, W is the diagonal weight matrix; The weighted observation equation is solved using an iterative threshold shrinkage algorithm or a fast iterative algorithm to solve the sparse optimization problem and obtain a sparse solution. ; The sparse solution is reconstructed through the inverse transformation of the sparse basis to obtain the enhanced amplitude vector. ; For the amplitude vector An inverse vectorization operation is performed to obtain an enhanced amplitude matrix. The enhanced amplitude matrix is then combined with the original phase matrix to obtain enhanced radar echo data.
6. The method according to claim 1, characterized in that, The prior-driven mapping model based on deep network fusion model includes: A priori-driven mapping model is established based on a deep network fusion model. The amplitude matrix and phase matrix are input as radar branches into the radar encoder of the priori-driven mapping model to obtain radar features. The human body prior data is input as an auxiliary branch into the prior encoder of the priori-driven mapping model to obtain prior features. Feature fusion is performed on the radar features and the prior features to obtain joint features. The joint features are input into the decoder to output the enhanced amplitude, and combined with the original phase to obtain the enhanced radar echo data.
7. The method according to claim 1, characterized in that, Applying constant false alarm rate (CFAR) detection to the enhanced radar echo data amplitude matrix generates a binary detection map, including: A constant false alarm rate (CFAR) detection algorithm is applied to the amplitude matrix of the enhanced radar echo data. The amplitude values of each radar grid point are compared with the background noise threshold to mark the regions where target objects reflect, and a binary detection map is generated.
8. The method according to claim 7, characterized in that, Kalman filtering is used to track the target object, and the motion trajectory information of the target object is obtained, including: Based on the binary detection map, adjacent detection response points are clustered to obtain candidate target clusters. The candidate target clusters are used as observation inputs to extract the center position, amplitude energy, and motion trend features of the target object, construct the observation vector of the target object, and use the Kalman filter algorithm to recursively estimate the state parameters of the target object. The state parameters include distance, radial velocity, and angle information. The state parameters are converted into position information in a plane coordinate system and combined with the velocity vector to form the motion trajectory information of the target object.
9. A millimeter-wave signal reconstruction system for indoor monitoring, used to implement the millimeter-wave signal reconstruction method for indoor monitoring as described in any one of claims 1-8, characterized in that, The system includes: The acquisition module acquires video frame data of target objects in the indoor area and raw I / Q data collected by millimeter-wave radar. It extracts key feature data of the target objects from the video frame data, performs spatiotemporal alignment on the video frame data and raw I / Q data, and maps them to the radar coordinate system using a rigid body transformation matrix to form prior human information. The preprocessing module preprocesses the raw I / Q data to obtain an amplitude matrix and a phase matrix. Based on the non-target region of the amplitude matrix, it calculates the average energy to obtain a noise baseline. Based on the noise baseline, it normalizes the amplitude matrix and combines it with the phase matrix to generate radar echo data. The enhancement module establishes a priori-driven mapping model based on a weighted reconstruction method of compressed sensing or a deep network fusion model, and performs enhancement processing on the radar echo data based on the priori-driven mapping model to obtain enhanced radar echo data. The generation module applies constant false alarm rate detection to the enhanced radar echo data amplitude matrix to generate a binary detection map, and uses Kalman filtering to track the target object and obtain the target object's motion trajectory information.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements a millimeter-wave signal reconstruction method for indoor monitoring as described in any one of claims 1-8.