Intermediate frequency signal phase demodulation method under low signal-to-noise ratio condition
Through VSPnet learning phase mapping function and using Gram transformation and U-Net network structure for phase demodulation, the problem of inaccurate phase demodulation under low signal-to-noise ratio conditions is solved, and high precision and noise anti-noise performance are improved.
Patent Information
- Application Number
- CN202510105128.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
AI Technical Summary
Under low signal-to-noise ratio conditions, existing intermediate frequency signal phase demodulation methods are susceptible to noise interference, resulting in dewinding errors and inaccurate phase demodulation results.
VSPnet is used to learn and establish a mapping function between the transition phase and the real phase. The one-dimensional phase signal is expanded into a two-dimensional matrix through Gram transformation, and phase demodulation is performed using the asymmetric U-Net network structure and the Swin-Transformer module.
It significantly improves noise resistance and phase demodulation accuracy, reduces network computing volume, and is significantly better than traditional phase difference algorithms and U-Net methods in low signal-to-noise ratio environments.
Smart Images

Figure CN119939134A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of non-contact vital sign detection technology of millimeter-wave radar, specifically relating to a method for phase demodulation of intermediate frequency signals under low signal-to-noise ratio conditions. Background Technology
[0002] Respiratory rate, heart rate, and other physiological indicators are fundamental vital signs of human life, directly reflecting a person's health status and aiding medical personnel in making accurate diagnoses. In recent years, with the high incidence of heart disease, people have paid particular attention to heart health, and heart rate is a crucial indicator of cardiac activity. In most cases, heart rate measurement methods utilize devices that directly contact the body, such as ECG electrodes or photoelectric volumetric pulse wave sensors. While these methods provide relatively reliable results, they can still cause discomfort and problems for users, and may even lead to severe allergic reactions.
[0003] In millimeter-wave radar for vital sign detection, the minute vibrations caused by a subject's heartbeat and respiration are recorded in the phase of the radar's intermediate frequency signal. Given the noise interference present in the environment and equipment, accurate phase extraction is crucial. Firstly, the micro-movements of the chest wall caused by respiration and heartbeat are mostly on the millimeter scale, easily obscured by background clutter and receiver thermal noise. Therefore, accurately extracting these minute phase changes under noise interference is key. Phase demodulation technology is used to recover the original phase information from the measured phase, especially in the presence of 2π ambiguity. Commonly used phase demodulation algorithms include phase difference algorithms, path unrolling algorithms, phase gradient unrolling algorithms, and U-Net methods. However, these traditional phase demodulation algorithms are sensitive to noise, and in noisy conditions, they are prone to dewinding errors, especially when processing complex signals.
[0004] The above background technology illustrates that existing phase demodulation methods for mid-frequency signals of vital signs still have shortcomings at low signal-to-noise ratios, and it is necessary to research and propose more reliable phase demodulation methods. Summary of the Invention
[0005] This invention provides a phase demodulation method for intermediate frequency signals under low signal-to-noise ratio conditions. A mapping function between the phase jump and the true phase is learned and established using VSPnet. This overcomes the limitations of traditional phase dewinding algorithms in handling noise and complex signals, achieving better noise immunity and more accurate phase demodulation.
[0006] The present invention provides a method for phase demodulation of intermediate frequency signals under low signal-to-noise ratio conditions, comprising the following steps:
[0007] 1.1 The mixer mixes the received human chest cavity reflection signal with the transmitted signal, and then samples it to obtain an intermediate frequency signal A[m,n] containing static interference and noise;
[0008] Where: n is the fast time, n = 1, 2, ..., N, and m is the slow time, m = 1, 2, ..., M;
[0009] 1.2 Eliminating static interference includes the following steps:
[0010] 1.2.1 Calculate the average value of the intermediate frequency signal along the fast time. As an estimate of static clutter;
[0011] 1.2.2 Subtract the estimated static clutter from the intermediate frequency signal to obtain the intermediate frequency signal with static interference removed.
[0012] 1.3 pairs The distance-dimensional Fourier transform is expressed as follows:
[0013]
[0014] Where: frequency u∈[1,N];
[0015] Accumulate F[m,u] over slow time Take the frequency component with the highest energy Phase extraction yields the phase jump containing information about the displacement of the human thoracic cavity. The expression is:
[0016]
[0017] Where: Im(F[m,U]) represents the imaginary part; Re(F[m,U]) represents the real part;
[0018] 1.4 Constructing VSPnet
[0019] VSPnet consists of three parts: Gram transform, phase demodulation network, and signal reconstruction; VSPnet first performs phase transformation on the input. Performing a Gram transformation yields a two-dimensional matrix G, where the element located in the i-th row and j-th column is:
[0020]
[0021] Where: i,j∈[1,M];
[0022] The phase demodulation network adopts an asymmetric U-Net network structure, which includes an encoding part, a decoding part, and a feature fusion skip connection; the signal reconstruction part is the output matrix of the phase demodulation network. Extracting the main diagonal elements yields the demodulation phase for denoising. in: Representation matrix The main diagonal elements;
[0023] 1.5 Encoding Section
[0024] It contains five encoding modules, each consisting of three convolutional layers and one Swin-Transformer module. Each convolutional layer comprises a 3×3 kernel with a stride of 2 and padding of 1, batch normalization, and a ReLU activation function. A shortcut connection is used between the first and third convolutional layers. The feature maps extracted by the convolutional layers are input into the Swin-Transformer module for layer normalization and multi-head window self-attention. The window size for the first and second encoding modules is 8×8, while the window size for the other three modules is 4×4. The multi-head window self-attention and shifted window multi-head self-attention processes include the following steps:
[0025] 1.5.1 For the input feature map G i We divide the window, where i∈[1,2,3,4,5] is the index of the encoding module, and obtain N. w A small window G i,j Window index j∈[1,2,...,N] w ];
[0026] 1.5.2 For each small window G i,j Perform multi-head self-attention, G i,j After query matrix Key matrix Sum matrix get Head indices h∈[1,2,3,4], each head undergoes individual self-attention computation to obtain enhanced features. The outputs of all heads are concatenated along the channels to obtain the multi-head self-attention output of a small window, expressed as:
[0027]
[0028] 1.5.3 The multi-head self-attention outputs of all small windows are merged to obtain G. i Multi-head self-attention output Z i , and G i After residual connections, layer normalization is performed; then a fully connected layer is applied to expand the channel dimension by 4 times; after passing through the GELU activation function, a fully connected layer is applied to compress the channel dimension; finally, G is obtained through residual connections. i The multi-head self-attention output of the window is expressed as:
[0029]
[0030] Where: LN is the layer normalization, W1∈R C×4C and W2∈R 4C×C These are the weight matrices of the two fully connected layers;
[0031] 1.5.4 Layer normalization and shifted window multi-head self-attention are performed, with the shift length being half the window size. Steps 1.5.2 and 1.5.3 are repeated for the newly divided window, with the overlapping regions being averaged during concatenation. Finally, the output of the encoding module is obtained.
[0032] 1.6 Decoding Section
[0033] It contains five decoding modules. Each decoding module consists of a dynamic convolutional attention module, a convolutional layer consisting of a 3×3 convolutional kernel with a stride of 2 and padding size of 1, batch normalization, and a ReLU activation function, and a transposed convolutional layer consisting of a 3×3 transposed convolution with a stride of 2 and padding size of 1, a ReLU activation function, and batch normalization. The dynamic convolutional attention module includes a dynamic channel attention module and a spatial attention module, specifically including the following steps:
[0034] 1.6.1 The dynamic channel attention module generates dynamic channel weights through global average pooling and MLP, expressed as follows:
[0035]
[0036] Where: s∈R C×1×1 AvgPool uses global average pooling. This is the input to the decoding module; W1∈R C×4C and W2∈R 4C×C These are the weight matrices of the two fully connected layers; σ is the Sigmond activation function.
[0037] 1.6.2 Integrating dynamic channel attention weights with The dynamic channel attention output is obtained by element-wise multiplication along the channel dimension.
[0038]
[0039] 1.6.3 will The input is fed into the spatial attention module, where max pooling and average pooling are performed along the channel dimension to obtain two single-channel feature maps. The two feature maps are then concatenated along the channel dimension to obtain a feature map with two channels.
[0040] 1.6.4 A feature map with two channels is transformed into a single-channel feature map using a 7×7 convolution kernel;
[0041] 1.6.5 Spatial attention weights are obtained through the Sigmond activation function, and these spatial attention weights are then combined with... Element-wise multiplication along the spatial dimension yields
[0042] 1.6.6 After passing through a 3×3 convolutional layer, the output of the decoding module is obtained by upsampling through a transposed convolutional layer.
[0043]
[0044] 1.7 Feature fusion skip connections, including the following steps:
[0045] 1.7.1 Dimension Matching: Using a 1×1 convolution kernel to... The channel is compressed by half, and bilinear interpolation is used for upsampling to double its spatial dimension, resulting in...
[0046] 1.7.2 and By directly adding them together, we obtain the output of the fusion skip connection for the i-th feature, expressed as:
[0047]
[0048] The output of the i-th feature fusion skip connection in the network Output of the 5-i decoding module in the network The data is concatenated along the channel dimension as input to the next decoding module. The input to the first decoding module is...
[0049]
[0050] 1.8 At the end of the network, a 1×1 convolutional kernel is used to compress the number of output channels to obtain...
[0051] The loss function for 1.9VSPnet is:
[0052]
[0053] in: β is the weight used to adjust the output quality, and B is the amount of phase signal used for training. and These are the phase transition and the demodulation phase of the VSPnet output, respectively.
[0054] μ G and It is G and The average value of the elements and It is G and The variance of element values It is G and The covariances of the element values, c1 and c2, are small constants close to zero;
[0055] 1.10 VSPnet Training
[0056] An adaptive moment estimation algorithm is used as the optimizer, and a learning rate scheduler based on monitoring metrics is used for training. When the training loss and validation loss remain relatively stable or tend to be flat in multiple consecutive training iterations, it indicates that the model has converged and the current training effect is optimal.
[0057] This invention expands a one-dimensional phase signal into a two-dimensional matrix through Gram transform, mapping phase information onto the main diagonal of the Gram matrix. Simultaneously, it mines the correlations between heartbeat signals to form the local structure of the signal, enriching the spatial modes of the phase signal and suppressing noise components without fixed modes. This enhances the network's ability to distinguish complex phase signals and noise. The asymmetric phase demodulation network can progressively extract and fuse multi-level features, significantly improving signal feature representation capabilities. This reduces network complexity and improves demodulation accuracy under low signal-to-noise ratio conditions.
[0058] The technical problem solved by this invention is that the phase signal reflecting heartbeat information in millimeter-wave radar vital sign detection systems is easily affected by noise interference. Existing phase demodulation algorithms are prone to accumulated errors and have weak ability to perceive details, resulting in inaccurate phase demodulation results. The method of this invention combines Gram transform to expand the sensitive one-dimensional signal into a two-dimensional signal with rich signal modes, solving the problem of noise sensitivity in millimeter-wave radar phase demodulation. Furthermore, the Swin-Transformer module and dynamic convolutional attention module address the difficulty in characterizing and extracting detailed features of heartbeat signals under low signal-to-noise ratio conditions.
[0059] The beneficial effects of this invention are as follows: Under low signal-to-noise ratio conditions, the performance of this invention is significantly better than that of the phase difference algorithm and the U-Net method. This invention reduces network computation while exhibiting strong noise resistance and high-precision phase demodulation, significantly reducing the deviation in heart rate estimation. Simulation data comparing different methods verifies the significant improvement in demodulation performance of this invention, demonstrating its broad application prospects in the field of millimeter-wave radar vital sign detection. Attached Figure Description
[0060] Figure 1 This is a flowchart of a phase demodulation method for intermediate frequency signals under low signal-to-noise ratio conditions;
[0061] Figure 2 A schematic diagram of VSPnet;
[0062] Figure 3 This is a structural diagram of the phase demodulation network;
[0063] Figure 4 Here is a diagram of the Swin-Transformer architecture;
[0064] Figure 5 This is a structural diagram of the dynamic convolutional attention module;
[0065] Figure 6 For feature fusion skip connection structure diagram;
[0066] Figure 7 This is a simulated respiratory signal diagram;
[0067] Figure 8 This is a simulated cardiac impact signal diagram;
[0068] Figure 9 The diagram shows the phase transition signal to be processed.
[0069] Figure 10 The demodulation result using the phase difference algorithm when SNR = 3dB is shown in the figure.
[0070] Figure 11 The demodulation results using VSPnet when SNR = 3dB are shown in the figure.
[0071] Figure 12 The demodulation results using U-Net are shown when SNR = 3dB.
[0072] Figure 13 The image shows a comparison of the demodulation results of three algorithms—phase difference algorithm, VSPnet, and Unet—and the spectrum of the micro-motion signal in the human chest cavity when SNR = 3dB.
[0073] Figure 14 To estimate the average deviation of heart rate using three algorithms—phase difference algorithm, VSPnet, and Unet—under different signal-to-noise ratios. Detailed Implementation
[0074] The present invention will now be described in conjunction with the accompanying drawings.
[0075] The present invention provides a method for phase demodulation of intermediate frequency signals under low signal-to-noise ratio conditions, comprising the following steps:
[0076] 1.1 The mixer mixes the received human chest cavity reflection signal with the transmitted signal, and then samples it to obtain an intermediate frequency signal A[m,n] containing static interference and noise;
[0077] Where: n is the fast time, n = 1, 2, ..., N, and m is the slow time, m = 1, 2, ..., M;
[0078] 1.2 Eliminating static interference includes the following steps:
[0079] 1.2.1 Calculate the average value of the intermediate frequency signal along the fast time. As an estimate of static clutter;
[0080] 1.2.2 Subtract the estimated static clutter from the intermediate frequency signal to obtain the intermediate frequency signal with static interference removed.
[0081] 1.3 pairs The distance-dimensional Fourier transform is expressed as follows:
[0082]
[0083] Where: frequency u∈[1,N];
[0084] Accumulate F[m,u] over slow time Take the frequency component with the highest energy Phase extraction yields the phase jump containing information about the displacement of the human thoracic cavity. The expression is:
[0085]
[0086] Where: Im(F[m,U]) represents the imaginary part; Re(F[m,U]) represents the real part;
[0087] 1.4 Constructing VSPnet
[0088] VSPnet consists of three parts: Gram transform, phase demodulation network, and signal reconstruction; VSPnet first performs phase transformation on the input. Performing a Gram transformation yields a two-dimensional matrix G, where the element located in the i-th row and j-th column is:
[0089]
[0090] Where: i,j∈[1,M];
[0091] The phase demodulation network adopts an asymmetric U-Net network structure, which includes an encoding part, a decoding part, and a feature fusion skip connection; the signal reconstruction part is the output matrix of the phase demodulation network. Extracting the main diagonal elements yields the demodulation phase for denoising. in: Representation matrix The main diagonal elements;
[0092] 1.5 Encoding Section
[0093] It contains five encoding modules, each consisting of three convolutional layers and one Swin-Transformer module. Each convolutional layer comprises a 3×3 kernel with a stride of 2 and padding of 1, batch normalization, and a ReLU activation function. A shortcut connection is used between the first and third convolutional layers. The feature maps extracted by the convolutional layers are input into the Swin-Transformer module for layer normalization and multi-head window self-attention. The window size for the first and second encoding modules is 8×8, while the window size for the other three modules is 4×4. The multi-head window self-attention and shifted window multi-head self-attention processes include the following steps:
[0094] 1.5.1 For the input feature map G i We divide the window, where i∈[1,2,3,4,5] is the index of the encoding module, and obtain N. w A small window G i,j Window index j∈[1,2,...,N] w ];
[0095] 1.5.2 For each small window G i,j Perform multi-head self-attention, G i,j After query matrix Key matrix Sum matrix get Head indices h∈[1,2,3,4], each head undergoes individual self-attention computation to obtain enhanced features. The outputs of all heads are concatenated along the channels to obtain the multi-head self-attention output of a small window, expressed as:
[0096]
[0097] 1.5.3 The multi-head self-attention outputs of all small windows are merged to obtain G. i Multi-head self-attention output Z i , and G i After residual connections, layer normalization is performed; then a fully connected layer is applied to expand the channel dimension by 4 times; after passing through the GELU activation function, a fully connected layer is applied to compress the channel dimension; finally, G is obtained through residual connections. i The multi-head self-attention output of the window is expressed as:
[0098]
[0099] Where: LN is the layer normalization, W1∈R C×4C and W2∈R 4C×C These are the weight matrices of the two fully connected layers;
[0100] 1.5.4 Layer normalization and shifted window multi-head self-attention are performed, with the shift length being half the window size. Steps 1.5.2 and 1.5.3 are repeated for the newly divided window, with the overlapping regions being averaged during concatenation. Finally, the output of the encoding module is obtained.
[0101] 1.6 Decoding Section
[0102] It contains five decoding modules. Each decoding module consists of a dynamic convolutional attention module, a convolutional layer consisting of a 3×3 convolutional kernel with a stride of 2 and padding size of 1, batch normalization, and a ReLU activation function, and a transposed convolutional layer consisting of a 3×3 transposed convolution with a stride of 2 and padding size of 1, a ReLU activation function, and batch normalization. The dynamic convolutional attention module includes a dynamic channel attention module and a spatial attention module, specifically including the following steps:
[0103] 1.6.1 The dynamic channel attention module generates dynamic channel weights through global average pooling and MLP, expressed as follows:
[0104]
[0105] Where: s∈R C×1×1 AvgPool uses global average pooling. This is the input to the decoding module; W1∈R C×4C and W2∈R 4C×C These are the weight matrices of the two fully connected layers; σ is the Sigmond activation function.
[0106] 1.6.2 Integrating dynamic channel attention weights with The dynamic channel attention output is obtained by element-wise multiplication along the channel dimension.
[0107]
[0108] 1.6.3 will The input is fed into the spatial attention module, where max pooling and average pooling are performed along the channel dimension to obtain two single-channel feature maps. The two feature maps are then concatenated along the channel dimension to obtain a feature map with two channels.
[0109] 1.6.4 A feature map with two channels is transformed into a single-channel feature map using a 7×7 convolution kernel;
[0110] 1.6.5 Spatial attention weights are obtained through the Sigmond activation function, and these spatial attention weights are then combined with... Element-wise multiplication along the spatial dimension yields
[0111] 1.6.6 After passing through a 3×3 convolutional layer, the output of the decoding module is obtained by upsampling through a transposed convolutional layer.
[0112]
[0113] 1.7 Feature fusion skip connections, including the following steps:
[0114] 1.7.1 Dimension Matching: Using a 1×1 convolution kernel to... The channel is compressed by half, and bilinear interpolation is used for upsampling to double its spatial dimension, resulting in...
[0115] 1.7.2 and By directly adding them together, we obtain the output of the fusion skip connection for the i-th feature, expressed as:
[0116]
[0117] The output of the i-th feature fusion skip connection in the network Output of the 5-i decoding module in the network The data is concatenated along the channel dimension as input to the next decoding module. The input to the first decoding module is...
[0118]
[0119] 1.8 At the end of the network, a 1×1 convolutional kernel is used to compress the number of output channels to obtain...
[0120] The loss function for 1.9VSPnet is:
[0121]
[0122] in: β is the weight used to adjust the output quality, and B is the amount of phase signal used for training. and These are the phase transition and the demodulation phase of the VSPnet output, respectively.
[0123] μ G and It is G and The average value of the elements and It is G and The variance of element values It is G and The covariances of the element values, c1 and c2, are small constants close to zero;
[0124] 1.10 VSPnet Training
[0125] An adaptive moment estimation algorithm is used as the optimizer, and a learning rate scheduler based on monitoring metrics is used for training. When the training loss and validation loss remain relatively stable or tend to be flat in multiple consecutive training iterations, it indicates that the model has converged and the current training effect is optimal.
[0126] The processing flow of this invention is as follows: Figure 1 As shown, the mixer mixes the received chest cavity reflection signal with the transmitted signal and then samples it to obtain an intermediate frequency (IF) signal containing static interference and noise. Static clutter is then removed from the IF signal. For the IF signal after static interference removal, a distance-dimensional Fourier transform is performed, and the amplitude of the transform is accumulated over slow time. The frequency component with the highest energy is then used to extract the phase jump caused by chest micro-movement. Will The demodulated phase is obtained by inputting it into VSPnet. The overall structure of VSPnet is as follows: Figure 2 As shown, the phase demodulation network is as follows Figure 3 As shown, the submodule structure is as follows: Figure 4 , Figure 5 , Figure 6 As shown.
[0127] Example:
[0128] The effects of the present invention will be illustrated by simulation experiments below, and the implementation process of the present invention will be further explained in conjunction with the accompanying drawings.
[0129] like Figure 7 As shown, a model of human micro-motion signals is constructed using sinusoidal signals with amplitudes of 0.50–1.00 mm and frequencies of 0.1–0.5 Hz to simulate chest cavity displacement caused by human respiration. Figure 8 As shown, the seven sub-waves in the cardiac impulse signal, including the P wave, R wave, and J wave, are simulated using Gaussian function modeling. 'a' represents the amplitude of the wavelet, 'μ' represents the center position of the wavelet, and 'σ' represents the standard deviation of the wavelet. Wavelets are superimposed on the time axis to accurately simulate human micro-motion signals. The amplitude of the J-peak is set to 0.05–0.10 mm, and the heart rate to 0.8–2.0 Hz. Key parameters such as heart rate, respiratory rate, amplitude of chest cavity micro-motion caused by heartbeat and respiration, and waveform of cardiac impact signal are recorded in Mat files during the modeling of human micro-motion signals. The synthesized human micro-motion signals are then input into a simulated radar system built using the Phased ArraySystem toolbox in MATLAB. The human body is positioned facing the radar at a distance of 1.0 m. Various noises are added to the radar system's receiver, including simulated receiver thermal noise, simulated phase noise, and simulated multipath reflection and random interference environmental noise. Thermal noise and environmental noise are modeled using Gaussian white noise, and phase noise is modeled using random phase jitter, thus realistically reproducing various noise interferences during radar signal reception. Static clutter is first filtered out from the intermediate frequency signal obtained by the simulated radar system, and then the noisy phase jump caused by the micro-movement of the human chest cavity is obtained.
[0130] A total of 15,000 pairs of human micro-motion signals and noisy phase transitions of length 128 were generated and used as the label and training set for VSPnet, respectively. The experimental platform of this invention uses a 12th generation Intel Core i7-12700KF CPU and an NVIDIA GeForce RTX 4080 GPU, and the programming language is Python. The noisy phase transitions to be processed are input into VSPnet for processing. VSPnet encodes the noisy phase transitions through Gram transform to obtain the corresponding two-dimensional matrix, which is used as the input to the phase demodulation network. After obtaining the output of the phase demodulation network, the demodulated phase is obtained through signal reconstruction.
[0131] like Figure 8 As shown, the simulated subject was positioned 1.0m from the radar, with a respiratory rate of 0.3Hz and a heart rate of 1.5Hz. Noise was added to achieve a signal-to-noise ratio of 3dB. After demodulating the extracted transition phase using a phase difference algorithm, it can be seen that the demodulated signal was severely affected by noise. Further measurements of the respiratory and heart rates yielded a respiratory rate of 17.4 beats / minute and a heart rate of 72.2 beats / minute. Figure 9 As shown, after demodulating the extracted transition phase using the VSPnet algorithm, it can be seen that VSPnet not only recovers the respiratory waveform well, but also recovers the subtle waveforms of the heartbeat component relatively well. Further calculation of the respiratory and heart rate yields a respiratory rate of 17.4 breaths / minute and a heart rate of 91.3 beats / minute. Figure 10As shown, after demodulating the extracted transition phase using U-Net, it can be seen that the waveform demodulated by the U-Net algorithm exhibits irregular disturbances, and the heartbeat component in the signal is difficult to distinguish. Further measurements of the respiratory and heart rate yielded a respiratory rate of 18.8 beats / minute and a heart rate of 74.7 beats / minute. The spectra of the demodulated signals from the three algorithms are shown below. Figure 11 As shown, when the received signal-to-noise ratio (SNR) is 3dB, the respiratory rate estimated by the phase difference algorithm and the VSPnet method is consistent with the reference value, while the U-Net method shows a deviation in its estimation of the respiratory rate. Similarly, the heart rate estimated by the VSPnet method is consistent with the reference value, while the U-Net method and the phase difference algorithm show significant deviations in their estimations. Generally, in millimeter-wave radar vital sign detection, a deviation of no more than two times per minute between the estimated and reference values of heart rate measurement is considered accurate. Multiple tests were conducted on the noisy phase transition under different SNR conditions using the three algorithms described above, combined with… Figure 14 As can be seen, in an environment with a signal-to-noise ratio of 3dB to 10dB, VSPnet maintains a relatively low level of estimation bias for heart rate, which is significantly better than the other two methods under the same conditions.
Claims
1. A method for phase demodulation of an intermediate frequency signal under low signal-to-noise ratio conditions, characterized in that it comprises the following steps: 1.1 The mixer mixes the received chest reflection signal with the transmitted signal, and then samples it to obtain the intermediate frequency signal A[m,n] containing static interference and noise; Where: n is the fast time, n = 1, 2..., N, m is the slow time, m = 1, 2..., M; 1.2 Eliminate static interference, including the following steps: 1.2.1 Calculate the average value of the intermediate frequency signal along the fast time as estimated static clutter; 1.2.2 Subtract the estimated static clutter from the intermediate frequency signal to obtain the intermediate frequency signal with static interference removed 1.3 pairs Do the distance dimension Fourier transform, the expression is: Where: frequency u∈[1,N]; Accumulate F[m,u] along slow time Take the frequency component with the largest energy Extract the phase and obtain the jump phase containing the human chest displacement information The expression is: Among them: Im(F[m,U]) represents the imaginary part; Re(F[m,U]) represents the real part; 1.4 Constructing VSPnet VSPnet consists of three parts: Gram transform, phase demodulation network and signal reconstruction. VSPnet first performs phase change analysis on the input. Performing Gram transformation, we get a two-dimensional matrix G, where the element in the i-th row and j-th column is: Where: i,j∈[1,M]; The phase demodulation network adopts an asymmetric U-Net network structure, which includes an encoding part, a decoding part and a feature fusion jump connection; the signal reconstruction part is the output matrix of the phase demodulation network The main diagonal elements of are extracted to obtain the denoised demodulated phase in: Representation Matrix The main diagonal elements of ; 1.5 Coding It contains five encoding modules, each of which contains three convolutional layers and a Swin-Transformer module. The convolutional layer consists of a 3×3 convolution kernel with a step size of 2 and a padding size of 1, batch normalization, and ReLU activation function. A shortcut connection is used between the first and third convolutional layers. The feature map extracted by the convolutional layer is input into the Swin-Transformer module for layer normalization and window multi-head self-attention. The window size of the first and second encoding modules is 8×8, and the window size of the other three encoding modules is 4×4. Among them, the window multi-head self-attention and shifted window multi-head self-attention include the following steps: 1.5.1 Input feature map G i Perform window division, i∈[1,2,3,4,5] is the encoding module index, and get N w Small window G i,j , window index j∈[1,2,...,N w ]; 1.5.2 For each small window G i,j Perform multi-head self-attention, G i,j After querying the matrix Key Matrix Sum Matrix get Head index h∈[1,2,3,4], each head performs a separate self-attention calculation to obtain enhanced features The outputs of all heads are concatenated along the channel to obtain the multi-head self-attention output of a small window, expressed as: 1.5.3 All the multi-head self-attention outputs of small windows are combined to obtain G i The multi-head self-attention output Z i , and G i After the residual connection, the layer is normalized; then the channel dimension is expanded by 4 times through the fully connected layer, and the channel dimension is compressed through the fully connected layer after the GELU activation function, and finally G is obtained through the residual connection. i The window multi-head self-attention output is expressed as: in: LN is layer normalization, W1∈R C×4C and W2∈R 4C×C They are the weight matrices of the two fully connected layers; 1.5.4 Pair Perform layer normalization and shift window multi-head self-attention, with the shift length being half the window size. Repeat steps 1.5.2 and 1.5.3 for the newly divided window, averaging the repeatedly calculated areas during splicing, and finally obtain the output of the encoding module. 1.6 Decoding part It contains five decoding modules, each of which consists of a dynamic convolutional attention module, a convolutional layer consisting of a 3×3, stride 2, padding size 1 convolution kernel, batch normalization, ReLU activation function, and a transposed convolutional layer consisting of a 3×3, stride 2, padding size 1 transposed convolution, ReLU activation function and batch normalization; the dynamic convolutional attention module contains a dynamic channel attention module and a spatial attention module, which specifically includes the following steps: 1.6.1 Dynamic channel attention module generates dynamic channel weights through global average pooling and MLP, expressed as: in: s∈R C×1×1 ; AvgPool is the global average pooling; is the input of the decoding module; W1∈R C×4C and W2∈R 4C×C are the weight matrices of the two fully connected layers respectively; σ is the Sigmond activation function; 1.6.2 Combine dynamic channel attention weights with Multiply element-wise along the channel dimension to get the dynamic channel attention output 1.6.3 The data is input into the spatial attention module, and the maximum pooling and average pooling are performed along the channel dimension to obtain two single-channel feature maps. The two feature maps are concatenated along the channel dimension to obtain a feature map with 2 channels. 1.6.4 The feature map with 2 channels is transformed into a single-channel feature map through a 7×7 convolution kernel; 1.6.5 The spatial attention weight is obtained through the Sigmond activation function, and the spatial attention weight is combined with Multiply element-wise along the spatial dimension to get 1.6.6 After passing through a 3×3 convolutional layer, it is upsampled through a transposed convolutional layer to obtain the output of the decoding module. 1.7 Feature fusion jump connection, including the following steps: 1.7.1 Dimension Matching: Use a 1×1 convolution kernel to The channel is compressed by half and upsampled using bilinear interpolation to double the spatial dimension, resulting in 1.7.2 and Directly add them together to get the output of the i-th feature fusion jump connection, which is expressed as: The output of the i-th feature fusion skip connection in the network And the output of the 5-ith decoding module in the network The concatenation along the channel dimension is used as the input of the next decoding module. In addition, the input of the first decoding module is 1.8 At the end of the network, a 1×1 convolution kernel is used to compress the number of output channels to obtain 1.9 The loss function of VSPnet is: in: β is the weight for adjusting the output quality, B is the number of phase signals used for training, and They are the jump phase and the demodulation phase of VSPnet output respectively; μ G and It is G and The average value of the elements, and It is G and The variance of the element values, It is G and The covariance of the element values, c1 and c2 are small constants close to zero; 1.10 VSPnet Training The adaptive moment estimation algorithm is used as the optimizer, and the learning rate scheduler based on monitoring indicators is used for training. When the training loss and validation loss keep small fluctuations or tend to be stable in multiple consecutive training iterations, it means that the model converges and the current training effect is the best.
Citation Information
Cited By
Method for repairing sound quality, electronic equipment, storage medium and program product
CN120510857A