Non-contact human respiration rate measurement method based on video and frequency modulated continuous wave radar information fusion

By fusing dual-mode data from video and frequency-modulated continuous wave radar, and utilizing multivariate singular spectrum analysis and signal-to-noise ratio weighting, the accuracy problem of respiratory rate monitoring under spontaneous large-amplitude body movements was solved, achieving high-precision respiratory rate measurement in multiple scenarios.

CN116530967BActive Publication Date: 2025-12-19HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310510531.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-12-19
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Existing non-contact respiratory rate monitoring technologies have low measurement accuracy in spontaneous large-amplitude body movements, and cannot work effectively, especially in poor lighting conditions. Furthermore, existing modal fusion methods lack generalization ability across multiple scenarios.

Method used

A dual-mode data fusion method based on video and frequency-modulated continuous wave radar is adopted. Through feature-level and decision-level fusion, multivariate singular spectrum analysis and signal-to-noise ratio weighting are used to achieve robust measurement of respiratory signals.

Benefits of technology

It improves the accuracy and robustness of respiratory rate measurement in various motion scenarios, expanding the application range of consumer-grade cameras and frequency-modulated continuous wave radar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116530967B_ABST
    Figure CN116530967B_ABST
Patent Text Reader

Abstract

The application discloses a kind of non-contact human respiratory rate measurement methods based on video and frequency modulation continuous wave radar information fusion, comprising:1 respectively processes video and radar bimodal data to obtain pixel motion trajectory and distance angle chart time sequence;2 video measurement part uses spectral subtraction and principal component analysis technique to carry out pretreatment, radar measurement part uses static clutter removal and average filtering technique to carry out pretreatment, two kinds of modal all use empirical mode decomposition technique to carry out single modal measurement;3 feature level fusion, after bimodal signal using multivariate singular spectrum analysis extraction shared respiratory signal is pretreated;4 decision level fusion, according to the signal-to-noise ratio weighted frequency domain estimation value summation of feature level fusion result and single modal measurement result obtains final measurement result.The application can measure respiration under a variety of spontaneous motion scene, measurement result is compared with single modal method and has significant improvement, can effectively expand the application range of video and radar fusion measurement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of non-contact physiological signal detection and analysis, and particularly relates to a non-contact respiration rate measurement method based on video and frequency-modulated continuous wave radar information fusion. BACKGROUND

[0002] Respiration rate (RR) is one of the most important vital signs for assessing the physical and mental state of a human being. The normal range of RR for a healthy adult is from 12 breaths per minute (bpm) to 24 bpm. An RR value outside the normal range is usually closely related to some diseases, such as apnea, congestive heart failure or cardiac arrest, and therefore it is of great significance to monitor the RR value. The traditional method of respiration monitoring is to use a contact sensor such as a respiratory belt or a respiratory mask to measure the respiration closely attached to the skin, which is easy to make the subject feel uncomfortable and also causes interference to some activities of the subject. Therefore, non-contact RR monitoring technology has attracted more and more attention. Among various non-contact respiration monitoring methods, the method of monitoring the respiration micro-movement based on computer vision or radar has become the mainstream, among which the monitoring method based on a consumer-level camera has become a popular research direction due to its universality and low cost, and the monitoring method based on frequency-modulated continuous wave (FMCW) radar has also attracted more and more attention due to its illumination robustness and high detection accuracy.

[0003] The video-based respiration measurement method measures the respiration activity of a human being using computer vision technology to cause the fluctuation of the shoulder and chest. In the video, such changes are reflected in the displacement and intensity change of the corresponding region of the pixel points. Common methods include optical flow method, Profile cross-correlation, photoplethysmography (PPG), etc. Compared with the video method, the RR measurement based on FMCW radar is not affected by the illumination condition. The respiration measurement method based on FMCW radar measures the phase change caused by the modulation of the radar echo signal by the human respiration movement according to the Doppler effect. It mainly relies on two kinds of images: range slow-time image and range-angle image. However, most of the current methods assume that the subject is in a static state or there is random body movement. In the case of spontaneous large-amplitude body movement, the respiration signal is severely contaminated, and the existing measurement methods may fail. In recent years, multi-modal fusion measurement has attracted more and more attention. Existing works include camera-assisted radar for motion denoising, however, this method cannot work normally in poor illumination conditions. In addition, there are also decision fusion strategies based on statistical methods, such as Bayesian estimation and maximum likelihood estimation. These model-based methods can well remove noise, but they are only limited to single scene use and have poor generalization ability. Therefore, it is very challenging to measure respiration in various spontaneous movement scenarios. SUMMARY

[0004] The present application is to solve the above technical deficiencies, provide a kind of non-contact human respiratory rate measurement method based on video and frequency modulation continuous wave radar information fusion, to be able to realize more robust breathing measurement in a variety of spontaneous motion scene, so as to improve measurement accuracy.

[0005] The present application is to solve the above technical deficiencies, provide a kind of non-contact human respiratory rate measurement method based on video and frequency modulation continuous wave radar information fusion, to be able to realize more robust breathing measurement in a variety of spontaneous motion scene, so as to improve measurement accuracy.

[0006] The present application is to solve the above technical deficiencies, provide a kind of non-contact human respiratory rate measurement method based on video and frequency modulation continuous wave radar information fusion, to be able to realize more robust breathing measurement in a variety of spontaneous motion scene, so as to improve measurement accuracy.

[0007] Step one, dual-mode data generation;

[0008] Step 1.1, video measurement part;

[0009] The feature point motion trajectory time series of left and right two shoulder regions of interest from the J frame video image of the subject are extracted;Wherein, the coordinate information of the N feature points of the j frame shoulder region of interest is denoted as Wherein, The coordinate information of the n feature point of the shoulder region of interest in the j frame video image is denoted as j∈[2,J];

[0010] The vertical pixel motion trajectory And horizontal pixel motion trajectory N=1,…,N;Wherein, And The horizontal coordinate and vertical coordinate of the n feature point of the j frame shoulder region of interest are respectively denoted as

[0011] Step 1.2, radar measurement part;

[0012] The distance-angle graph time series {RM t ,t=1,2,…,T} in the detection region of the subject T frame radar original data is obtained;Wherein, RM t Denotes the distance-angle graph of the t frame;

[0013] Step two, dual-mode data preprocessing and single mode measurement;

[0014] Step 2.1, video measurement part;

[0015] Step 2.1.1, after band-pass filtering to vertical pixel motion trajectory Y(n) and horizontal pixel motion trajectory X(n), respectively, then using fast Fourier transform to calculate the frequency spectrum f X(n) Of X(n), the frequency spectrum f Y(n) Of Y(n), so as to obtain the residual frequency spectrum f' of denoising vertical pixel motion trajectory Y(n) by using spectrum subtractionY(n) If f' Y(n) If the amplitude is negative, then let f' Y(n) =0; otherwise, use the inverse Fast Fourier Transform to transform f' Y(n) Converted into a time-domain signal Y'(n);

[0016] Step 2.1.2: Use principal component analysis to extract the waveforms in Y'(n) whose cumulative contribution rate is greater than the threshold and sum them in the time domain to obtain the summation results of the left and right shoulders, which are denoted as X(1) and X(2) respectively.

[0017] Step 2.1.3: Decompose the waveform using the complete set empirical mode decomposition method of serial adaptive noise to obtain several intrinsic mode functions; select several intrinsic mode functions with the highest signal-to-noise ratio and use fast Fourier transform to obtain their dominant frequencies respectively; finally select the dominant frequency that appears most frequently and denote it as RR1, and its corresponding signal-to-noise ratio is denoteed as SNR1.

[0018] Step 2.2, Radar Measurement Section;

[0019] Step 2.2.1: Use the static clutter removal method to process the time series of the range-angle map {RM}. t The range-angle map {RM'} of frame T after static clutter removal is obtained by processing the data in the range {t = 1, 2, ..., T}. t ,t=1,2,…,T};where,RM' t This is the distance-angle diagram for the t-th frame after static clutter removal;

[0020] Step 2.2.2: Use a constant false alarm rate detector to check {RM' t The target location coordinates are obtained by processing the set {(PX, t=1,2,…,T}. t ,PY t ),t=1,2,…,T},(PX t ,PY t ) represents RM' t The set of target location coordinates;

[0021] Let the target position coordinate set {(PX t ,PY t The set of amplitudes corresponding to t = 1, 2, ..., T is: in, (PX) t ,PY t The corresponding amplitude set;

[0022] Step 2.2.3, RM' t Target location coordinate set (PX) t ,PYt For each coordinate point within the range, take its two adjacent points along the distance direction, and then take the average of the three coordinate points as the new amplitude for each coordinate point, thus obtaining (PX). t ,PY t The corresponding new amplitude set This leads to a new distance-angle map RM. t ;

[0023] Step 2.2.4: Based on the new amplitude set Calculate the new distance-angle map RM” t The phase value of each target position in the T frames is obtained by concatenating the phase values ​​of the i-th target position frame by frame to obtain the i-th segment of phase change signal. I segments of phase change signal were acquired. in, This represents the phase value of the i-th target position in the t-th frame, and each phase change signal contains T frames;

[0024] Step 2.2.5: Phase unwinding is used to reduce the phase value of each frame to the range of -π to π. Then, differential, smoothing, and bandpass filtering are used to process the reduced phase value to obtain the noise-removed I-segment phase change signal.

[0025] Step 2.2.6: Select the two segments with the highest signal-to-noise ratio in the processed I-segment phase change signal and denote them as X(3) and X(4);

[0026] Step 2.2.7: Process the noise-removed I-segment phase change signal according to the process in step 2.1.3 to obtain the main frequency RR6 and its corresponding signal-to-noise ratio SNR6;

[0027] Step 3: Feature-level fusion;

[0028] Step 3.1: After normalizing the video preprocessing results X(1) and X(2) with the radar preprocessing results X(3) and X(4), four normalized input signals are obtained, and the r-th input signal is denoted as . in, Let F represent the value of the r-th input signal at the F-th point, where F represents the sample size, and r = 1, 2, 3, 4;

[0029] Step 3.2: Let M be the length of the embedded window, and Embedding the r-th input signal X'(r) into the Hankel trajectory matrix yields the r-th embedded Hankel trajectory matrix H of dimension M×K. (r) Where K represents the number of subsequences, and K = F - M + 1;

[0030] All Hankel trajectory matrices {H (r) , r = 1, 2, …, 4} embedded are embedded into a vertical form H of dimension M x K x 4;

[0031] Step 3.3, singular value decomposition is performed on H to obtain 4 x M eigenvalues λ 1, λ 2, …, λ e , …, λ 4M and their corresponding eigenvectors U 1, U 2, …, U e , …, U 4M ; wherein λ e represents the e th eigenvalue, U e represents the eigenvector corresponding to the e th eigenvalue λ e .

[0032] Suppose the rank of H is k, calculate the right singular vector corresponding to the g th eigenvalue λ g and calculate the g th dimension reduction matrix g = 1, 2, …, k, the rank of H g is 1, and the dimension of H g is 4M x K;

[0033] Step 3.4, convert H g into the g th respiratory pulse signal , thereby obtaining k respiratory pulse signals , wherein represents the amplitude of the g th respiratory pulse signal at the F th point;

[0034] Divide the k respiratory pulse signals into , wherein represents the w th respiratory pulse signal; represents the operation of taking the integer part of k / 4 downwards;

[0035] Select the first group of respiratory pulse signals as the feature level fusion result;

[0036] Step 3.5, use fast Fourier transform to calculate the main frequency RR2 of , the main frequency RR3 of , the main frequency RR4 of , the main frequency RR5 of , the signal-to-noise ratio SNR2 of , the signal-to-noise ratio SNR3 of , the signal-to-noise ratio SNR4 of , and the signal-to-noise ratio SNR5 of ​The signal-to-noise ratio SNR5 of the fourth breathing signal;

[0037] Step four, decision level fusion;

[0038] The signal-to-noise ratios {SNR s After summing up, the sum of the signal-to-noise ratios is denoted as SNR sum ; wherein SNR s represents the signal-to-noise ratio corresponding to the s-th breathing signal;

[0039] For the main frequency RR s corresponding to the s-th breathing signal, the weight value corresponding to the s-th breathing signal is SNR s / SNR sum , s = 1, 2, …, 6; thereby calculating the two-level fusion result

[0040] The electronic device comprises a memory and a processor, and is characterized in that the memory is used to store a program supporting the processor to execute the non-contact human respiratory rate measurement method, and the processor is configured to execute the program stored in the memory.

[0041] The computer readable storage medium stores a computer program, and the computer program is executed by the processor to execute the steps of the non-contact human respiratory rate measurement method.

[0042] Compared with the prior art, the beneficial effects of the present application are reflected in:

[0043] 1. The present application proposes a motion robust non-contact RR measurement method with two-level fusion, in the feature fusion level, the multiple singular spectrum analysis (MSSA) method is introduced as the feature fusion strategy, the method of feature level fusion measurement of breathing by combining the frequency modulated continuous wave radar and the camera is first proposed, the problem of inter-modal mismatch is effectively solved, and the robustness of the RR measurement result is improved;

[0044] 2. In the decision fusion level, the results of feature level fusion of the frequency modulated continuous wave radar and the camera and the single modal result are jointly used as the input of decision level fusion, and all RR results are weighted and summed according to the signal-to-noise ratio (SNR) of the breathing signal, thereby further improving the reliability of the target RR value;

[0045] 3. The present application evaluates the proposed RR fusion measurement method in various motion scenes. The experimental results prove the robustness of the RR measurement in multiple scenes, which effectively expands the application range of the fusion of the consumer-level camera and the frequency modulated continuous wave radar in the RR measurement. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 This is a flowchart of the method of the present invention;

[0047] Figure 2 This is a schematic diagram illustrating the location of the region of interest on the shoulder in a video frame, as described in this invention.

[0048] Figure 3 This is a comparison diagram of the distance angle diagrams generated and processed by this invention;

[0049] Figure 4a This is a comparison chart of the Bland-Altman analysis of four comparison methods in the speaking scenario of this invention;

[0050] Figure 4b This is a comparison chart of the Bland-Altman analysis of four comparison methods in a walking scenario according to the present invention;

[0051] Figure 4c This is a comparison chart of the Bland-Altman analysis of four comparison methods in the running scenario according to the present invention;

[0052] Figure 4d This is a comparison chart of the Bland-Altman analysis of four comparison methods under changing scenarios according to the present invention;

[0053] Figure 5 This is a box plot of the root mean square error at each fusion stage of the present invention;

[0054] Figure 6 This is a line graph showing the root mean square error and average Pearson coefficient γ of the present invention under different window lengths. Detailed Implementation

[0055] In this embodiment, a non-contact human respiratory rate measurement method based on the fusion of video and FM continuous wave radar information involves denoising the video and FM continuous wave radar signals separately, then fusing their features using a multivariate singular spectrum analysis method, followed by decision fusion based on signal-to-noise ratio weighting. The specific process is as follows: Figure 1 As shown, it includes the following steps:

[0056] Step 1: Generation of dual-modal data;

[0057] Step 1.1, Video Measurement Section;

[0058] First, J frames of video images are acquired, and the bounding rectangle of the face is located using the Viola-Jones face detector. Next, based on the geometric relationship between the human face and shoulders, the distance between the subject and the camera, and the video resolution, the regions of interest (ROIs) of both shoulders in the first frame of the video image are located, and the coordinate information of N feature points in each shoulder ROI is determined using the Good Feature to Track technique. in, The coordinates of the j-th feature point in a single shoulder region of interest in the first frame of the video image are given, where j ∈ [2, J]. Then, the Kanade Lucas-Tomasi (KLT) algorithm is used to track the above N feature points.

[0059] Vertical pixel motion trajectories are generated from the coordinates of the nth feature point in the shoulder region of interest in frame J. and horizontal pixel motion trajectory n = 1, ..., N; where, and These represent the x and y coordinates of the nth feature point in the region of interest (ROI) of the shoulder in frame j, respectively; the ROI defined in the first frame of the video is as follows: Figure 2 As shown.

[0060] Step 1.2, Radar Measurement Section;

[0061] First, the raw binary data of the radar for T frames is acquired. Then, the channel parsing method is used to acquire the data from the first frame to the Tth frame, and the in-phase channel and quadrature channel data are merged into a complex signal. Then, the first frame data is extracted and reconstructed to obtain a three-dimensional data block V1 of size n_Rx*n_samples*n_chirps. The Tth frame data is then acquired using the same method. t The data blocks of a frame are Vt, and a total of T frames of data blocks {V} are ultimately obtained. t ,t=1,2,…,T}. Where, n_Rx represents the number of receiving antennas, n_samples represents the number of sampling points per chirp, and n_chirps represents the number of chirps per frame;

[0062] Next, for the first frame of 3D data block V1, an intermediate point is selected along the n_chirps dimension. A cross-section containing this point and perpendicular to the n_chirps dimension is taken, with a size of n_Rx*n_samples, denoted as MP1. Then, a Fast Fourier Transform is applied along the n_samples dimension to coherently accumulate MP1. Then, a 180-point Fast Fourier Transform is applied along the n_Rx dimension to MP1 to obtain the first frame's distance angle map, denoted as RM1. The distance angle map of the t-th frame is obtained using the same method, denoted as RM. t Finally, a total of T frames of distance-angle maps {RM} were obtained. t ,i=1,2,…,T}.

[0063] Step 2: Dual-modal data preprocessing and single-modal measurement;

[0064] Step 2.1, Video Measurement Section;

[0065] Step 2.1.1, band-pass filter the vertical pixel motion trajectory Y(n) and the horizontal pixel motion trajectory X(n) respectively using a 0.2-0.5 Hz Butterworth filter, then calculate the frequency spectrum f X(n) of X(n) and the frequency spectrum f Y(n) of Y(n) using fast Fourier transform respectively, subsequently, obtain the residual frequency spectrum f Y(n) of the de-noised vertical pixel motion trajectory Y(n) using spectral subtraction; Y(n) if the amplitude of f Y(n) is negative, set f Y(n) =0; if not, directly convert f t into a time domain signal Y'(n) using inverse fast Fourier transform;

[0066] Step 2.1.2, extract the waveform whose cumulative contribution rate is greater than 95% in Y'(n) using principal component analysis and sum in time domain, obtain the sum results of the left and right shoulders respectively, denoted as C(1) and C(2);

[0067] Step 2.1.3, decompose C(1) and C(2) using the complete ensemble empirical mode decomposition method of serial adaptive noise to obtain a plurality of intrinsic mode functions, select a plurality of intrinsic mode functions with the highest signal-to-noise ratio, use fast Fourier transform to calculate the main frequency of each intrinsic mode function, and finally select the main frequency with the highest occurrence frequency as RR1, and the corresponding signal-to-noise ratio as SNR1.

[0068] Step 2.2, radar measurement part;

[0069] Step 2.2.1, first, use a static clutter removal method to subtract the amplitude of the corresponding points of the second frame range-angle map RM2 and the first frame range-angle map RM1, if there is a point with a difference of zero, set the amplitude of the point to zero, obtain a new range-angle map RM'1, in this way, obtain the t-th frame range-angle map denoted as RM'1, finally, obtain T frame range-angle maps after static clutter removal {RM' t ,i=1,2,…,T};

[0070] Step 2.2.2, process {RM' t ,t=1,2,…,T} using a constant false alarm detector, obtain a target position coordinate set {(PX t ,PY t ),t=1,2,…,T}, (PX t ,PY t ) represents the target position coordinate set of RM' t ;

[0071] Let the target position coordinate set {(PX t ,PY tThe set of amplitudes corresponding to t = 1, 2, ..., T is: in, (PX) t ,PY t The corresponding amplitude set;

[0072] Step 2.2.3, RM' t Target location coordinate set (PX) t ,PY t Within each coordinate point, take the average of its two adjacent points along the distance direction, and use this average as the new amplitude for each coordinate point, thus obtaining (PX). t ,PY t The corresponding new amplitude set This leads to a new distance-angle map RM. t ;

[0073] Step 2.2.4: Based on the new amplitude set Calculate the new distance-angle map RM” t The phase value of each target position in the T frames is obtained by concatenating the phase values ​​of the i-th target position frame by frame to obtain the i-th segment of phase change signal. I segments of phase change signal were acquired. in, This represents the phase value of the i-th target position in the t-th frame, and each phase change signal contains T frames;

[0074] Step 2.2.5: Phase unwrapping is used to reduce the phase value to the range of -π to π. Then, differential, smoothing, and bandpass filtering are used to process the reduced phase value to obtain the noise-removed I-segment phase change signal.

[0075] Step 2.2.6: Select the two segments with the highest signal-to-noise ratio in the preprocessed I-segment phase change signal and denote them as X(3) and X(4);

[0076] Step 2.2.7: Process the noise-removed I-segment phase change signal according to the process in Step 2.1.3 to obtain the main frequency RR6 and its corresponding signal-to-noise ratio SNR6; the distance angle diagrams of the constant false alarm rate detector before and after processing are shown below. Figure 3 As shown.

[0077] Step 3: Feature-level fusion;

[0078] Step 3.1: After normalizing the video preprocessing results X(1) and X(2) with the radar preprocessing results X(3) and X(4), four normalized input signals are obtained, and the r-th input signal is denoted as . in, Let F represent the value of the r-th input signal at the F-th point, where F represents the sample size, and r = 1, 2, 3, 4;

[0079] Step 3.2: Let M be the length of the embedded window, and Embedding the r-th input signal X'(r) into the Hankel trajectory matrix yields the r-th embedded Hankel trajectory matrix H of dimension M×K. (r) Where K represents the number of subsequences, and K = F - M + 1;

[0080] Using multivariate singular spectrum analysis, all embedded Hankel trajectory matrices {H (r) The embedding of the group r = 1, 2, ..., 4 is a vertical form H with dimension M × K × 4;

[0081] Step 3.3: Perform singular value decomposition on H to obtain 4×M eigenvalues, denoted as λ1, λ2, ..., λ e ,...,λ 4M and their corresponding eigenvectors U1, U2, ..., U e ,...,U 4M ; where λ e U represents the e-th eigenvalue. e λ represents the e-th eigenvalue. e The corresponding feature vector;

[0082] Let the rank of H be k, calculate the g-th eigenvalue λ. g The corresponding right singular vector And calculate the g-th dimension reduction matrix. g = 1, 2, ..., k, H g The rank of H is 1. g The dimension is 4M×K;

[0083] Step 3.4: Use the anti-angle average method to calculate H. g Converted to the g-th segment respiratory pulse signal Thus, k segments of respiratory pulse signals are obtained. in, This represents the amplitude corresponding to the F-th point of the g-th respiratory pulse signal; when g = r, express Approximate value;

[0084] k-segment respiratory pulse signals Divided into in, This represents the w-th respiratory pulse signal; This indicates that the result of k / 4 is rounded down.

[0085] Select the first group of respiratory pulse signals As a feature-level fusion result;

[0086] Step 3.5: Calculate using Fast Fourier Transform. main frequency RR2, main frequency RR3, main frequency RR4, The main frequency is RR5. Signal-to-noise ratio (SNR2), Signal-to-noise ratio (SNR) 3 Signal-to-noise ratio (SNR) 4 The signal-to-noise ratio (SNR) is 5.

[0087] Step 4: Decision-level integration;

[0088] Signal-to-noise ratio {SNR} s Summing the sums of the values ​​of s = 1, 2, ..., 6, we get the sum of the signal-to-noise ratios, denoted as SNR. sum Among them, SNR s This represents the signal-to-noise ratio corresponding to the s-th type of respiratory signal;

[0089] For the dominant frequency RR corresponding to the s-th type of respiratory signal s The weight corresponding to the s-th type of respiratory signal is SNR. s / SNR sum s = 1, 2, ..., 6; thus, the two-level fusion result is calculated.

[0090] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0091] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0092] like Figures 4a-4d As shown, the comparative experiment was designed to demonstrate the effectiveness of two-level fusion under four spontaneous motion scenarios (stationary, walking, running, and changing) set in the BSIPL-RR dataset through Bland-Altman analysis. The four comparison methods are as follows: (1) Video RR measurement results; (2) Frequency-modulated continuous wave radar RR measurement results; (3) Feature-level fusion results: the SNR weighted sum of four multivariate singular spectrum analysis results is used as the feature-level fusion result; (4) Decision-level fusion results: the SNR weighted sum of two single-mode measurement results and four multivariate singular spectrum analysis results (a total of six results) is used as the decision-level fusion result, i.e., RR fusion; From the Bland-Altman plot, the two-stage fusion method has the smallest error with the reference RR value, which is reflected in the narrowest 95% consistency range and the more concentrated sample points near 0 bpm. There are still some outliers for each scene using the two-stage fusion method, but compared with the two single modal measurement methods and the feature-level fusion method, the number of outliers is significantly reduced;

[0093] Figure 5 A box plot is shown, which compares the root mean square error (RMSE) of each fusion measurement stage, and the blue box in the figure represents the description group of each RMSE through its quartile, and the red line in each box represents the median of each stage RMSE; from the figure, the two-stage fusion method improves the accuracy of the results and the stability of the system;

[0094] Figure 6 The influence of the window length of multivariate singular spectrum analysis on the measurement error RMSE and the average Pearson coefficient is shown.

Claims

1. A non-contact human respiration rate measurement method based on video and frequency-modulated continuous wave radar information fusion, characterized in that, Comprising the following steps: Step one, dual-mode data generation; Step 1.1, video measurement part; The feature point motion trajectory time sequence of the left and right two shoulder regions of interest is extracted from the J-frame video image of the subject; wherein the coordinate information of N feature points of the jth frame shoulder region of interest is denoted as wherein, The coordinate information of the nth feature point of the shoulder region of interest in the jth frame video image is denoted as j∈[2, J]. Vertical pixel motion trajectories are generated from the coordinate information of the n-th feature point of the region of interest of the shoulder of the j-th frame and horizontal pixel motion trajectories n = 1,..., N; where, and denote the horizontal and vertical coordinates of the n-th feature point of the region of interest of the shoulder of the j-th frame, respectively. Step 1.2, radar measurement part; {RM t , t = 1, 2,..., T} from the T-frame radar raw data of the subject, wherein RM t represents the range-angle map of the t-th frame. Step two, dual-mode data preprocessing and single-mode measurement; Step 2.1, video measurement part; Step 2.1.1, after band-pass filtering the vertical pixel motion trajectory Y(n) and the horizontal pixel motion trajectory X(n) respectively, the frequency spectrum f X(n) of X(n) and the frequency spectrum f Y(n) of Y(n) are calculated by using fast Fourier transform respectively, so that the residual frequency spectrum f′ Y(n) of the de-noised vertical pixel motion trajectory Y(n) is obtained by using spectrum subtraction; if the amplitude of f′ Y(n) is negative, then f′ Y(n) = 0; otherwise, the time domain signal Y′(n) is converted from f′ Y(n) by using inverse transform of fast Fourier transform. Step 2.1.2, using principal component analysis to extract the waveform in Y'(n) whose cumulative contribution rate is greater than the threshold value and sum in time domain, obtaining the sum results of left and right shoulder, respectively recorded as X(1) and X(2); Step 2.1.3, using serial adaptive noise complete ensemble empirical mode decomposition method to decompose the waveform, obtaining several intrinsic mode functions; selecting several intrinsic mode functions with the highest signal-to-noise ratio to use fast Fourier transform to calculate their main frequency, finally selecting the main frequency with the highest frequency as RR1, and its corresponding signal-to-noise ratio as SNR1; Step 2.2, radar measurement part; Step 2.2.1, process the time series of range-angle maps {RM t , t = 1, 2,..., T} using a static clutter removal method to obtain T frames of static clutter removed range-angle maps {RM' t , t = 1, 2,..., T}; wherein RM' t is the t frame of static clutter removed range-angle map. Step 2.2.

2. Process {RM' (t), t = 1, 2,..., T} with constant false alarm detector to get target position coordinate set {(PX t , PY t ), t = 1, 2,..., T}, where (PX t , PY t ) represents target position coordinate set of RM' t . t ​ Let the target position coordinate set {(PX t , PY t ), t = 1, 2,..., T} correspond to the amplitude set wherein, represents the amplitude set corresponding to (PX t , PY t ); Step 2.2.3, RM′ t Target location coordinate set (PX) t PY t For each coordinate point within the range, take its two adjacent points along the distance direction, and then take the average of the three coordinate points as the new amplitude for each coordinate point, thus obtaining (PX). t PY t The corresponding new amplitude set This leads to a new distance-angle diagram RM″. t ; Step 2.2.4, calculating new amplitude set according to new distance-angle map RM" calculating new distance-angle map RM" t The phase value of each target position is connected by frame, and the phase value of the i-th target position of the T-th frame is connected by frame to obtain the i-th phase change signal I phase change signals are obtained wherein, The phase value of the i-th target position of the t-th frame is connected by frame to obtain the i-th phase change signal Step 2.2.5, using phase unwrapping to reduce the phase value of each frame within the range of-π to π, and then using difference, smoothing and band-pass filtering to process the reduced phase value, obtaining the I section phase change signal after removing noise; Step 2.2.6, selecting the two signals with the highest signal-to-noise ratio in the processed I section phase change signal as X(3) and X(4); Step 2.2.7, according to the process of step 2.1.3, processing the I section phase change signal after removing noise to obtain the main frequency RR6 and its corresponding signal-to-noise ratio SNR6; Step three, feature level fusion; Step 3.1, after the normalization operation of the video pre-processing results X(1) and X(2) and the radar pre-processing results X(3) and X(4), four normalized input signals are obtained, and the rth input signal is recorded as wherein, X(F, r) represents the value of the rth input signal at the Fth point, F represents the sample amount, and r = 1, 2, 3, 4. Step 3.2, let M be the length of the embedded window, and The rth input signal X'(r) is embedded into the Hankel trajectory matrix to obtain an rth embedded Hankel trajectory matrix H of dimension M x K (r) where K represents the number of constituent subsequences, and K = F - M + 1; All Hankel trajectory matrices {H (r) , r = 1, 2,..., 4} embedded are embedded into a vertical form H with dimension M x K x 4 using the multivariate singular spectrum analysis method. Step 3.3, singular value decomposition is performed on H to obtain 4 x M eigenvalues denoted as λ1, λ2,..., λ e , and their corresponding eigenvectors U1, U2,..., U 4M ,..., I e ,..., I 4M ; where λ e represents the e-th eigenvalue, U e represents the eigenvector corresponding to the e-th eigenvalue λ e ; Let H have rank k, compute the gth eigenvalue λ g The corresponding right singular vector And compute the gth reduced dimension matrix H g has rank 1, H g has dimension 4M x K; Step 3.4, calculate H g Converts into the gth breath pulse signal Thus, the kth breath pulse signal is obtained Wherein, The amplitude corresponding to the Fth point of the gth breath pulse signal k respiratory pulse signals are divided into wherein, represents the wth respiratory pulse signal; represents an operation of rounding down the result of k / 4.​ selecting a first set of respiratory pulse signals as a feature level fusion result; Step 3.5: Calculate using Fast Fourier Transform. main frequency RR2, main frequency RR3, main frequency RR4, The main frequency is RR5. Signal-to-noise ratio (SNR2), Signal-to-noise ratio (SNR) 3 Signal-to-noise ratio (SNR) 4 The signal-to-noise ratio (SNR) is 5. Step four, decision level fusion; The signal-to-noise ratio {SNR s , s = 1, 2,..., 6} is summed to obtain the sum of signal-to-noise ratios, denoted as SNR sum ; wherein SNR s represents the signal-to-noise ratio corresponding to the s-th respiratory signal. The main frequency RR corresponding to the s-th respiratory signal s The weight value SNR corresponding to the s-th respiratory signal s / SNR sum s = 1, 2,..., 6; thereby calculating the two-level fusion result 2. An electronic device comprising a memory and a processor, characterized in that The memory is used to store the program supporting the processor to execute the non-contact human respiratory rate measurement method of claim 1, and the processor is configured to execute the program stored in the memory.

3. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to execute the steps of the non-contact human respiratory rate measurement method of claim 1.