Respiration monitoring method based on depth camera and posture self-adaption

By employing a depth camera and attitude-adaptive respiratory monitoring method, combined with multi-scale wavelet decomposition and attention mechanisms, the problem of low accuracy and signal interruption in traditional respiratory monitoring methods in complex environments is solved. This enables dynamic identification and accurate assessment of individual respiratory zones, improving the robustness and adaptability of respiratory monitoring.

CN121059142APending Publication Date: 2025-12-05UNIV OF SCI & TECH BEIJING +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511260279.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing respiratory monitoring methods are not very accurate in complex indoor environments, are sensitive to changes in posture, and are interrupted when obstructed. They are difficult to adapt to individual differences and dynamically identify, and traditional methods are not adaptable to multiple application scenarios, resulting in discrepancies between the evaluation results and the actual situation.

Method used

A respiratory monitoring method based on depth camera and posture adaptation is adopted. Human images are acquired through a depth camera, and the three-dimensional spatial coordinates of skeletal key points are obtained using a posture recognition algorithm. Multi-scale wavelet decomposition and attention mechanism are combined to extract respiratory-induced micro-motion signals. Dynamic localization and signal compensation of the respiratory region are achieved through a random forest regression model and a spatial weighted correction strategy.

Benefits of technology

It enables accurate assessment of individual respiratory zone exposure levels under multiple scenarios and postures, improving the robustness and accuracy of monitoring. It has non-contact, posture-adaptive ROI positioning capabilities, can adapt to complex scenarios and individual differences, and enhances the extraction accuracy and signal continuity of respiratory signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121059142A_ABST
    Figure CN121059142A_ABST
Patent Text Reader

Abstract

The invention discloses a respiration monitoring method based on a depth camera and posture self-adaption, and belongs to the technical field of environmental health monitoring and intelligent perception. The method comprises the following steps: acquiring a human body image corresponding to a target to be subjected to respiration monitoring by using the depth camera; using a posture recognition algorithm to obtain human skeleton key point coordinates, and based on the human skeleton key point coordinates, obtaining respiration area center point coordinates; depth time sequence signals of the chest and abdomen key points are extracted, and on the basis, breath-induced micro-motion signals are obtained; detecting whether the signal has abnormal jump or not; and when abnormal hopping of the signal is detected, signal compensation is carried out to ensure that a signal waveform is close to a real signal in visual continuity and frequency characteristics. By the adoption of the scheme, the precision and robustness of respiration monitoring in a complex scene can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of environmental health monitoring and intelligent sensing technology, in particular to a breathing monitoring method based on a depth camera and posture self-adaptation. BACKGROUND

[0002] With the acceleration of urbanization and the extension of indoor activities, people are increasingly concerned about the health risks of indoor air pollution exposure. Existing exposure assessment methods mostly use fixed monitoring points or wearable biological monitoring devices. Although these methods have certain detection capabilities, they have limitations such as limited spatial coverage, high cost, and difficulty in popularization. Especially, the wearable device belongs to direct contact measurement, which may interfere with the behavior habits of the detected person, thereby affecting the respiratory rhythm, causing differences between the sampled signal and the true breathing in the natural state, and affecting the accuracy of the evaluation results. In addition, there is still a lack of dynamic identification and continuous tracking means for the human respiratory zone, which makes it difficult to reflect the actual inhalation risk of individuals in different spatial positions. At the same time, most methods ignore the dynamic diffusion path of pollutants in three-dimensional space affected by ventilation, thermal convection and human activity, causing deviations between the exposure estimation results and the actual situation. In multi-scene applications (such as kitchens, desks, and vehicle interiors), fixed or wearable devices are also difficult to adapt to complex spatial structures and behavior states, and traditional methods rely on manual settings in terms of scale adjustment and pollution source modeling, limiting their feasibility in terms of intelligence, continuity and wide deployment.

[0003] In summary, traditional exposure assessment methods (such as fixed monitoring instruments and wearable sensors) usually cannot reflect individual differences and are difficult to capture the dynamic spatial relationship between the human respiratory zone and the pollution source in real time. In addition, the existing methods have obvious deficiencies in the accuracy and adaptability of the assessment of the diffusion path and exposure level of pollutants in complex scenarios (such as kitchen fumes, office printer volatile substances, and vehicle exhaust), making it difficult to meet the needs of personalized and refined exposure risk analysis. In addition, traditional breathing monitoring methods usually use fixed human region positioning, which is difficult to adapt to changes in human posture and obstructed environments, resulting in loss or discontinuity of breathing signals; while the existing multi-scale wavelet decomposition method can separate noise to some extent, but it cannot accurately distinguish between true breathing signals and interference noise when the environment fluctuates greatly, reducing the accuracy and stability of the monitoring. SUMMARY

[0004] The present application provides a breathing monitoring method based on a depth camera and posture self-adaptation to solve the technical problem of poor accuracy and stability of existing exposure assessment methods for breathing monitoring.

[0005] To solve the above technical problems, the present application provides the following technical solutions: In one aspect, the present application provides a breathing monitoring method based on a depth camera and posture adaptation, which comprises: acquiring a human body image corresponding to a target to be monitored for breathing by using a depth camera; based on the human body image, acquiring three-dimensional space coordinates of preset human body skeleton key points by using a preset posture recognition algorithm, and obtaining three-dimensional space coordinates of a center point of a breathing area based on the acquired three-dimensional space coordinates of the human body skeleton key points; wherein the breathing area refers to a three-dimensional space with the mouth and nose as the center; based on the human body image, extracting a depth time sequence signal of a preset chest and abdomen key point located in the breathing area, and obtaining a breathing-induced micro-motion signal based on the depth time sequence signal.

[0006] Further, after obtaining the breathing-induced micro-motion signal, the method further comprises: detecting whether the breathing-induced micro-motion signal has abnormal jumps; when detecting that the breathing-induced micro-motion signal has abnormal jumps, performing signal compensation on the micro-motion signal with abnormal jumps to ensure that the signal waveform is close to the true signal in terms of visual continuity and frequency characteristics.

[0007] Further, the obtaining of the three-dimensional space coordinates of the center point of the breathing area based on the acquired three-dimensional space coordinates of the human body skeleton key points comprises: using the three-dimensional space coordinates of the human body skeleton key points to form a feature vector; inputting the feature vector into a pre-trained breathing area prediction model, and outputting the three-dimensional space coordinates of the center point of the breathing area from the breathing area prediction model; correcting the three-dimensional space coordinates of the center point of the breathing area output by the breathing area prediction model by using a spatial weighting correction strategy to obtain corrected three-dimensional space coordinates of the center point of the breathing area.

[0008] Further, the human body skeleton key points include shoulders, a spinal center, and a pelvis.

[0009] Further, the breathing area prediction model is a random forest regression model.

[0010] Further, the correction of the three-dimensional space coordinates of the center point of the breathing area output by the breathing area prediction model by using the spatial weighting correction strategy to obtain the corrected three-dimensional space coordinates of the center point of the breathing area comprises: calculating the corrected three-dimensional space coordinates of the center point of the breathing area by using the following formula:

[0011] wherein, ; wherein, represents the coordinate of the center point of the corrected breathing region; represents the coordinate of the center point of the breathing region predicted in a certain time before the current frame after time averaging; represents the candidate coordinate of the center point of the breathing region predicted in the current frame, is the single predicted point coordinate participating in the weighted correction calculation, is used to compare with the historical average position to determine the distance from the historical average to determine the weight; m represents the total number of candidate coordinates of the center point of the breathing region participating in the weighted correction calculation, that is, how many participate in this weighted average operation, representing the size of the predicted point set used for correction calculation at present; represents the Euclidean distance between two points (here, the coordinate points ) and the historical average position ), which measures the straight-line distance between two points in space, and the formula is to calculate the distance between ) and ) in space; is a preset value, which is used to ensure the numerical stability of the weight calculation, that is, to ensure that when ) is very close to ), the denominator will not be 0 to cause meaningless calculation, so that the weighted correction strategy can be stably executed.

[0012] Further, based on the depth time sequence signal, a breathing-induced micro-motion signal is obtained, comprising: performing multi-scale wavelet decomposition on the depth time sequence signal, decomposing the depth time sequence signal according to frequency levels through multi-scale wavelet decomposition to obtain a multi-scale wavelet decomposition result; for the multi-scale wavelet decomposition result, using an attention mechanism to adaptively weight and fuse different scale components to obtain a breathing-induced micro-motion signal.

[0013] Further, the attention mechanism is represented as: ; wherein, ; wherein, is the weight of the jth layer scale; is the attention score of the jth layer scale; is the attention score of the kth layer scale. When calculating the weight of the jth layer scale , participate in the exponential summation operation of the attention scores of all layer scales to determine the relative importance of different layer scales in adaptive weighted fusion; K represents the total number of scales in the multi-scale wavelet decomposition result, that is, the total number of scale components participating in adaptive weighted fusion; is an attention projection vector; represents the transpose of a matrix; is a linear weight matrix of the jth layer scale; is a detail coefficient of the jth layer scale obtained by multi-scale wavelet decomposition; is a bias term.

[0014] Further, the detection of whether the respiratory-induced micro-motion signal has an abnormal jump comprises: The fluctuation is detected using the standard deviation of the inter-frame respiratory region average depth change; if the inter-frame respiratory region average depth change corresponding to consecutive N frames exceeds the standard deviation, it is judged that there is an abnormal jump.

[0015] Further, the signal compensation of the micro-motion signal with an abnormal jump comprises: The long short-term memory network is used to predict the occluded signal according to the historical signal, so as to realize signal compensation.

[0016] In another aspect, the present application also provides an electronic device comprising a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to realize the above-mentioned method.

[0017] In another aspect, the present application also provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, which is loaded and executed by the processor to realize the above-mentioned method.

[0018] The technical scheme provided by the present application aims at the problems of low accuracy of respiratory monitoring, sensitivity to posture changes, signal interruption during occlusion, etc. in a complex indoor environment, and combines a depth camera with a multi-modal intelligent algorithm to realize respiratory monitoring, which can significantly improve the robustness of dynamic monitoring and the accuracy of individual exposure evaluation; the beneficial effects brought by it at least include: Firstly, in terms of posture recognition and respiratory region modeling, the present application establishes a respiratory region dynamic mapping mechanism through bone key point extraction combined with random forest, which can accurately predict the respiratory region center position according to real-time posture changes, and reduces the interference of factors such as bone jitter on respiratory region ROI positioning through a spatial weighting correction strategy. Compared with the traditional monitoring method based on fixed region, the present application realizes non-contact and posture-adaptive ROI positioning, has strong individual difference adaptation ability and posture change robustness, breaks through the limitations of traditional static ROI assumption and fixed position recognition, and still has stable tracking ability under non-standard postures such as lateral lying, turning over and reclining.

[0019] Secondly, in the aspect of respiratory micro-motion signal extraction, the traditional method relies on low-pass filtering or fixed window average strategy, which cannot effectively deal with high-frequency noise interference and non-periodic motion background. The present application introduces multi-scale wavelet decomposition technology to separate the signal according to the frequency level, and combines the attention mechanism to adaptively weight each scale component, which can emphasize the low-frequency respiratory dominant component and suppress the high-frequency environmental noise. Compared with the traditional wavelet fusion method, this method has stronger scale selectivity and signal discrimination ability, which can effectively improve the extraction accuracy of weak respiratory signal in complex scenes and improve the distinguishability of micro-motion signal in complex background.

[0020] Thirdly, in the aspect of occlusion compensation and waveform continuity, the traditional depth signal processing is sensitive to occlusion, which is easy to cause the interruption of respiratory signal. The present application realizes the continuous prediction and post-processing smoothing of respiratory signal during occlusion through LSTM time series modeling and interpolation repair algorithm, which guarantees the continuity of the monitoring signal in visual and frequency domain features, improves the robustness of the method in practical application, and enhances the practicability and scene adaptability of the method.

[0021] In summary, the present application realizes the accurate evaluation of individual respiratory zone exposure level under multi-scene, multi-pose and dynamic conditions by constructing a pose-adaptive ROI recognition mechanism, introducing a multi-scale wavelet-attention fusion strategy and an occlusion-robust LSTM prediction compensation scheme, which breaks through the limitations of existing methods such as poor spatial adaptability, signal distortion and single analysis dimension, and has higher practical value and engineering promotion potential. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0023] Figure 1 is the execution flow diagram of the respiratory monitoring method based on depth camera and pose adaptation provided by the embodiment of the present application; Figure 2 is the human body skeleton key point recognition diagram under different poses provided by the embodiment of the present application; Figure 3 is the key point driven respiratory zone prediction function diagram provided by the embodiment of the present application; Figure 4 is the structural block diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0025] First of all, it should be noted that in the embodiments of the present application, the words such as "exemplarily", "for example" and the like are used to represent as an example, illustration or description. Any embodiment or design scheme described as "exemplary" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of the word "exemplarily" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be either one of the two.

[0026] First embodiment

[0027] The present embodiment provides a breathing monitoring method based on a depth camera and a posture adaptation, proposes a dynamic posture-breathing region mapping model based on a skeleton key point and an LSTM time sequence prediction mechanism, and combines a multi-scale wavelet decomposition and an attention weighted fusion algorithm to realize dynamic tracking of a breathing region, enhancement of a micro-motion signal and continuous repair of a signal under a shielding condition, effectively improving the accuracy and robustness of breathing monitoring in a complex scene. It is suitable for individual pollutant exposure dynamic evaluation in various spatial scenes and scales.

[0028] The method can be realized by an electronic device, and the execution flow of the method is as shown in Figure 1 The method comprises the following steps: S1, acquiring a human body image corresponding to a target to be subjected to breathing monitoring by using a depth camera; S2, acquiring a three-dimensional space coordinate of a preset human body skeleton key point by using a preset posture recognition algorithm based on the human body image, and obtaining a three-dimensional space coordinate of a center point of a breathing region based on the three-dimensional space coordinate of the human body skeleton key point; wherein the breathing region refers to a dynamic three-dimensional space with the mouth and nose as the center, which is a key spatial range for calculating individual pollutant exposure dose; Specifically, the present embodiment acquires three-dimensional coordinate information of a human body skeleton key point by using a depth camera and a posture recognition algorithm, and predicts a position of a center point of a ROI of a breathing region by a random forest regression model to realize posture-adaptive ROI positioning; the specific implementation process is as follows: S21, using the three-dimensional space coordinates of the human body skeleton key points to form a feature vector; In the present embodiment, 16 main nodes in the human body skeleton key points are selected to form a three-dimensional space coordinate set, including positions such as a head, a shoulder, a spine, a pelvis, and limb joints, as shown in formula (1). Figure 2The text indicates the functional names of human skeletal feature points in different postures (sitting, standing, walking) and their index numbers in the node set. Among them, the nodes closely related to respiratory zone prediction are the shoulder, spine center, and pelvis. Therefore, in this embodiment, the feature vector is composed of the coordinates of the three key points: shoulder, spine center, and pelvis. Figure 3 This paper demonstrates the modeling process for predicting the center location of the respiratory zone based on skeletal key points. Two human skeletal structures are represented by red and green dots, respectively, indicating their respective skeletal key points. Gray line segments between these dots represent physiological structural connections. Orange connecting lines illustrate the spatial mapping relationship between the two individuals at corresponding joint points, which can be used to compare respiratory zone prediction results under different body postures.

[0029] (1)

[0030] S22, the feature vector is input into the pre-trained respiratory region prediction model, and the respiratory region prediction model outputs the three-dimensional spatial coordinates of the center point of the respiratory region, represented as: (2) in, The prediction function (in this embodiment, a Random Forest (RF) regression model is used) is trained using the function... The coordinates of the center point C of the respiratory region can be predicted. ROI The implementation process is as follows: 1) Data preparation and input feature construction First, a depth camera and pose estimation algorithm are used to acquire full-body skeletal information, extracting multiple key skeletal points, including the shoulder, spine, and pelvis. Each point contains its (x, y, z) coordinates in three-dimensional space. Under standard configuration, three key points (shoulder, spine, pelvis) are selected to form a 3 × 3 = 9-dimensional feature vector, which is used as the input to the model.

[0031] 2) Target output setting: The center point of the respiratory zone, either accurately marked or collected. As the target output in supervised learning, the model learns the regression relationship from the pose feature X to the position of that point.

[0032] (3)

[0033] 3) Construction of the random forest regression model: Random forest regression model is derived from Each independent regression decision tree is constructed, each tree is constructed based on different random sample subsets and feature subsets during the training process, and a prediction result of the respiratory region center point position is respectively output, and finally the regression output of the overall model is obtained by averaging the prediction coordinates of all trees.

[0034] (4)

[0035] In the formula: : input features of key skeleton points; : overall regression function, parameters are determined by training samples; : three-dimensional coordinates of the respiratory region center point; : number of trees in the forest; : the first regression tree prediction model.

[0036] The embodiment takes the real labeled respiratory region center coordinates as the target output, realizes robust mapping learning of the non-linear space structure by constructing multiple decision trees and using random subsets for splitting judgment during the training process. The RF model has strong generalization ability under the conditions of posture diversity and individual differences of human body structure, and has low overfitting risk in the small sample training scene.

[0037] Based on the above, the embodiment uses the human body skeleton key point coordinates (a total of 16 points) collected by the depth camera, constructs a random forest regression model taking the shoulder, spine, pelvis and other key points as input, and dynamically predicts the respiratory region center position. Non-contact, posture-adaptive respiratory region positioning is realized, which has strong individual difference adaptation ability and posture change robustness, and breaks through the limitations of traditional static ROI hypothesis and fixed position recognition.

[0038] In addition, it should be noted that when performing dynamic prediction of the respiratory region, although the random forest regression algorithm based on the skeleton key points is adopted in the embodiment, replacing it with a lightweight convolutional neural network (such as MobileNet) or a support vector regression (SVR) algorithm can also realize the mapping relationship learning of the posture-respiratory region. Therefore, the embodiment does not limit the specific method for performing dynamic prediction of the respiratory region.

[0039] Further, to improve the robustness of point cloud positioning, in the embodiment, S2 further includes: S23, correcting the three-dimensional space coordinates of the center point of the respiratory region output by the respiratory region prediction model by using a spatial weighting correction strategy to obtain the three-dimensional space coordinates of the center point of the corrected respiratory region. (5) wherein, (6) wherein, represents the coordinate of the center point of the corrected breathing region; is a historical average breathing region position, representing the coordinate value of the center point of the breathing region obtained by time averaging the center point coordinates of the breathing region predicted in a certain time (such as the past 10 frames) before the current frame t, that is, the result obtained by averaging the center point coordinates of the breathing region predicted in a certain time (such as the past 10 frames); the historical average value is used to calculate the distance between the current prediction point and the historical average position. In this way, the prediction point closer to the historical average value is given a higher weight and is considered more reliable, thereby enhancing the anti-interference ability of the human skeleton tracking error. The weighting correction mechanism can effectively smooth the fluctuation of the center point of the breathing region, and improve the stability and accuracy of the exposure evaluation. represents the candidate coordinate of the center point of the breathing region predicted in the current frame, which is a single prediction point coordinate participating in the weighted correction calculation, and is used to compare the distance with the historical average position to determine the weight; m represents the total number of candidate coordinates of the center point of the breathing region participating in the weighted correction calculation, that is, how many participate in the weighted average operation, representing the size of the prediction point set used for correction calculation at present; represents the Euclidean distance between two points (here, the coordinate points ( ) and the historical average position ( ) ), which measures the straight line distance between the two points in space. In the formula, the distance between ( ) and ( ) in space is calculated; is a preset value used to ensure the numerical stability of the weight calculation, that is, to ensure that when the distance between ( ) and ( ) is extremely close, the denominator will not be 0 to cause meaningless calculation, so that the weighted correction strategy can be stably executed.

[0040] Based on the above, the embodiment introduces a spatial weighting strategy, which distributes weights according to the distance between the current prediction candidate point and the historical average position, and fuses multiple prediction results to generate the final breathing region position, which can effectively suppress the abnormal drift caused by frame jitter and key point disturbance, and improve the stability and consistency of the recognition result between consecutive frames.

[0041] S3, based on the human image, extracts a preset chest and abdomen key point depth time sequence signal located in the breathing region, and obtains a breathing-induced micro-motion signal based on the depth time sequence signal; Specifically, the embodiment performs multiscale wavelet decomposition on the depth time sequence signal of the chest and abdomen key points, and introduces an attention mechanism to weight and fuse signals of each frequency band, thereby enhancing the respiratory-induced micro-motion signal and suppressing environmental noise interference. The specific implementation process is as follows: S31, performing multiscale wavelet decomposition on the depth time sequence signal, and obtaining multiscale wavelet decomposition results by decomposing the depth time sequence signal according to frequency levels through multiscale wavelet decomposition; It should be noted that multiscale wavelet decomposition is a time-frequency localization method that can effectively extract local features of signals at different frequencies. In respiratory monitoring, chest and abdominal micro-motions caused by respiratory movements are low-frequency periodic changes, while environmental disturbances, sensor noise, etc. are usually high-frequency or non-periodic components. The method is used to extract respiratory-induced small periodic displacements from depth signals of the chest and abdominal regions of the human body collected by a depth camera. The original signal is decomposed according to frequency levels through multiscale wavelet decomposition, and different scale components are adaptively weighted and fused through an attention mechanism, thereby enhancing the respiratory dominant frequency component and suppressing environmental noise and non-respiratory disturbances. Finally, a high signal-to-noise ratio respiratory displacement time sequence signal is obtained, providing basic data for subsequent respiratory frequency estimation and behavior modeling. The signal contains respiratory-induced periodic low-frequency displacement (target component), environmental disturbance, device noise, etc. (non-target component).

[0042] The wavelet decomposition process is as follows: Let be the vertical direction (Z-axis) depth coordinate change of the selected key point in the respiratory region ROI within a unit time, with a unit of meters (m), and the depth time sequence is , reflecting the small body surface undulation of the chest and abdomen caused by respiratory movements.

[0043] (7)

[0044] The original depth signal is sequentially passed through a low-pass filter bank to extract low-frequency components and downsample, and finally obtain the Jth layer approximation coefficient , which is used to represent the slow trend of respiratory-induced displacement; at the same time, the high-frequency disturbance of each layer is extracted and downsampled through the corresponding high-pass filter g(j), and the jth layer detail coefficient is obtained, which is used to capture the fast-changing disturbance information. Both of them are obtained through multi-layer convolution and down-sampling iteration, and the process is shown in the following formula: (8) (9) It is decomposed into scales using Discrete Wavelet Transform (DWT): (10) wherein: : the sequence of approximation coefficients of the bottom layer (Jth layer), reflecting the low-frequency trend change of the signal, corresponding to the overall respiratory trend fluctuation.

[0045] : the sequence of detail coefficients at the jth scale, which is a time-varying component of a certain frequency component in the original depth signal, corresponding to the respiratory-induced micro-motion disturbance.

[0046] Parameters and are the decomposition results of the original depth signal sequence in the time domain through Discrete Wavelet Transform (DWT), reflecting the change characteristics of the signal at different frequency scales; their values vary with time and are calculated through layer-by-layer filtering and downsampling, used to separate the time characteristics of the respiratory main frequency and the background disturbance.

[0047] In addition, it should be noted that in addition to using wavelet decomposition and attention mechanism weighted fusion, methods with time-frequency analysis capability such as Empirical Mode Decomposition (EMD) and Short-Time Fourier Transform (STFT) can also be used to decompose the signal and perform scale weighting. Therefore, the specific algorithm for signal decomposition is not limited in this embodiment.

[0048] S32, for the multi-scale wavelet decomposition result, an attention mechanism is used to adaptively weight and fuse different scale components to obtain a respiratory-induced micro-motion signal.

[0049] It should be noted that in order to realize noise elimination and weight modeling, this embodiment introduces an attention mechanism to adaptively weight each scale, the input is the wavelet detail coefficient at each scale j, representing the micro-motion signal at different frequency scales, and the output is the attention weight , which determines the contribution of the jth layer wavelet detail coefficient in the final micro-motion enhanced signal, and then weights the important frequency components and suppresses the interference components: (11) wherein, (12) In the formula, is the weight (attention parameter) of the jth scale, indicating the contribution strength of the jth layer. The attention score is given at the j-th level. The attention score is given at the k-th level. The weights at the j-th level are calculated... )hour,( ) participates in the exponential summation of attention scores for all scales, used to determine the relative importance of different scales in adaptive weighted fusion; K represents the total number of scales in the multi-scale wavelet decomposition result, which is the total number of scale components participating in adaptive weighted fusion; This is the attention projection vector; Represents the transpose of a matrix; Let be the linear weight matrix at the j-th level scale; These are the detail coefficients at the j-th scale obtained through multi-scale wavelet decomposition; This is a bias term. Parameter , , All three are trainable parameters used to generate the attention score at the j-th layer scale. Among them, all three can be trained through backpropagation.

[0050] The final fused signal is: (13) in, The output is a micro-motion signal after multi-scale fusion, which is used for subsequent respiratory rate detection (through time-domain / frequency-domain periodic analysis); respiratory behavior temporal feature modeling (such as inhalation and exhalation duration); and in this embodiment, it is also used as an input signal for the subsequent occlusion compensation module.

[0051] Based on the above, this embodiment employs wavelet multi-scale decomposition to extract multi-band micro-motion information from depth time-series signals in the chest and abdominal regions. Combined with an attention mechanism, signals at different scales are weighted and fused to enhance the dominant respiratory frequency component and suppress background noise. This method effectively improves the extraction accuracy of weak respiratory signals in complex scenes and enhances the system's adaptability to low signal-to-noise ratio environments.

[0052] S4, detect whether there are abnormal jumps in the micromotion signal induced by breathing; Specifically, in this embodiment, the occlusion detection logic is as follows: For a given frame (e.g., frame t), suppose the chest and abdomen ROI region in this frame contains n depth data points, and the depth value of each point is... The formula for calculating the average depth of a single-frame ROI is:

[0053] in, is the average depth of the t-th frame ROI, which is the final single-frame average depth result, i.e., the inter-frame ROI average depth; n is the number of depth data points in the t-th frame ROI region, used to determine the base for summation and averaging; is the depth value of the i-th depth data point in the t-th frame ROI region, which is the original data participating in the average calculation; Summation operation is performed on the depth values of the first to n-th depth data points in the t-th frame ROI region.

[0054] The inter-frame ROI average depth change is used for volatility detection. If there is an abnormal jump (such as occlusion or signal mutation) in consecutive N frames, the occlusion compensation mode is triggered, and S5 is executed. It should be noted that the standard deviation of the inter-frame ROI average depth change represents the normal fluctuation amplitude. If the average depth change is much larger than the normal fluctuation amplitude represented by the standard deviation, and such a situation of exceeding the reasonable range occurs for consecutive N frames, it is determined that there is an abnormal jump of the respiratory-induced micro-motion signal that needs to be processed (to avoid single-frame noise misjudgment).

[0055] S5, when detecting that the respiratory-induced micro-motion signal has an abnormal jump, performing signal compensation on the micro-motion signal with the abnormal jump to ensure that the signal waveform is close to the true signal in terms of visual continuity and frequency characteristics.

[0056] Specifically, when detecting occlusion or signal mutation, the embodiment starts the waveform prediction and interpolation repair module based on the long short-term memory network (LSTM), uses the long short-term memory network to predict the occluded signal according to the historical signal, and realizes signal compensation. Ensure the continuity and reliability of the respiratory signal in the occlusion scene.

[0057] Among them, the long short-term memory network (LSTM, Long Short-Term Memory) is a classic neural network structure for processing time series problems, which has the ability to "remember long-term trends and filter short-term interference", and is particularly suitable for predicting weak periodic signals such as respiratory waveforms. The LSTM modeling process is as follows: 1) Input sequence preparation: Construct the sequence of micro-motion signals of the previous T frames: (14) Among them, the input sequence for prediction is a historical signal window of length T represents the original micro-motion signal of the t-th frame.

[0058] 2) Predict the current frame: (15) wherein, denotes the t-th frame signal predicted by LSTM, θ denotes the weight parameters of LSTM, and the network structure is:

[0059] 3) Signal repair: If the occlusion time is t1 t2, the output predicted sequence is: (17) wherein, denotes the occlusion segment waveform sequence predicted by LSTM. Then, the spline interpolation function and the filter are used for smoothing processing: (18) wherein, is the final interpolated and filtered signal estimation value, which ensures that the waveform is close to the real signal in terms of visual continuity and frequency characteristics.

[0060] Based on the above, the occlusion detection and prediction compensation module is designed in the embodiment. When occlusion or key point failure occurs, a time series prediction model based on long short-term memory network (LSTM) is called, and an interpolation and filtering strategy is combined to reconstruct the missing stage of the respiratory waveform data, so as to ensure the continuity and availability of the system output. This mechanism improves the robustness of the method in the face of uncertain factors such as occlusion and jitter in actual application.

[0061] In addition, it should be noted that, in terms of occlusion compensation and sequence prediction, in addition to the LSTM model used in the embodiment, a gated recurrent unit (GRU), a Transformer structure, or a Kalman filter and Bayesian prediction method based on probability statistics can also be used. For this, the embodiment does not limit the specific model.

[0062] In summary, the embodiment introduces a depth camera and human posture recognition technology to construct a multi-scale and multi-scene respiratory zone recognition method based on a depth camera, realizes accurate prediction and real-time tracking of the position of the human respiratory zone in complex postures and dynamic scenes, can automatically extract and track the three-dimensional spatial position of the respiratory zone, combines multi-scale spatiotemporal modeling and respiratory micro-motion signal enhancement algorithms, realizes accurate recognition and dynamic stable tracking of the respiratory zone under multi-scene and multi-posture conditions, and improves the accuracy and adaptability of individual pollution exposure assessment. The method has the advantages of non-contact, dynamic and personalized recognition, can adapt to various postures such as lying, sitting and moving, and various spatial scenes such as kitchens, offices and vehicles, and is widely applicable to the fields of intelligent health management, indoor air safety monitoring and occupational exposure assessment. Thus, the adaptability, real-time performance and robustness of respiratory zone positioning are significantly improved, the limitations of traditional technology in static, single posture and low resolution scenes are broken through, and key technical support is provided for non-contact physiological monitoring and intelligent health sensing.

[0063] Second embodiment

[0064] The embodiment provides an electronic device, such as Figure 4 As shown in the figure, the electronic device comprises a processor and a memory; wherein the processor and the memory can be connected through a communication bus; the memory stores at least one instruction, which is loaded and executed by the processor to realize the method of the first embodiment. In addition, the electronic device can further comprise a transceiver, and the processor and the transceiver can be connected through a communication bus, and the transceiver is used for communicating with other devices.

[0065] Next, the specific introduction of each component of the electronic device will be given in combination with Figure 4 The specific introduction of each component of the electronic device will be given in combination with The processor is the control center of the electronic device. The electronic device can include multiple processors. Each of the processors can be a single-CPU or a multi-CPU. The processor can be one processor or a collective term of multiple processing elements. For example, the processor can be one or more central processing units (CPUs), other general purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present application, such as one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or the like. The general purpose processor can be a microprocessor or any conventional processor, or the like. The processor can perform various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0066] In a specific implementation, as an embodiment, the processor can include one or more CPUs, such as CPU0 and CPU1 shown in FIG. 6, of course, this is only an exemplary description. Figure 4

[0067] The memory is used to store software programs for implementing the solution of the present application, and is controlled by the processor to perform the implementation. The specific implementation can refer to the method embodiments described above, and will not be described here.

[0068] ​Optionally, the memory may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory may be integrated with the processor or exist independently, and may be accessed through the interface circuit of the electronic device ( Figure 4 (Not shown in the image) is coupled to the processor; however, this embodiment of the invention does not impose specific limitations on this.

[0069] The transceiver may include a receiver and a transmitter. Figure 4 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver can be integrated with the processor or exist independently, and can be connected through the interface circuit of the electronic device (…). Figure 4 (Not shown in the image) is coupled to the processor, and this embodiment of the invention does not specifically limit this.

[0070] In addition, it should be noted that, Figure 4 The structure of the electronic device shown is not intended to limit the device. Actual devices may include more or fewer components than shown, or combine certain components, or have different component arrangements. Furthermore, the technical effects achieved by this electronic device when performing the method of the first embodiment described above can be referenced to the technical effects described in the first embodiment; therefore, they will not be repeated here.

[0071] Third Embodiment

[0072] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc. The instruction stored therein can be loaded and executed by a processor in a terminal.

[0073] Moreover, it should be noted that the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present application can take the form of an entirely or partially hardware embodiment, an entirely or partially software embodiment, or an embodiment combining software and hardware aspects. Furthermore, when implemented in software, the embodiments of the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, a computer diskette, an optical storage medium, a magnetic storage medium, and a semiconductor memory device). The computer program product includes one or more computer instructions that when loaded and executed by a computer, cause the computer to carry out the processes or functions described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, such as from a website, a computer, a server, or a data center to another website, computer, server, or data center through a wired (for example, infrared, wireless, microwave, or the like) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device, such as a server, data center, or the like, including one or more collections of available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.

[0074] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device that implements the flow or flows and / or block or blocks specified in the flowcharts and / or block diagrams. Figure 1 The flow or flows and / or block or blocks Figure 1 The device that implements the function specified in the flow or flows and / or block or blocks.

[0075] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product that includes an instruction device that implements the flow or flows and / or block or blocks specified in the flowcharts and / or block diagrams. Figure 1 The flow or flows and / or block or blocks Figure 1the functions specified in the individual block or blocks. Such computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable devices to generate a computer-implemented process, thus the instructions executed on the computer or other programmable devices provide processes for implementing the functions specified in the flowchart block(s) or block(s). Figure 1 the functions specified in the individual block or blocks. Such computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable devices to generate a computer-implemented process, thus the instructions executed on the computer or other programmable devices provide processes for implementing the functions specified in the flowchart block(s) or block(s). Figure 1 the functions specified in the individual block or blocks. Such computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable devices to generate a computer-implemented process, thus the instructions executed on the computer or other programmable devices provide processes for implementing the functions specified in the flowchart block(s) or block(s).

[0076] It should also be noted that, in the present document, the terms such as first and second, etc. are merely used to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal device. Without more limitations, the element defined by the statement "including a…", does not exclude the presence of other identical elements in the process, method, article or terminal device including the element. In addition, the term "and / or" is merely a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the existence of A alone, the existence of A and B together, and the existence of B alone, where A and B can be singular or plural. In addition, the character " / " in the present document generally represents an "or" relationship between the preceding and following associated objects, but it can also represent an "and / or" relationship, which can be understood in the context before and after. "At least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including single item or any combination of multiple items. For example, at least one of a, b or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0077] In addition, it can be understood that in various embodiments of the present application, the size of the sequence number of the above processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0078] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or in a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0079] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of functional modules / units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms. The units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present, or two or more units can be integrated in one unit.

[0080] If the method is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0081] Finally, it should be noted that the above description is only the preferred embodiment of the application, it should be pointed out that although the preferred embodiment of the application has been described, for those skilled in the art, once the basic creative concept of the application is known, several improvements and refinements can be made without departing from the principles of the application, and these improvements and refinements should also be considered as the protection scope of the application. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the application.

Claims

1. A method for respiration monitoring based on depth camera and pose adaptation, characterized in that, The method comprises the following steps: acquiring a human body image corresponding to a target to be monitored for respiration by using a depth camera; based on the human body image, acquiring three-dimensional spatial coordinates of preset human skeleton key points by using a preset posture recognition algorithm, and obtaining three-dimensional spatial coordinates of a center point of a respiration region based on the acquired three-dimensional spatial coordinates of the human skeleton key points, wherein the respiration region refers to a three-dimensional space with the mouth and nose as the center; based on the human body image, extracting a depth time sequence signal of a preset chest and abdomen key point located in the respiration region, and obtaining a respiration-induced micro-motion signal based on the depth time sequence signal.

2. The depth camera and posture adaptive based respiration monitoring method of claim 1, wherein, After obtaining the respiration-induced micro-motion signal, the method further comprises: detecting whether the respiration-induced micro-motion signal has an abnormal jump; when it is detected that the respiration-induced micro-motion signal has an abnormal jump, performing signal compensation on the micro-motion signal with the abnormal jump to ensure that the signal waveform is close to the true signal in terms of visual continuity and frequency characteristics.

3. The depth camera and posture adaptive based respiration monitoring method of claim 1, wherein, The method of obtaining the three-dimensional spatial coordinates of the center point of the respiration region based on the acquired three-dimensional spatial coordinates of the human skeleton key points comprises: using the three-dimensional spatial coordinates of the human skeleton key points to form a feature vector; inputting the feature vector into a pre-trained respiration region prediction model to output the three-dimensional spatial coordinates of the center point of the respiration region from the respiration region prediction model; using a spatial weighting correction strategy to correct the three-dimensional spatial coordinates of the center point of the respiration region output by the respiration region prediction model to obtain the corrected three-dimensional spatial coordinates of the center point of the respiration region.

4. The depth camera and posture adaptive based respiration monitoring method of claim 3, wherein, The human skeleton key points include shoulders, a spinal center, and a pelvis.

5. The depth camera and posture adaptive based respiration monitoring method of claim 3, wherein, The respiration region prediction model is a random forest regression model.

6. The depth camera and posture adaptive based respiration monitoring method as claimed in claim 3, wherein, The method of using a spatial weighting correction strategy to correct the three-dimensional spatial coordinates of the center point of the respiration region output by the respiration region prediction model to obtain the corrected three-dimensional spatial coordinates of the center point of the respiration region comprises: using the following formula to calculate the corrected three-dimensional spatial coordinates of the center point of the respiration region: wherein, ; In the formula, represents the coordinate of the center point of the corrected breathing region; represents the coordinate value of the center point of the breathing region predicted in a certain time before the current frame after time averaging; represents the candidate coordinate of the center point of the breathing region predicted in the current frame; m represents the total number of candidate coordinates of the center point of the breathing region participating in the weighted correction calculation; represents the Euclidean distance between two points; is a preset value.

7. The depth camera and posture adaptive based respiration monitoring method of claim 1, wherein, The method of obtaining the respiration-induced micro-motion signal based on the depth time sequence signal comprises: performing multi-scale wavelet decomposition on the depth time sequence signal, decomposing the depth time sequence signal according to frequency levels by multi-scale wavelet decomposition to obtain a multi-scale wavelet decomposition result; adopting an attention mechanism to adaptively weight and fuse different scale components for the multi-scale wavelet decomposition result to obtain the respiration-induced micro-motion signal.

8. The depth camera and posture adaptive based respiration monitoring method of claim 7, wherein, The attention mechanism is represented as: ; wherein, ; wherein, is the weight of the j-th level of scale; is the attention score of the j-th level of scale; is the attention score of the k-th level of scale; K represents the total number of levels of scale in the multi-scale wavelet decomposition result; is the attention projection vector; denotes the transpose of a matrix; is the linear weight matrix of the j-th level of scale; is the detail coefficient of the j-th level of scale obtained by multi-scale wavelet decomposition; is the bias term.

9. The depth camera and posture adaptive based respiration monitoring method as claimed in claim 1, wherein, The method of detecting whether the respiration-induced micro-motion signal has an abnormal jump comprises: using a standard deviation of inter-frame respiration region average depth change for volatility detection; if the inter-frame respiration region average depth change corresponding to consecutive N frames exceeds the standard deviation, it is determined that there is an abnormal jump.

10. The depth camera and posture adaptive based respiration monitoring method of claim 1, wherein, The method of performing signal compensation on the micro-motion signal with the abnormal jump comprises: using a long short-term memory network to predict a blocked signal according to historical signals to realize signal compensation.