Sleep monitoring method and system based on non-contact video data sequence

Through the sleep monitoring method of contactless video data sequence, the YOLOv11 network and the multi-physiological indicator fusion model are used to solve the discomfort and subjective deviation of traditional sleep monitoring, achieving high-precision sleep monitoring and staging, which is suitable for hospitals, nursing homes and families.

CN120531338AActive Publication Date: 2025-08-26YANGZHOU CHENGKE MEDICAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510704661.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-26
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional polysomnography monitoring technology (PSG) has contact sensors that cause discomfort, affect sleep status, and there are subjective differences in manual interpretation, resulting in evaluation bias.

Method used

The contactless video data sequence is used, and the face and chest abdomen are detected using the YOLOv11 network, combined with the RPPG signal extraction model, the respiratory signal extraction model and the blood oxygen saturation model, and multimodal fusion analysis is performed through the fusion recognition model of multiple physiological indicators to achieve sleep stage stage.

Benefits of technology

Achieve high-precision sleep monitoring, reduce sleep interference, and reduce subjective deviations in assessment. It is suitable for sleep health monitoring in hospitals, nursing homes and families.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120531338A_ABST
    Figure CN120531338A_ABST
Patent Text Reader

Abstract

The invention discloses a sleep monitoring method and system based on a non-contact video data sequence. The method comprises the steps that S1, video data streams of human body sleep are collected through a camera system; utilizing a YOLOv11 network to obtain face video data and thoracoabdominal video data; s2, the RPPG signal extraction model extracts and obtains an RPPG signal by using the face video data; s3, the breathing signal extraction model extracts thoracic and abdominal micro-motion change characteristics in the thoracic and abdominal video data by using an optical flow method, and noise filtering processing is carried out to obtain thoracic and abdominal motion signals as breathing signals; s4, the oxyhemoglobin saturation extraction model detects and outputs an oxyhemoglobin saturation signal by using the RPPG signal; and S5, performing multi-modal fusion analysis on the multi-physiological index fusion recognition model according to time slice T1 division to obtain a long-time-sequence sleep stage staging result. According to the invention, the physiological index signals are extracted and recognized by adopting the non-contact video data sequence, and the sleep stage staging result with a long time sequence is obtained, so that high-precision sleep monitoring and sleep stage recognition are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sleep monitoring analysis and processing, and in particular to a sleep monitoring method and system based on a contactless video data sequence. Background Art

[0002] In the field of sleep science, polysomnography (PSG), the traditional gold standard, has long held a central position in sleep staging assessment. This technology monitors sleep by collecting multiple physiological parameters: on the one hand, it records bioelectrical signals such as EEG activity, eye movement potentials, muscle electrical signals, and cardiac activity; on the other hand, it simultaneously acquires physiological indicators such as breathing patterns, chest and abdominal movements, blood oxygen saturation, and snoring characteristics. However, PSG monitoring requires strict adherence to electrode placement specifications. The use of these direct-contact sensors often causes significant discomfort to the subject, disrupting their normal sleep state and ultimately causing the data to deviate from their true sleep state. Furthermore, during data analysis, physicians must manually interpret PSG recordings based on standardized interpretation manuals developed by the American Academy of Sleep Medicine. Interpretations of the same set of monitoring data may vary from physician to physician, and this subjective factor inevitably introduces assessment bias. Summary of the Invention

[0003] The purpose of the present invention is to provide a sleep monitoring method and system based on a contactless video data sequence. The YOLOv11 network is used to capture the face and / or chest and abdomen of a person in a video data stream and obtain facial video data and chest and abdomen video data based on the human body structure relationship. The RPPG signal extraction model obtains a remote photoelectric volume pulse wave signal based on inter-frame difference processing and time-series physiological signal feature extraction processing, and separates it to obtain a blood oxygen saturation signal. The respiratory signal extraction model extracts the chest signal and the abdominal signal, and performs weighted fusion processing to obtain a chest and abdominal motion signal as a respiratory signal. The multi-physiological indicator fusion recognition model uses the RPPG signal, respiratory signal, and blood oxygen saturation signal to perform multimodal fusion analysis to obtain a long-term sleep stage classification result, providing technical support for human sleep quality detection and analysis.

[0004] The purpose of the present invention is achieved through the following technical solutions:

[0005] A sleep monitoring method based on a non-contact video data sequence, the method comprising:

[0006] S1. Collect a video data stream of a sleeping person through a camera system; use the YOLOv11 network to detect, segment, and track the face, chest, and abdomen of the person in the video data stream, and obtain facial video data and chest and abdomen video data;

[0007] S2. Constructing an RPPG signal extraction model, which uses facial video data to extract RPPG signals;

[0008] S3. Construct a respiratory signal extraction model. The respiratory signal extraction model uses the optical flow method to extract the micro-motion change characteristics of the chest and abdomen in the chest and abdomen video data and simultaneously uses the frequency domain adaptive notch filtering method to filter out noise. The sparse representation time domain filtering method is used to perform sparse representation and eliminate non-respiratory signals, thereby obtaining the chest and abdomen motion signal as the respiratory signal.

[0009] S4. Construct a blood oxygen saturation extraction model. The blood oxygen saturation extraction model uses the RPPG signal to extract blood oxygen signal features and detect and output the blood oxygen saturation signal;

[0010] S5. Construct a multi-physiological indicator fusion recognition model. The multi-physiological indicator fusion recognition model includes four sleep stages: wakefulness W, rapid eye movement REM, light sleep N1 / N2, and deep sleep N3. The multi-physiological indicator fusion recognition model uses RPPG signals, respiratory signals, and blood oxygen saturation signals to perform multimodal fusion analysis according to time slice T1 to obtain long-term sleep stage classification results.

[0011] In order to better implement the present invention, the YOLOv11 network is trained using a face dataset. The YOLOv11 network includes a backbone network, a feature enhancement network and a detection head. The backbone network includes a C3k2 module and is used to maintain the receptive field and extract basic features; the feature enhancement network includes an SPPF-C2PSA joint module and is used to obtain global information and adopt a channel space dual attention mechanism for feature enhancement; the detection head is constructed based on a dual-branch decoupling structure of classification and regression. The detection head uses a LoU-aware mechanism to predict the center point offset and width and height, and at the same time adds a dynamic label allocation strategy and dynamically adjusts the positive and negative sample thresholds through the LoU-aware mechanism; the detection head of the YOLOv11 network outputs the detection and segmentation results of the face, chest and abdomen and tracks them.

[0012] Preferably, the camera system includes at least one infrared camera, and the camera area of ​​the camera system covers the triangular area of ​​the subject's face and the chest and abdominal movement area. If the camera system includes one infrared camera, the YOLOv11 network identifies and segments the video data stream of the infrared camera into facial video data and chest and abdominal video data; if the camera system includes multiple infrared cameras, the multiple infrared cameras are fused and combined in time sequence into a video data stream covering the entire area, and the YOLOv11 network identifies and segments the video data stream into facial video data and chest and abdominal video data.

[0013] Preferably, the RPPG signal extraction model processing method includes:

[0014] S21, using a difference layer and a batch normalization layer to perform noise processing on the facial video data, including illumination noise and motion noise; the difference layer calculates the inter-frame difference signal of two consecutive video frames and determines the difference frame, and the batch normalization layer normalizes the difference frame to the same scale and performs noise processing;

[0015] S22. The self-attention mechanism is used to migrate the network. The network is first normalized by a custom normalization module and processed by several two-dimensional convolutional layers to gradually extract the temporal changes in physiological signal features from low-level to high-level. At the same time, the attention mechanism is used to enhance the weights of important temporal physiological signal features, and finally the RPPG signal is output.

[0016] Preferably, the respiratory signal extraction model extracts the respiratory signal method comprising:

[0017] S31, using the optical flow method to extract the chest and abdomen micro-motion change features in the chest and abdomen video data, the chest and abdomen micro-motion change features include chest signals and abdominal signals ;

[0018] S32, chest signal Adopt frequency domain adaptive notch filter Suppressing frequency domain noise ; Abdominal signal Use sparse decomposition to remove motion artifacts to obtain :

[0019] S32, weighted fusion of the processed chest signal and abdomen signal to generate a respiratory signal :

[0020] ,in and are the weights of chest signal and abdomen signal respectively.

[0021] Preferably, the method for extracting features of chest and abdominal micro-motion changes is as follows:

[0022] S311, using polynomials to estimate the neighborhood information of each pixel in the image, converting the pixel gray value into coordinates, the pixel area information The expression is as follows:

[0023] ,in Indicates the coordinates corresponding to the conversion of pixel grayscale values, Represents the approximation of the second-order grayscale derivative of the image, Represents the approximation of the first-order derivative of image grayscale, Indicates the grayscale of the corresponding neighborhood center, and T indicates transposition;

[0024] S312, obtaining preliminary displacement vector between pixel frames , the expression is as follows:

[0025] ,in is the weight function for integrating pixel neighborhoods using weighted least squares, is the average value of the first-order derivative difference of the grayscale of adjacent frame images;

[0026] S313. Construct an 8-parameter parametric displacement model to describe the displacement motion. The expression is as follows:

[0027]

[0028]

[0029] , where d is the pixel displacement vector, P is the motion parameter vector, S is the design matrix, and the weighted least squares solution is performed to obtain the final pixel displacement , the expression is as follows:

[0030] .

[0031] Preferably, the method for obtaining the blood oxygen saturation signal is as follows: the blood oxygen saturation extraction model extracts the time series signal features of the RPPG signal, and performs dimension reduction mapping on the RPPG signal through the convolution layer to obtain the feature ,feature After being processed by the Gaussian error linear unit activation function, batch normalization and feature splicing to strengthen the feature propagation mechanism, the spatiotemporal dimension features are compressed through the global pooling operation. , the predicted blood oxygen saturation signal mapped by the fully connected layer.

[0032] Preferably, the time slice T1 is 30 seconds, and the multi-physiological index fusion recognition model includes a signal decoder, a period mixer and a sequence mixer. The signal decoder downsamples the signal to 1 / 2 of the original length, and uses a tensor reshaping operation and a time-distributed fully connected layer to finally generate a feature vector sequence; the period mixer fuses the feature vectors of multiple modalities into a unified representation of each 30-second time period, and the different modal features are fused through a linear transformation and activation function; the sequence mixer fuses the feature vectors of the multiple modalities into a unified representation of each 30-second time period. The sleep stage classification is performed on 30-second time slices to output sleep stage labels and obtain long-term sleep stage classification results.

[0033] A sleep monitoring system for implementing a sleep monitoring method includes a camera system, a YOLOv11 network, an RPPG signal extraction model, a respiratory signal extraction model, a blood oxygen saturation extraction model, and a multi-physiological index fusion recognition model. The camera system is used to collect a video data stream of a human body sleeping and input it into the YOLOv11 network; the YOLOv11 network is used to detect, segment, and track the face, chest, and abdomen of the video data stream, and obtain facial video data and chest and abdomen video data; the RPPG signal extraction model uses the facial video data to extract the RPPG signal; the respiratory signal extraction model uses the optical flow method to extract the chest and abdomen micro-motion change characteristics in the chest and abdomen video data and At the same time, a frequency domain adaptive notch filtering method is used to remove noise, and a sparse representation time domain filtering method is used for sparse representation and elimination of non-respiratory signals, so as to obtain chest and abdominal movement signals as respiratory signals; the blood oxygen saturation extraction model uses RPPG signals to extract blood oxygen signal features and detects and outputs blood oxygen saturation signals; the multi-physiological indicator fusion recognition model includes four sleep stage stages: wakefulness W, rapid eye movement REM, light sleep N1 / N2, and deep sleep N3. The multi-physiological indicator fusion recognition model uses RPPG signals, respiratory signals, and blood oxygen saturation signals to divide them into time slices T1 for multimodal fusion analysis to obtain long-term sleep stage classification results.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] (1) The present invention uses the YOLOv11 network to capture the face and / or chest and abdomen of a person in a video data stream and obtains facial video data and chest and abdomen video data based on the human body structure relationship. The RPPG signal extraction model obtains a remote photoelectric volume pulse wave signal based on inter-frame difference processing and temporal physiological signal feature extraction processing and separates it to obtain a blood oxygen saturation signal. The respiratory signal extraction model extracts the chest signal and the abdominal signal and performs weighted fusion processing to obtain a chest and abdominal motion signal as a respiratory signal. The multi-physiological indicator fusion recognition model uses the RPPG signal, respiratory signal, and blood oxygen saturation signal to perform multimodal fusion analysis to obtain a long-term sleep stage classification result, thereby achieving high-precision sleep monitoring and sleep stage recognition, and providing technical support for human sleep quality detection and analysis.

[0036] (2) The present invention uses contactless video data sequences to extract RPPG signals, respiratory signals, and blood oxygen saturation signals. The multi-physiological index fusion recognition model is used to fuse and process the data to obtain long-term sleep stage classification results. The sleep monitoring process reduces sleep interference, and the detection and evaluation are not subject to human subjective factors. The subjective bias of the evaluation is low, and it has broad clinical application prospects. It can be widely used in hospitals, nursing homes, and home sleep health monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0038] Below in conjunction with embodiment, the present invention is described in further detail:

[0039] Example

[0040] like Figure 1 As shown, a sleep monitoring method based on a contactless video data sequence includes:

[0041] S1. Capture a video data stream of a sleeping person using a camera system. Use the YOLOv11 network to detect, segment, and track the face, chest, and abdomen of the subject in the video data stream, and obtain facial video data and chest and abdomen video data. The camera system includes at least one infrared camera, and the camera area of ​​the camera system covers the subject's facial triangular area and the chest and abdominal motion area. If the camera system includes one infrared camera, the YOLOv11 network identifies and segments the infrared camera's video data stream into facial video data and chest and abdomen video data. If the camera system includes multiple infrared cameras, the multiple infrared cameras are combined in time sequence to form a video data stream that covers the entire system, and the YOLOv11 network identifies and segments the video data stream into facial video data and chest and abdomen video data. The camera system consists of several infrared cameras fixed vertically to the center of the bed's long axis via adjustable brackets. Installation requirements must meet the following: Horizontally, the optical center of the infrared camera must be located 0.8 to 1.2 meters from the bed's edge; vertically, the lens center must be located 1.0 to 1.3 meters above the mattress surface; and the lens optical axis must be tilted 10 to 15 degrees from the horizontal plane. Observation images must meet the following criteria: 1. Complete coverage of the subject's facial triangular area (the area connecting the center of the eyebrows to the tip of the nose); 2. Simultaneous capture of thoracic and abdominal motion from the sternal notch to the umbilicus; and 3. Minimum 85% coverage of a standard double bed (1.8 x 2.0 meters).

[0042] In some embodiments, the YOLOv11 network is trained using a face dataset. The network comprises a backbone network, a feature enhancement network, and a detection head. The backbone network includes a C3k2 module, which is used to maintain the receptive field and extract basic features. The feature enhancement network includes a joint SPPF-C2PSA module (comprising an SPPF layer and a C2PSA module), which is used to acquire global information (primarily the SPPF layer, which acquires global information through multi-scale pooling) and utilizes a channel-space dual-attention mechanism for feature enhancement (primarily the C2PSA module). The detection head is constructed based on a dual-branch decoupled architecture for classification and regression. It uses a LoU-aware mechanism to predict center point offset, width, and height. It also incorporates a dynamic label assignment strategy and dynamically adjusts the positive and negative sample thresholds through the LoU-aware mechanism. The detection head of the YOLOv11 network outputs detection, segmentation, and tracking results for the face, chest, and abdomen. This embodiment can use a transfer learning strategy to fine-tune the model, which can effectively improve the face detection performance. This embodiment can only detect the face and obtain the face detection frame, and then quickly predict the chest and abdomen area of ​​interest based on the proportional relationship between the face detection frame and the human body structure, and then obtain the chest and abdomen detection frame. Of course, the detection and segmentation of the face, chest and abdomen can also use the YOLOv11 network, and then use the human body structure proportional relationship for verification.

[0043] S2. Construct an RPPG signal extraction model. The RPPG signal extraction model uses facial video data to extract RPPG signals (also known as remote photoplethysmography signals). In some embodiments, the RPPG signal extraction model processing method includes:

[0044] S21. Use the difference layer and batch normalization layer to perform noise processing on the facial video data, including illumination noise and motion noise. The difference layer calculates the inter-frame difference signal of two consecutive video frames and determines the difference frame. The batch normalization layer normalizes the difference frame to the same scale and performs noise processing. The RPPG signal extraction model of this embodiment can also use a time domain shift module to enhance the sequence feature extraction capability. Preferably, the difference layer establishes an optical model and calculates the inter-frame difference signal of two consecutive video frames. The expression is as follows:

[0045] ,in is the difference signal, is the light intensity, is the specular reflection intensity, is the diffuse reflection intensity, The sensor noise signal is modulated by specular reflection, diffuse reflection, and sensor noise. The batch normalization layer normalizes the difference frames to the same scale, significantly amplifying subtle changes in skin pixels and enabling more accurate capture of subtle changes in physiological signals.

[0046] S22. The self-attention mechanism migration network is first normalized by a custom normalization module and processed by several two-dimensional convolutional layers to gradually extract temporal physiological signal features from low-level to high-level. The attention mechanism is used to enhance the weights of important temporal physiological signal features, ultimately outputting the RPPG signal. The self-attention mechanism migration network of this embodiment can also use an average pooling layer to downsample the feature map, reducing the dimensionality and computational complexity, preventing overfitting, and improving the model's generalization capability.

[0047] S3. Construct a respiratory signal extraction model. The respiratory signal extraction model uses optical flow to extract the characteristics of chest and abdomen micro-motion changes in chest and abdomen video data and simultaneously uses frequency domain adaptive notch filtering to filter out noise. In some embodiments, the chest and abdomen micro-motion change feature extraction method is as follows:

[0048] S311, using polynomials to estimate the neighborhood information of each pixel in the image, converting the pixel gray value into coordinates, the pixel area information The expression is as follows:

[0049] ,in Indicates the coordinates corresponding to the conversion of pixel grayscale values , Represents the approximation of the second-order grayscale derivative of the image, Represents the approximation of the first-order derivative of image grayscale, Indicates the grayscale of the corresponding neighborhood center, and T indicates transposition.

[0050] S312, obtaining preliminary displacement vector between pixel frames , the expression is as follows:

[0051] ,in is the weight function for integrating pixel neighborhoods using weighted least squares, is the average value of the first-order derivative difference of the grayscale of the adjacent frames. In each pixel neighborhood, the image grayscale can be approximated by a local quadratic polynomial, and the fitting parameters of the two frames before and after satisfy the following relationship: ,Due to the presence of noise and local changes in the image, direct solution is unstable, so the weighted least squares method is used to integrate the neighborhood pixel information and construct the following error term:

[0052] ,in Represents the preliminary displacement vector between pixel frames , Indicates that weighted least squares method is used to integrate neighborhood pixels. Indicates the pixel displacement change between frames.

[0053] S313. Construct an 8-parameter parametric displacement model to describe the displacement motion. The expression is as follows:

[0054]

[0055]

[0056] , where d is the pixel displacement vector, P is the motion parameter vector, and S is the design matrix. The neighborhood pixels are integrated according to the weighted least squares method. The error term of the weighted least squares for pixel i is expressed as follows: Then perform weighted least squares solution to obtain the final pixel displacement , the expression is as follows:

[0057] .

[0058] A sparse representation time domain filtering method is used to perform sparse representation and eliminate non-respiratory signals, thereby obtaining chest and abdominal motion signals as respiratory signals. In some embodiments, the respiratory signal extraction model extracts the respiratory signal by:

[0059] S31, using the optical flow method to extract the chest and abdomen micro-motion change features in the chest and abdomen video data, the chest and abdomen micro-motion change features include chest signals and abdominal signals , so the optical flow signal set , .

[0060] S32, chest signal Adopt frequency domain adaptive notch filter Suppressing frequency domain noise .Abdominal signal Use sparse decomposition to remove motion artifacts to obtain : The expression is as follows:

[0061] ,in is the Fourier transform, is the inverse Fourier transform, and Represent the dictionary and sparse coefficients in sparse representation respectively.

[0062] S32, weighted fusion of the processed chest signal and abdomen signal to generate a respiratory signal :

[0063] ,in and are the weights of chest signal and abdomen signal respectively, .

[0064] S4. Construct a blood oxygen saturation extraction model. The blood oxygen saturation extraction model uses the RPPG signal to extract blood oxygen signal features and detects and outputs the blood oxygen saturation signal. In some embodiments, the method for obtaining the blood oxygen saturation signal is as follows: the blood oxygen saturation extraction model extracts the time series signal features of the RPPG signal. The blood oxygen saturation extraction model integrates a prediction network that includes DenseNet and ConvMixer design concepts and adaptively reconstructs the one-dimensional time series signal features. The network includes three core processing stages: feature projection, dense convolution, and feature aggregation. The entire network is a regression function: , Indicates the predicted blood oxygen value, is the network parameter. The RPPG signal is mapped to a reduced dimension through the convolution layer to obtain the feature , preferably, the convolution kernel size of the convolution layer is , the step length is , satisfying the requirement that the step size used in the projection process is equal to the convolution kernel size, that is, Special configuration: , is the number of channels, .feature After being processed by the Gaussian error linear unit activation function, batch normalization and feature splicing to strengthen the feature propagation mechanism, the features are: After being processed by the Gaussian Error Linear Unit (GELU) activation function, : , Then perform batch normalization: , by constructing low-dimensional representations, we can effectively capture the key fluctuation patterns in the signal.

[0065] Then, the spatiotemporal dimension features are compressed through global pooling operation , specifically, the spatiotemporal dimension features The method is as follows: using a fully connected hierarchical structure, the output of each convolutional layer is concatenated with the features of all subsequent layers to form an enhanced feature propagation mechanism. The output of each layer is as follows:

[0066] Each basic convolutional module contains a residual structure composed of depthwise separable convolutions, followed by activation layers and batch normalization layers. The module internally concatenates the input features with the batch normalized output channels, expressed as: , and then perform point-by-point convolution operations to form a multi-level feature fusion path: The design of the present invention not only preserves the fine structure of local features, but also enhances the interaction ability of cross-layer information. The spatiotemporal dimension features are compressed by global pooling operation. , the expression is as follows:

[0067]

[0068] The blood oxygen saturation signal predicted by the fully connected layer mapping is expressed as follows: The entire processing flow of the above method realizes the nonlinear mapping from the original waveform to the physiological parameters through the step-by-step abstract signal representation conversion, which improves the robustness of the feature expression while ensuring the computational efficiency.

[0069] S5. Construct a multi-physiological indicator fusion recognition model. The multi-physiological indicator fusion recognition model includes four sleep stages: wakefulness W, rapid eye movement sleep REM, light sleep N1 / N2, and deep sleep N3. The multi-physiological indicator fusion recognition model uses RPPG signals, respiratory signals, and blood oxygen saturation signals to divide them into time slices T1 (in this embodiment, time slice T1 is 30 seconds for example) and performs multimodal fusion analysis to obtain long-term sleep stage classification results. In some embodiments, the multi-physiological indicator fusion recognition model includes a signal decoder, a period mixer, and a sequence mixer. The signal decoder downsamples the signal to 1 / 2 of the original length, and each modality The original signal is represented as: , the number of modalities is M. For each input modality, an independent CNN encoder architecture is used, which consists of stacked residual layers (each residual layer contains three convolutional layers and a subsequent maximum pooling layer, which can downsample the signal to 1 / 2 of the original length. The structure is as follows:

[0070] ,in, Represents a nonlinear activation function (such as ReLU or GELU). This module downsamples the signal to 1 / 2 of its original length.

[0071] By using tensor reshaping operations and time-distributed fully connected layers, a feature vector sequence is finally generated. The feature vector sequence expression is: The epoch mixer fuses the feature vectors of multiple modalities into a unified representation of each 30-second time segment. The fusion input expression is: ,in Represents a vector concatenation operation.

[0072] Different modal features are fused through a linear transformation and activation function: The sequence mixer pair contains The sleep stage classification is performed on 30-second time slices to output sleep stage labels and obtain long-term sleep stage classification results. Apply Transformer or Temporal Convolution modules for global modeling: ,in is the number of sleep stages W, REM, N1 / N2, and N3, and finally outputs the sleep stage prediction label for each 30-second segment:

[0073] ,in It indicates the four sleep stages: wakefulness W, rapid eye movement REM, light sleep N1 / N2, and deep sleep N3. Indicates the total number of sleep stages, , realizing the conversion from multimodal long time series signals to classification results of each 30s video segment, while ensuring the adequacy of feature expression, it also improves cross-modal modeling and global reasoning capabilities.

[0074] A sleep monitoring system for implementing a sleep monitoring method includes a camera system, a YOLOv11 network, a RPPG signal extraction model, a respiratory signal extraction model, a blood oxygen saturation extraction model, and a multi-physiological indicator fusion recognition model. The camera system is used to capture a video data stream of a sleeping person and input it into the YOLOv11 network. The YOLOv11 network is used to detect, segment, and track the face, chest, and abdomen in the video data stream, generating facial and chest and abdominal video data. The RPPG signal extraction model extracts the RPPG signal from the facial video data. The respiratory signal extraction model uses optical flow to extract micro-motion characteristics of the chest and abdomen in the chest and abdomen video data and simultaneously employs frequency-domain adaptive notch filtering to remove noise. A sparse representation time-domain filtering method is used for sparse representation and to eliminate non-respiratory signals, generating a chest and abdomen motion signal as a respiratory signal. The blood oxygen saturation extraction model uses the RPPG signal to extract blood oxygen signal characteristics and detect and output the blood oxygen saturation signal. The multi-physiological indicator fusion recognition model includes four sleep stage classifications: wakefulness W, rapid eye movement REM, light sleep N1 / N2, and deep sleep N3. The multi-physiological indicator fusion recognition model uses RPPG signals, respiratory signals, and blood oxygen saturation signals to perform multimodal fusion analysis according to time slice T1 to obtain long-term sleep stage classification results.

[0075] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A sleep monitoring method based on a non-contact video data sequence, characterized by: The methods include: S1. Collect a video data stream of a sleeping person through a camera system; use the YOLOv11 network to detect, segment, and track the face, chest, and abdomen of the person in the video data stream, and obtain facial video data and chest and abdomen video data; S2. Constructing an RPPG signal extraction model, which uses facial video data to extract RPPG signals; S3. Construct a respiratory signal extraction model. The respiratory signal extraction model uses the optical flow method to extract the micro-motion change characteristics of the chest and abdomen in the chest and abdomen video data and simultaneously uses the frequency domain adaptive notch filtering method to filter out noise. The sparse representation time domain filtering method is used to perform sparse representation and eliminate non-respiratory signals, thereby obtaining the chest and abdomen motion signal as the respiratory signal. S4. Construct a blood oxygen saturation extraction model. The blood oxygen saturation extraction model uses the RPPG signal to extract blood oxygen signal features and detect and output the blood oxygen saturation signal; S5. Construct a multi-physiological indicator fusion recognition model. The multi-physiological indicator fusion recognition model includes four sleep stages: wakefulness W, rapid eye movement REM, light sleep N1 / N2, and deep sleep N3. The multi-physiological indicator fusion recognition model uses RPPG signals, respiratory signals, and blood oxygen saturation signals to perform multimodal fusion analysis according to time slice T1 to obtain long-term sleep stage classification results.

2. The sleep monitoring method based on non-contact video data sequence according to claim 1, characterized in that: The YOLOv11 network is trained using a face dataset. The YOLOv11 network includes a backbone network, a feature enhancement network, and a detection head. The backbone network includes a C3k2 module and is used to maintain the receptive field and extract basic features. The feature enhancement network includes an SPPF-C2PSA joint module and is used to obtain global information and adopt a channel-space dual-attention mechanism for feature enhancement. The detection head is constructed based on a dual-branch decoupling structure of classification and regression. The detection head uses a LoU-aware mechanism to predict the center point offset and width and height, while incorporating a dynamic label allocation strategy and dynamically adjusting the positive and negative sample thresholds through the LoU-aware mechanism. The detection head of the YOLOv11 network outputs detection and segmentation results for the face, chest and abdomen, and tracks them.

3. The sleep monitoring method based on non-contact video data sequence according to claim 1 or 2, characterized in that: The camera system includes at least one infrared camera, and the camera area of ​​the camera system covers the triangular area of ​​the subject's face and the chest and abdominal movement area. If the camera system includes one infrared camera, the YOLOv11 network identifies and segments the video data stream of the infrared camera into facial video data and chest and abdominal video data; if the camera system includes multiple infrared cameras, the multiple infrared cameras are fused and combined in time sequence into a video data stream covering the entire area, and the YOLOv11 network identifies and segments the video data stream into facial video data and chest and abdominal video data.

4. The sleep monitoring method based on non-contact video data sequence according to claim 1, characterized in that: The RPPG signal extraction model processing method includes: S21, using a difference layer and a batch normalization layer to perform noise processing on the facial video data, including illumination noise and motion noise; the difference layer calculates the inter-frame difference signal of two consecutive video frames and determines the difference frame, and the batch normalization layer normalizes the difference frame to the same scale and performs noise processing; S22. The self-attention mechanism is used to migrate the network. The network is first normalized by a custom normalization module and processed by several two-dimensional convolutional layers to gradually extract the temporal changes in physiological signal features from low-level to high-level. At the same time, the attention mechanism is used to enhance the weights of important temporal physiological signal features, and finally the RPPG signal is output.

5. The sleep monitoring method based on non-contact video data sequence according to claim 1, characterized in that: The respiratory signal extraction model extracts the respiratory signal method comprising: S31, using the optical flow method to extract the chest and abdomen micro-motion change features in the chest and abdomen video data, the chest and abdomen micro-motion change features include chest signals and abdominal signals ; S32, chest signal Adopt frequency domain adaptive notch filter Suppressing frequency domain noise ; Abdominal signal Use sparse decomposition to remove motion artifacts to obtain : S32, weighted fusion of the processed chest signal and abdomen signal to generate a respiratory signal : ,in and are the weights of chest signal and abdomen signal respectively.

6. The sleep monitoring method based on non-contact video data sequence according to claim 1 or 5, characterized in that: The method for extracting the characteristics of chest and abdomen micro-motion changes is as follows: S311, using polynomials to estimate the neighborhood information of each pixel in the image, converting the pixel gray value into coordinates, the pixel area information The expression is as follows: ,in Indicates the coordinates corresponding to the conversion of pixel grayscale values, Represents the approximation of the second-order grayscale derivative of the image, Represents the approximation of the first-order derivative of image grayscale, Indicates the grayscale of the corresponding neighborhood center, and T indicates transposition; S312, obtaining preliminary displacement vector between pixel frames , the expression is as follows: ,in is the weight function for integrating pixel neighborhoods using weighted least squares, is the average value of the first-order derivative difference of the grayscale of adjacent frame images; S313. Construct an 8-parameter parametric displacement model to describe the displacement motion. The expression is as follows: ; ; , where d is the pixel displacement vector, P is the motion parameter vector, S is the design matrix, and the weighted least squares solution is performed to obtain the final pixel displacement , the expression is as follows: 。 7. The sleep monitoring method based on non-contact video data sequence according to claim 1, characterized in that: The method for obtaining the blood oxygen saturation signal is as follows: the blood oxygen saturation extraction model extracts the time series signal features of the RPPG signal, and performs dimension reduction mapping on the RPPG signal through the convolution layer to obtain the features ,feature After being processed by the Gaussian error linear unit activation function, batch normalization and feature splicing to strengthen the feature propagation mechanism, the spatiotemporal dimension features are compressed through the global pooling operation. , the predicted blood oxygen saturation signal mapped by the fully connected layer.

8. The sleep monitoring method based on non-contact video data sequence according to claim 1, characterized in that: The time slice T1 is 30 seconds. The multi-physiological index fusion recognition model includes a signal decoder, a period mixer and a sequence mixer. The signal decoder downsamples the signal to 1 / 2 of the original length, and uses a tensor reshaping operation and a time-distributed fully connected layer to finally generate a feature vector sequence; the period mixer fuses the feature vectors of multiple modalities into a unified representation of each 30-second time period, and the different modal features are fused through a linear transformation and activation function; the sequence mixer fuses the feature vectors of the multiple modalities into a unified representation of each 30-second time period. The sleep stage classification is performed on 30-second time slices to output sleep stage labels and obtain long-term sleep stage classification results.

9. A sleep monitoring system for implementing the sleep monitoring method of claim 1, characterized in that: The invention comprises a camera system, a YOLOv11 network, an RPPG signal extraction model, a respiratory signal extraction model, a blood oxygen saturation extraction model and a multi-physiological index fusion recognition model. The camera system is used to collect a video data stream of a human body sleeping and input it into the YOLOv11 network; the YOLOv11 network is used to detect, segment and track the face, chest and abdomen of the video data stream, and obtain facial video data and chest and abdomen video data; the RPPG signal extraction model uses the facial video data to extract the RPPG signal; the respiratory signal extraction model uses the optical flow method to extract the chest and abdomen micro-motion change characteristics in the chest and abdomen video data and simultaneously adopts the frequency domain adaptive trapping The wave filtering method is used to filter out noise, and the sparse representation time domain filtering method is used for sparse representation and elimination of non-respiratory signals, so as to obtain chest and abdominal movement signals as respiratory signals; the blood oxygen saturation extraction model uses RPPG signals to extract blood oxygen signal features and detects and outputs blood oxygen saturation signals; the multi-physiological indicator fusion recognition model includes four sleep stage stagings including wakefulness W, rapid eye movement sleep REM, light sleep N1 / N2, and deep sleep N3. The multi-physiological indicator fusion recognition model uses RPPG signals, respiratory signals, and blood oxygen saturation signals to divide them into time slices T1 for multimodal fusion analysis to obtain long-term sleep stage staging results.

Citation Information

Patent Citations

  • Method and system for non-contact sleep monitoring

    CN104834946A

  • Intelligent sleep monitoring system for three-dimensional multi-dimensional data

    CN114652274A

  • Long-distance non-contact physiological parameter detection method, system and device

    CN117158926A

  • Physiological feature extraction model training method and remote heart rate measurement method

    CN119028003A

  • Non-contact rPPG signal extraction method and system based on space-time attention

    CN119851327A

Cited By

  • Breathing state detection method and device, terminal and storage medium

    CN120884276A

  • A respiratory state detection method, device, terminal and storage medium

    CN120884276B

  • Method, device, equipment, medium and product for predicting task related to sleep

    CN121080919A

  • Children mouth breathing monitoring method, device and equipment and medium

    CN121176862A