Sleep monitoring model training method, sleep monitoring method and equipment
By combining a large-sample unlabeled dataset with a small-sample labeled dataset, and using a pre-trained teacher network model to perform knowledge transfer training on the pure radar feature extractor, the problem of low training efficiency of the millimeter-wave radar sensor sleep monitoring model was solved, and efficient and comfortable sleep monitoring was achieved.
Patent Information
- Application Number
- CN202511267366.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
In the existing technology, the neural network model training efficiency of millimeter wave radar sensors used for sleep monitoring is low, and it is difficult to efficiently obtain labeled training samples, resulting in poor sleep status monitoring performance.
By combining a large-sample unlabeled dataset and a small-sample labeled dataset, the pure radar feature extractor is trained for knowledge transfer using a pre-trained teacher network model. Pulse wave data is used to assist learning, and a small amount of high-labeled value dataset is introduced for fine-tuning. Multi-loss weighted optimization is calculated to correct the feature deviation of the pure radar feature extractor on single-modal radar data.
It achieves efficient and comfortable sleep monitoring, reduces labeling costs, enhances the feature representation capability of radar data by the pure radar modality sleep monitoring model, and can accurately capture key discriminant information of sleep stages and respiratory events.
Smart Images

Figure CN120744690A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radar signal processing, and in particular to a sleep monitoring model training method, a sleep monitoring method and a device. Background Art
[0002] Sleep is a core physiological activity for maintaining human health and is crucial for nervous system repair, metabolic regulation, and immune balance. Sleep disorders are a widespread and serious problem worldwide. They not only affect daytime well-being (such as memory loss and sleepiness) but also increase long-term health risks such as cardiovascular disease and depression, driving a growing demand for accurate and convenient sleep monitoring.
[0003] The current mainstream sleep monitoring technology is polysomnography (PSG), known as the "gold standard." By attaching electrodes and sensors such as electroencephalogram (EEG), electrooculogram (EOG), and electromyography (EMG) to the patient's body surface, it simultaneously collects multi-dimensional physiological signals such as brain waves, respiration, and heart rate, providing a basis for sleep staging and disease diagnosis. However, a limitation of traditional polysomnography is that the monitoring process requires attaching multiple electrodes and connecting sensors to the patient's body surface. The fixation and connection of these components are easily affected by factors such as wear time and changes in body position, resulting in poor contact between electrodes and skin or interruption of sensor signals. Due to the need to maintain device stability, patients are restricted from turning over and adjusting their body position during sleep, and their sleep state deviates from natural patterns, resulting in distorted monitoring data and difficulty in accurately reflecting the true extent of sleep disorders.
[0004] The application of new sensors, such as millimeter-wave radar sensors, has made non-contact, low-invasive sleep monitoring possible. The high-frequency radio signals transmitted and received by millimeter-wave radar sensors can provide information about a person's breathing, heartbeat, and body movements. This information can be used to determine sleep stages, making it suitable for sleep monitoring in both home and medical settings. Its advantages are non-contact, non-intrusive, and highly accurate monitoring.
[0005] To determine sleep stages based on human body information detected by millimeter-wave radar sensors, one feasible solution is to use a neural network model to identify the data. The conventional model training approach for this application scenario is to simultaneously collect PSG and radar data, using the PSG monitoring results as labels to train the neural network model. However, the PSG data collection process is complex, costly, and significantly disruptive to the subjects, making it difficult to quickly obtain a large number of samples.
[0006] Since it is difficult to efficiently obtain labeled training samples, the training efficiency of the neural network model is low, which in turn affects the performance of the model in monitoring sleep status. Summary of the Invention
[0007] In view of this, a first aspect of the present invention provides a sleep monitoring model training method, comprising: Acquiring datasets S1 and S2, wherein dataset S1 includes radar data and pulse wave data synchronously collected by a millimeter-wave radar device and a pulse sensor, and dataset S2 includes radar data, pulse wave data, and PSG data synchronously collected by the millimeter-wave radar device, the pulse sensor, and a PSG device, as well as sleep stage labels and / or respiratory event labels annotated based on the PSG data, wherein the data volume of dataset S2 is smaller than that of dataset S1; A pure radar feature extractor M3 is trained using a pre-trained teacher network model M2 and the data set S1. The training process includes the teacher network model M2 fusing the radar data and pulse wave data in the data set S1 and outputting fused feature data. The pure radar feature extractor M3 extracts features from the radar data in the data set S1 and outputs radar feature data. The loss is then calculated based on the fused feature data and the radar feature data. The data set S2 is used to train a dual-modal sleep monitoring model M4 and a pure radar modality sleep monitoring model M5, wherein the dual-modal sleep monitoring model M4 includes a trained teacher network model M2 and a first recognition layer, and the pure radar modality sleep monitoring model M5 includes a trained pure radar feature extractor M3 and a second recognition layer. The training process includes the first recognition layer determining the first sleep stage and / or respiratory event information based on the fused feature data output by the teacher network model M2, the second recognition layer determining the second sleep stage and / or respiratory event information based on the radar feature data output by the pure radar feature extractor M3, and calculating the loss based on the first sleep stage and / or respiratory event information, the second sleep stage and / or respiratory event information, and the sleep stage label and / or respiratory event label.
[0008] Optionally, the pre-trained teacher network model M2 is obtained as follows: An unsupervised learning method is adopted to train an autoencoder network model M1 using the data set S1. The autoencoder network model M1 includes an encoder, a fusion layer and a decoder. The training process includes the encoder encoding the radar data and pulse wave data in the data set S1 respectively, and converting them into radar feature vectors and pulse wave feature vectors of the same dimension; the fusion layer fuses the radar feature vectors and pulse wave feature vectors to obtain fused feature data; the decoder decodes the fused feature data to obtain reconstructed radar data and reconstructed pulse wave data; the radar data, the pulse wave data, the reconstructed radar data and the reconstructed pulse wave data are used to calculate the loss; the encoder and the fusion layer in the trained autoencoder network model M1 constitute a pre-trained teacher network model M2.
[0009] Optionally, the loss is calculated using the radar data, the pulse wave data, the reconstructed radar data and the reconstructed pulse wave data in the following manner: : ; in, Represents radar data, Indicates pulse wave data, Indicates that at the same time and The reconstructed radar data obtained by inputting the model, Indicates that at the same time and The reconstructed pulse wave data obtained by inputting the model, Indicates that only The reconstructed pulse wave data obtained by inputting the model, 、 、 is the weighting coefficient, express and The mean square error loss, express and The mean square error loss, express and The mean square error loss.
[0010] Optionally, the encoder includes a radar data encoder and a pulse wave data encoder, and the decoder includes a radar data decoder and a pulse wave data decoder. The radar data encoder is used to encode the radar data in the data set S1 to obtain the radar feature vector, and the pulse wave data encoder is used to encode the pulse wave data in the data set S1 to obtain the pulse wave feature vector. The radar data decoder is used to decode the fused feature data to obtain the reconstructed radar data, and the pulse wave data decoder is used to decode the fused feature data to obtain the reconstructed pulse wave data.
[0011] Optionally, the first recognition layer includes a first fully connected layer and a first probability activation layer, the first fully connected layer is used to perform linear transformation and nonlinear mapping on the fused feature data output by the teacher network model M2 to obtain a first classification feature vector that matches the number of sleep stages and / or respiratory event categories, and the first probability activation layer is used to calculate the probability distribution of each sleep stage and / or respiratory event category based on the first classification feature vector to obtain the first sleep stage and / or respiratory event information.
[0012] Optionally, the second recognition layer includes a second fully connected layer and a second probability activation layer, the second fully connected layer is used to perform linear transformation and nonlinear mapping on the radar feature data output by the pure radar feature extractor M3 to obtain a second classification feature vector that matches the number of sleep staging and / or respiratory event categories, and the second probability activation layer is used to calculate the probability distribution of each sleep staging and / or respiratory event category based on the second classification feature vector to obtain second sleep staging and / or respiratory event information.
[0013] Optionally, the sleep monitoring model training method provided by the present invention further includes: preprocessing the data set S1 and the data set S2 respectively, wherein the preprocessed radar data is a two-dimensional range-time spectrum, including three channels of low-frequency energy spectrum, high-frequency energy spectrum and phase Doppler spectrum, and the preprocessed pulse wave data is a two-dimensional frequency-time spectrum, including one channel of energy spectrum.
[0014] A second aspect of the present invention provides a sleep monitoring method based on millimeter wave radar, comprising: Acquire radar data of target objects; The pure radar modality sleep monitoring model M5 trained using any of the sleep monitoring model training methods described above performs feature recognition on the radar data to obtain target sleep stage and / or respiratory event information.
[0015] A third aspect of the present invention provides a sleep monitoring model training device, which includes: a processor and a memory connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the processor to perform the above-mentioned sleep monitoring model training method.
[0016] The fourth aspect of the present invention provides a sleep monitoring device based on millimeter wave radar, which includes: a processor and a memory connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the processor to execute the above-mentioned sleep monitoring method based on millimeter wave radar. The sleep monitoring model training method provided by the present invention first obtains a large sample data set S1 (radar + pulse wave) collected synchronously without annotations, and combines it with a small amount of annotated sample data set S2 (radar + pulse wave + sleep stage label + respiratory event label), taking into account both data scale and annotation accuracy; secondly, the pre-trained teacher network model M2 is used to perform knowledge transfer training on the pure radar feature extractor M3, and the information contained in the pulse wave data is used to help the pure radar feature extractor M3 better learn and understand the radar data, thereby enhancing the feature representation capability of the pure radar feature extractor M3 for radar data, without the need for PSG. This method, which uses either equipment or manual interpretation, addresses the insufficient sample size issue inherent in traditional methods due to the low efficiency and high cost of PSG annotation. A small, highly annotated dataset S2 is then introduced to fine-tune the dual-modality sleep monitoring model M4 and the radar-only sleep monitoring model M5. Feature and classification losses are calculated, and multi-loss weighted optimization is used to correct for feature deviations in the radar-only feature extractor M3 based on unimodal radar data. This allows the radar-only feature extractor M3 to accurately capture key discriminative information about sleep stages and / or respiratory events (e.g., the association between respiratory abnormalities, body movement intensity, and sleep stages) using radar data alone. This fine-tuning process requires only a small amount of annotated data, significantly reducing annotation costs. By leveraging the collaborative strategy of "unlabeled big data pre-training + small sample annotation fine-tuning" and the low overhead of the radar-only modality, this method addresses the issue of insufficient annotated data while avoiding the high cost and high interference of traditional multimodal equipment, achieving efficient and comfortable sleep monitoring.
[0017] The millimeter-wave radar-based sleep monitoring method provided by the present invention collects data only through the non-contact millimeter-wave radar, avoiding the physical discomfort caused to the target object by traditional contact equipment, and is more suitable for long-term and comfortable sleep monitoring needs; and through cross-modal knowledge transfer (using pulse wave data to assist learning) during model training, the pure radar modality sleep monitoring model's feature representation ability for single-modality radar data is enhanced. Without relying on high-load pulse wave or PSG equipment, accurate sleep monitoring can be achieved only through non-contact millimeter-wave radar, enabling it to more accurately capture key discrimination information of sleep stages and / or respiratory events, ultimately achieving efficient and reliable contactless sleep monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 Flowchart of a sleep monitoring model training method in an embodiment of the present invention; Figure 2 : is a structural diagram of the teacher network model M2 in an embodiment of the present invention; Figure 3 : is a structural diagram of the autoencoder network model M1 in an embodiment of the present invention; Figure 4 1 is a structural diagram of a dual-modality sleep monitoring model M4 and a pure radar modality sleep monitoring model M5 in an embodiment of the present invention; Figure 5 4 is a flowchart of a sleep monitoring method based on millimeter wave radar in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] In the description of the present invention, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and should not be understood as indicating or implying relative importance.
[0022] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0023] like Figure 1 As shown, an embodiment of the present invention provides a sleep monitoring model training method, which is executed by an electronic device such as a computer or a server, and specifically includes: S11. Obtain data set S1 and data set S2, where data set S1 includes radar data and pulse wave data synchronously collected by a millimeter-wave radar device and a pulse sensor, and data set S2 includes radar data, pulse wave data, and PSG data synchronously collected by a millimeter-wave radar device, a pulse sensor, and a PSG device, as well as sleep stage labels and / or respiratory event labels marked based on the PSG data. The amount of data in data set S2 is smaller than that in data set S1.
[0024] Millimeter-wave radar and pulse sensors are used to synchronously collect sleep monitoring data of several targets for a corresponding period of time (such as one night). The monitoring data of one person for one night is called a data sample. There is no need to manually label the collected data samples. The collection of all collected data samples constitutes a data set S1.
[0025] Millimeter-wave radar, pulse sensors, and polysomnography (PSG) are used to synchronously collect sleep monitoring data for several target durations (e.g., one night). One night's monitoring data for a person is called a data sample, and each collected data sample is manually labeled (professional technicians interpret and label the PSG data) with sleep stages (e.g., wakefulness, non-rapid eye movement (NREM) sleep stages N1-N3, and rapid eye movement (REM) sleep stages) and / or respiratory events (e.g., obstructive sleep apnea, central sleep apnea, and hypopnea events). Finally, the labeled radar data and pulse wave data are integrated with the corresponding sleep stage labels and / or respiratory event labels (stage / event categories) to form dataset S2 for model training. Due to the low efficiency of collecting polysomnography and labeling in dataset S2, the sample size of dataset S1 is much larger than that of dataset S2.
[0026] S12, use the pre-trained teacher network model M2 and data set S1 to train the pure radar feature extractor M3. The training process includes the teacher network model M2 fusing the radar data and pulse wave data in the data set S1 and outputting the fused feature data, and the pure radar feature extractor M3 extracting features from the radar data in the data set S1 and outputting the radar feature data, and then calculating the loss based on the fused feature data and the radar feature data.
[0027] like Figure 2 As shown, the teacher network model M2 is a pre-trained model. A feature distillation method is used to train a pure radar feature extractor M3 using dataset S1. The inputs to the teacher network model M2 are radar data and pulse wave data, while the input to the student network model M3 is radar data only. The output of the teacher network model M2 is feature data, and the output of the pure radar feature extractor M3 is also feature data. The loss function used in training the pure radar feature extractor M3 is mean squared error loss. During training, the parameters of the teacher network model M2 are fixed, and the feature extraction capabilities of the teacher network model M2 are transferred to the pure radar feature extractor M3.
[0028] S13. Use the data set S2 to train a dual-modal sleep monitoring model M4 and a pure radar modality sleep monitoring model M5, wherein the dual-modal sleep monitoring model M4 includes a trained teacher network model M2 and a first recognition layer, and the pure radar modality sleep monitoring model M5 includes a trained pure radar feature extractor M3 and a second recognition layer. The training process includes the first recognition layer determining the first sleep stage and / or respiratory event information based on the fused feature data output by the teacher network model M2, the second recognition layer determining the second sleep stage and / or respiratory event information based on the radar feature data output by the pure radar feature extractor M3, and calculating the loss based on the first sleep stage and / or respiratory event information, the second sleep stage and / or respiratory event information, the sleep stage label and / or respiratory event label.
[0029] The dual-modal sleep monitoring model M4 and the pure radar modality sleep monitoring model M5 are fine-tuned using the dataset S2. Specifically, the pulse wave data and radar data in the dataset S2 are used to train the dual-modal sleep monitoring model M4, and only the radar data in the dataset S2 is used to train the pure radar modality sleep monitoring model M5. During the training process, flexible parameter adjustment strategies can be adopted for the dual-modal sleep monitoring model M4 and the pure radar modality sleep monitoring model M5: the parameters of the encoder, fusion layer, and pure radar feature extractor M3 can be directly fixed (to avoid destroying the learned multi-modal / single-modal feature extraction capabilities), or the impact of training on the parameters can be reduced by lowering the learning rate or using other learning rate adjustment strategies (such as weight attenuation). Figure 4As shown, the dual-modal sleep monitoring model M4 includes a trained teacher network model M2 and a first recognition layer, and the pure radar modality sleep monitoring model M5 includes a trained pure radar feature extractor M3 and a second recognition layer. The first recognition layer determines the first sleep stage and / or respiratory event information (prediction label 2) based on the fusion feature data output by the teacher network model M2, and the second recognition layer determines the second sleep stage and / or respiratory event information (prediction label 1) based on the radar feature data output by the pure radar feature extractor M3, and then calculates the second sleep stage and / or respiratory event information output by the pure radar modality sleep monitoring model M5. The first classification loss (classification loss 1) and the second classification loss (classification loss 2) are calculated based on the difference (e.g., cross-entropy loss) between the predicted label (predicted label 1), the first sleep stage and / or respiratory event information (predicted label 2), and the true labeled information (sleep stage label and / or respiratory event label). This is used to measure the accuracy of the two in distinguishing sleep stages and / or respiratory events. Furthermore, a feature loss (e.g., mean squared error) is calculated using the fused feature data output by the teacher network model M2 and the radar feature data output by the pure radar feature extractor M3 to ensure that the unimodal features of the pure radar feature extractor inherit the multimodal discriminative information of the teacher network model. The feature loss, the first classification loss, and the second classification loss are then weighted and summed to obtain the total loss. This process is repeated until the total loss converges to a first preset threshold, completing the training of the pure radar modality sleep monitoring model M5.
[0030] This embodiment first obtains a large sample data set S1 (radar + pulse wave) collected synchronously without annotations, and combines it with a small amount of annotated sample data set S2 (radar + pulse wave + sleep stage label + respiratory event label), taking into account both data scale and annotation accuracy; secondly, the pre-trained teacher network model M2 is used to perform knowledge transfer training on the pure radar feature extractor M3, and the information contained in the pulse wave data is used to help the pure radar feature extractor M3 better learn and understand the radar data, thereby enhancing the feature representation ability of the pure radar feature extractor M3 for radar data, without the need for PSG equipment or manual judgment. This method addresses the insufficient sample size issue inherent in traditional methods due to the low efficiency and high cost of PSG annotation. A small, highly annotated dataset S2 is then introduced to fine-tune the dual-modality sleep monitoring model M4 and the radar-only sleep monitoring model M5. Feature and classification losses are calculated, and a multi-loss weighted optimization algorithm is used to correct for the characteristic bias of the radar-only feature extractor M3 on unimodal radar data. This allows the radar-only feature extractor M3 to accurately capture key discriminant information about sleep stages and / or respiratory events (e.g., the association between respiratory abnormalities and body movement intensity and sleep stages) using radar data alone. This fine-tuning process requires only a small amount of annotated data, significantly reducing annotation costs. By leveraging the collaborative strategy of "unlabeled big data pre-training + small sample annotation fine-tuning" and the low overhead of the radar-only modality, this method addresses the issue of insufficient annotated data while avoiding the high cost and high interference of traditional multimodal equipment, achieving efficient and comfortable sleep monitoring.
[0031] In some optional implementations of this embodiment, the pre-trained teacher network model M2 in step S12 is obtained as follows: An unsupervised learning method is used to train the autoencoder network model M1 using the dataset S1. The autoencoder network model M1 includes an encoder, a fusion layer and a decoder. The training process includes the encoder encoding the radar data and pulse wave data in the dataset S1 respectively, and converting them into radar feature vectors and pulse wave feature vectors of the same dimension. The fusion layer fuses the radar feature vector and the pulse wave feature vector to obtain fused feature data. The decoder decodes the fused feature data to obtain reconstructed radar data and reconstructed pulse wave data. The loss is calculated using the radar data, pulse wave data, reconstructed radar data and reconstructed pulse wave data. The encoder and fusion layer in the trained autoencoder network model M1 constitute the pre-trained teacher network model M2.
[0032] like Figure 3As shown, the autoencoder network model M1 comprises an encoder, a fusion layer, and a decoder. The encoder encodes the radar data and pulse wave data in dataset S1 separately, generating radar feature vectors and pulse wave feature vectors, which are then fed into the fusion layer. Because the fusion layer requires joint processing of the radar and pulse wave feature vectors (e.g., through a cross-attention fusion mechanism or feature alignment), only if the two have the same dimensionality can the fusion layer effectively integrate multimodal information and avoid information loss or computational errors caused by dimensionality mismatch. Specifically, the fusion layer employs a cross-attention mechanism to determine the correlation weights between features from different modalities (radar and pulse wave). This weighted summation is then performed to generate fused feature data. The decoder then reconstructs the fused feature data to produce reconstructed radar data and reconstructed pulse wave data. The reconstructed radar data and pulse wave data have the same dimensionality as the radar data in dataset S1, and the reconstructed pulse wave data have the same dimensionality as the pulse wave data in dataset S1.
[0033] The fused feature data is then fed into the decoder, which can adopt a symmetrical structure with the encoder or another structure. The decoder output maintains the same data dimensions as the encoder input. The decoder and encoder share the same structure and dimensions to achieve reverse information recovery (such as convolution transposition and attention alignment) through symmetric structures. This ensures that the compressed features are strictly aligned with the original input in terms of data dimensions, thereby guaranteeing the integrity and accuracy of the reconstructed data.
[0034] The model parameters of the autoencoder network model M1 are optimized by calculating the loss function of the radar data, the pulse wave data, and the reconstructed radar data and the reconstructed pulse wave data until the loss converges to the second preset threshold value, and the training of the autoencoder network model M1 is completed. The encoder and the fusion layer in the trained autoencoder network model M1 are separated to form a teacher network model M2 with multimodal joint feature extraction capability. The structure of the teacher network model M2 is as follows: Figure 2 As shown in the figure, only the encoder and fusion layers need to be retained to form the teacher network model M2. This is because the training goal of the autoencoder network model M1 is to teach the encoder to extract high-quality features from multimodal data (radar + pulse wave) and verify the validity of these features through the decoder's reconstruction task. After training, the encoder has learned the "multimodal feature extraction capability," and the fusion layer has mastered the logic of "integrating multimodal information." The decoder's core role is to assist in training (verifying feature integrity by reconstructing the original data) and does not directly participate in feature extraction. Therefore, the teacher network model only needs to retain the encoder (responsible for feature extraction) and the fusion layer (responsible for multimodal information integration) to achieve joint feature extraction of multimodal data, without the need for a separate decoder.
[0035] This embodiment uses unsupervised learning to train an autoencoder network model M1 using an unlabeled multimodal dataset S1. The encoder learns key features from radar and pulse wave data (such as respiratory rate, body motion, and heart rate) and maps them into feature vectors of the same dimension. The fusion layer integrates the multimodal features through mechanisms such as cross-attention. The decoder verifies the integrity of the feature extraction by reverse-reconstructing the fused features into radar and pulse wave data. A loss constraint (minimizing the reconstruction loss) ensures that the features retain core information. Finally, the encoder and fusion layers form a teacher network model M2. Therefore, the teacher network model M2 is trained solely on the unlabeled, readily available dataset S1 (without relying on costly PSG labeled data). This eliminates reliance on labeled data and lays a solid foundation for feature learning in the radar-only feature extractor M3, significantly enhancing its comprehensive representation of sleep stages and / or respiratory events.
[0036] In some optional implementations of this embodiment, the loss is calculated using radar data, pulse wave data, reconstructed radar data, and reconstructed pulse wave data in the following manner: : ; in, Represents radar data, Indicates pulse wave data, Indicates that at the same time and The reconstructed radar data obtained by inputting the model, Indicates that at the same time and The reconstructed pulse wave data obtained by inputting the model, Indicates that only The reconstructed pulse wave data obtained by inputting the model, 、 、 is the weighting coefficient, express and The mean square error loss, express and The mean square error loss, express and The mean square error loss.
[0037] The loss function used in this embodiment includes two types of loss: multimodal reconstruction error loss and single-modal reconstruction error loss. The two types of loss are weighted summed to obtain the total loss function. 、 ) and the single-modal cross-modal reconstruction error loss ( ) can not only constrain the reconstruction accuracy of radar and pulse wave data respectively to improve the reliability of multimodal feature extraction, but also force the model to learn the correlation between multimodalities through cross-modal reconstruction tasks, thereby enhancing the fusion representation ability of the autoencoder network model M1 for multimodal information and providing a more robust feature basis for the subsequent knowledge transfer of the pure radar feature extractor M3.
[0038] In some optional implementations of this embodiment, the encoders in the above-mentioned autoencoder network model M1, teacher network model M2 and dual-modal sleep monitoring model M4 include a radar data encoder and a pulse wave data encoder, and the decoders include a radar data decoder and a pulse wave data decoder. The radar data encoder is used to encode the radar data in the data set S1 to obtain a radar feature vector, the pulse wave data encoder is used to encode the pulse wave data in the data set S1 to obtain a pulse wave feature vector, the radar data decoder is used to decode the fused feature data to obtain reconstructed radar data, and the pulse wave data decoder is used to decode the fused feature data to obtain reconstructed pulse wave data.
[0039] like Figure 2-4 As shown, the encoders of the autoencoder network model M1, the teacher network model M2, and the dual-modal sleep monitoring model M4 can be divided into a radar data encoder and a pulse wave data encoder. These encoders can employ network structures such as convolutional neural networks and Transformers. Their outputs have the same dimensionality (i.e., the radar data and pulse wave data are mapped into feature spaces of the same dimensionality by their respective encoders). The resulting radar feature vectors and pulse wave feature vectors are then fed into the fusion layer. The decoder in the autoencoder network model M1 can be divided into a radar data decoder and a pulse wave data decoder. The fused feature data output from the fusion layer is fed into the radar data decoder and the pulse wave data decoder for decoding. The radar data decoder can adopt a symmetrical structure with the radar data encoder or other structures. The output of the radar data decoder maintains the same data dimensionality as the input of the radar data encoder. The pulse wave data decoder can adopt a symmetrical structure with the pulse wave data encoder or other structures. The output of the pulse wave decoder maintains the same data dimensionality as the input of the pulse wave data encoder.
[0040] The radar data encoder and pulse wave data encoder in this embodiment can independently process data of different modalities (such as radar data and pulse wave data). Based on the characteristics of radar data and pulse wave data, dedicated networks such as convolution or Transformer can be used to efficiently extract unimodal discriminative features (such as body motion details and heart rate variability) to avoid information confusion. The two encoders output feature vectors of the same dimension, ensuring that the fusion layer can effectively integrate multimodal information through methods such as cross-attention, avoiding losses caused by dimensionality mismatch. The decoder adopts a symmetrical or flexible structure with the encoder, and its output has the same dimension as the original input, ensuring that the M1 autoencoder network model M1 has accurate dual-modal feature extraction capabilities. The radar data encoder and pulse wave data encoder respectively extract unimodal key information (such as respiratory cycle and heart rate variability) based on their respective modal characteristics. The decoder verifies feature integrity through a closed-loop "encoding-decoding" task. If the decoded data is highly consistent with the original input, it indicates that the feature retains core information, ultimately providing an accurate and reliable feature foundation for subsequent sleep staging and respiratory event detection.
[0041] In some optional implementations of this embodiment, the first recognition layer in step S13 includes a first fully connected layer and a first probability activation layer. The first fully connected layer is used to perform linear transformation and nonlinear mapping on the fused feature data output by the teacher network model M2 to obtain a first classification feature vector that matches the number of sleep stages and / or respiratory event categories. The first probability activation layer is used to calculate the probability distribution of each sleep stage and / or respiratory event category based on the first classification feature vector to obtain first sleep stage and / or respiratory event information.
[0042] like Figure 4 As shown, the dual-modal sleep monitoring model M4 includes not only the encoder and fusion layer of the trained teacher network model M2, but also a first recognition layer. This first recognition layer comprises a first fully connected layer and a first probabilistic activation layer (softmax layer). More complex structures may also be employed. The first recognition layer adjusts the feature dimensions of the fused feature data using the fully connected layer, then calculates the probability distribution of each sleep stage or respiratory event using the first probabilistic activation layer (e.g., a softmax layer), ultimately outputting the first sleep stage and / or respiratory event information. By adjusting the feature dimensions through the fully connected layer and calculating the probabilities through the probabilistic activation layer, the first recognition layer efficiently converts the multimodal fused feature data into probabilistic results for sleep stages or respiratory events.
[0043] In some optional implementations of this embodiment, the second recognition layer in step S13 includes a second fully connected layer and a second probability activation layer. The second fully connected layer is used to perform linear transformation and nonlinear mapping on the radar feature data output by the pure radar feature extractor M3 to obtain a second classification feature vector that matches the number of sleep staging and / or respiratory event categories. The second probability activation layer is used to calculate the probability distribution of each sleep staging and / or respiratory event category based on the second classification feature vector to obtain second sleep staging and / or respiratory event information.
[0044] like Figure 4 As shown, the radar-only sleep monitoring model M5 includes not only the trained radar-only feature extractor M3 but also a second recognition layer. This layer consists of a second fully connected layer and a second probabilistic activation layer (softmax layer). More complex structures are also possible. The second recognition layer adjusts the feature dimensions of the radar feature data to match the number of sleep stage or respiratory event categories through the second fully connected layer. The second probabilistic activation layer (such as a softmax layer) then calculates the probability distribution of each sleep stage or respiratory event, ultimately outputting the second sleep stage and / or respiratory event information. Through the second fully connected layer and the second probabilistic activation layer, the second recognition layer efficiently completes the end-to-end mapping from single-modal radar features to sleep stage and / or respiratory event probabilities. Although the input of the radar-only sleep monitoring model M5 consists solely of radar data, its core capabilities (such as the extraction of key features such as respiratory cycles and body movements) derive from knowledge transfer learning from the teacher network model M2 to general feature representations related to sleep staging and / or respiratory events. By transferring this knowledge, the radar-only feature extractor M3 extracts unimodal discriminant features that are highly correlated with the fused features of the teacher network model M2 using only radar data. The second recognition layer then further maps these radar features into specific sleep staging or respiratory event probabilities, ultimately outputting secondary sleep staging and / or respiratory event information. This design avoids reliance on pulse wave or PSG data while retaining the discriminative capabilities of the multimodal model through transfer learning, significantly improving the accuracy and practicality of the radar-only model for sleep staging and / or respiratory event monitoring using unimodal data.
[0045] A sleep monitoring model training method provided by an embodiment of the present invention further includes preprocessing datasets S1 and S2, respectively, wherein the preprocessed radar data is a two-dimensional range-time spectrogram comprising three channels: a low-frequency energy spectrum, a high-frequency energy spectrum, and a phase Doppler spectrum. The preprocessed pulse wave data is a two-dimensional frequency-time spectrogram comprising one channel: an energy spectrum.
[0046] When training the teacher network model M2, the radar-only feature extractor M3, the dual-modal sleep monitoring model M4, and the radar-only sleep monitoring model M5, the radar and pulse wave data in datasets S1 and S2 must be preprocessed. This preprocessing converts the raw radar and pulse wave data into time-frequency spectrograms that are easier for the models to process. The radar data is converted into a two-dimensional range-time spectrogram, specifically consisting of three channels: a low-frequency energy spectrum (respiratory frequency), a high-frequency energy spectrum (body motion), and a phase Doppler spectrum (respiratory dynamics). These channels capture sleep-related information such as respiratory cycles and body motion frequency. The pulse wave data is converted into a two-dimensional frequency-time spectrogram, specifically consisting of a single energy spectrum channel, which captures heart rate variability. After preprocessing, both types of data are unified in format (a two-dimensional time-frequency spectrogram). Together, these two types of data enhance the distinguishability of features related to sleep stages and / or respiratory events, providing higher-quality input for the dual-modal sleep monitoring model and improving training efficiency and feature discriminability.
[0047] This embodiment converts the raw training data into a two-dimensional spectrogram with a joint time-frequency representation through preprocessing. This effectively filters out noise interference and enhances the distinguishability of key features related to sleep stages and / or respiratory events (such as respiratory cycles and heart rate variability). This provides higher-quality input for feature learning of dual-modal sleep monitoring models and pure radar modality sleep monitoring models, thereby improving the discriminability of subsequent features and model training efficiency.
[0048] In some optional implementations of this embodiment, the parameters of the teacher network model M2 are fixed during the training process, and the specific process of migrating the feature extraction capability of the teacher network model M2 to the pure radar feature extractor M3 is as follows: Step 1: Obtain the model parameters of the teacher network model.
[0049] The trained teacher network model is capable of extracting high-quality features from radar and pulse wave multimodal data. The purpose of obtaining its parameters (such as the encoder and fusion layer weights) is to transfer this multimodal feature extraction capability as initial knowledge to the radar-only feature extractor M3, thus avoiding the inefficiency of training the radar-only feature extractor M3 from scratch.
[0050] Step 2: Use the teacher network model to extract features from the radar data and pulse wave data in the dataset S1 to obtain fusion features.
[0051] like Figure 2 As shown in the figure, the teacher network model is used to extract features from radar data and pulse wave data, generating a fused feature. This feature is obtained through joint training on multimodal data and provides a clear learning objective for the radar-only feature extractor: the radar-only feature extractor needs to extract a unimodal feature from the unimodal radar data that is highly correlated with the multimodal feature.
[0052] Step 3: Based on the radar data in the dataset S1, the pure radar feature extractor M3 is trained with model parameters to obtain target features.
[0053] The radar-only feature extractor M3 is initialized using the parameters of the teacher network model (such as the encoder weights) as initial values. Feature distillation is then used to train M3 using only radar data to extract target features, specifically, unimodal discriminative features from a radar-only perspective. Initialization inherits the general feature extraction logic learned by the teacher network model M2, avoiding the inefficiency of training M3 from random parameters. Training only with radar data forces the model to focus on the discriminative features of unimodal data, rather than relying on auxiliary information from the pulse wave.
[0054] Step 4: Calculate the mean square error loss of the fusion feature and the target feature until the mean square error loss converges to a third preset threshold, completing the training of the pure radar feature extractor M3.
[0055] The mean squared error loss is calculated between the fused features (multimodal features) output by the teacher network model and the target features (unimodal features) output by the pure radar feature extractor. Backpropagation is then used to optimize the parameters of the pure radar feature extractor until the loss converges to a third preset threshold. This process, through "feature alignment," ensures that the unimodal features extracted by the pure radar feature extractor are as close as possible to the discriminative features of the multimodal model, thereby inheriting the teacher network model's key logic for distinguishing sleep stages or respiratory events.
[0056] This embodiment obtains the multimodal feature extraction parameters of a trained teacher network model M2 to provide high-quality initialization for a pure radar feature extractor M3. The teacher network model M2 then extracts multimodal fusion features as learning targets, clarifying the target features that the pure radar feature extractor M3 needs to approximate. The pure radar feature extractor M3 is then initialized with the teacher network model parameters and trained using only radar data to extract unimodal target features. Using a mean squared error loss constraint, the unimodal features of the pure radar feature extractor M3 are aligned with the multimodal benchmark features. Ultimately, the pure radar feature extractor M3 learns discriminative features that are highly consistent with the multimodal model using only radar data. This process eliminates the need to collect pulse wave data, significantly reducing the cost and equipment load of simultaneous multimodal data acquisition. While ensuring the pure radar feature extractor M3's high ability to discriminate sleep stages and / or respiratory events, it significantly improves the practicality and scalability of the unimodal monitoring solution.
[0057] like Figure 5 As shown, an embodiment of the present invention provides a sleep monitoring method based on millimeter wave radar, which is executed by an electronic device such as a computer or a server, and specifically includes: S21, acquiring radar data of the target object.
[0058] In practical applications, millimeter-wave radar equipment can be used to perform non-contact sleep monitoring on the target subject, directly collecting radar data (such as respiratory rate, body movement, and other time-series signals) during sleep. Compared to traditional contact devices (such as PSG and portable sensors), millimeter-wave radar does not need to be attached to the human body, avoiding the discomfort caused by electrode contact, making it more suitable for long-term, non-intrusive sleep monitoring scenarios.
[0059] S22 , using the pure radar modality sleep monitoring model M5 trained using the above sleep monitoring model training method to perform feature recognition on the radar data to obtain target sleep stage and / or respiratory event information.
[0060] This step calls the trained pure radar modality sleep monitoring model M5 (consisting of a pure radar feature extractor and a second recognition layer) to process the radar data obtained in step S21: First, the pure radar feature extractor extracts high-level semantic features that are strongly correlated with sleep staging and / or respiratory events from the radar data (such as the stability of the respiratory cycle, the law of signal fluctuation caused by body movement, and other single-modal discrimination information); then, the second recognition layer maps the high-level semantic features to preset sleep staging labels (such as awake, deep sleep, light sleep, REM period, etc.) or respiratory event labels (such as obstructive sleep apnea, central sleep apnea, hypopnea events, etc.), and finally outputs the target sleep staging and / or respiratory event information (that is, the model's predicted label for the current target object's sleep staging or respiratory event).
[0061] The millimeter-wave radar-based sleep monitoring method provided by the present invention collects data only through the non-contact millimeter-wave radar, avoiding the physical discomfort caused to the target object by traditional contact equipment, and is more suitable for long-term and comfortable sleep monitoring needs; and through cross-modal knowledge transfer (using pulse wave data to assist learning) during model training, the pure radar modality sleep monitoring model's feature representation ability for single-modality radar data is enhanced. Without relying on high-load pulse wave or PSG equipment, accurate sleep monitoring can be achieved only through non-contact millimeter-wave radar, enabling it to more accurately capture key discrimination information of sleep stages and / or respiratory events, ultimately achieving efficient and reliable contactless sleep monitoring.
[0062] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0063] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0064] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0066] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A sleep monitoring model training method, characterized in that: include: Acquiring datasets S1 and S2, wherein dataset S1 includes radar data and pulse wave data synchronously collected by a millimeter-wave radar device and a pulse sensor, and dataset S2 includes radar data, pulse wave data, and PSG data synchronously collected by the millimeter-wave radar device, the pulse sensor, and a PSG device, as well as sleep stage labels and / or respiratory event labels annotated based on the PSG data, wherein the data volume of dataset S2 is smaller than that of dataset S1; A pure radar feature extractor M3 is trained using a pre-trained teacher network model M2 and the data set S1. The training process includes the teacher network model M2 fusing the radar data and pulse wave data in the data set S1 and outputting fused feature data. The pure radar feature extractor M3 extracts features from the radar data in the data set S1 and outputs radar feature data. The loss is then calculated based on the fused feature data and the radar feature data. The data set S2 is used to train a dual-modal sleep monitoring model M4 and a pure radar modality sleep monitoring model M5, wherein the dual-modal sleep monitoring model M4 includes a trained teacher network model M2 and a first recognition layer, and the pure radar modality sleep monitoring model M5 includes a trained pure radar feature extractor M3 and a second recognition layer. The training process includes the first recognition layer determining the first sleep stage and / or respiratory event information based on the fused feature data output by the teacher network model M2, the second recognition layer determining the second sleep stage and / or respiratory event information based on the radar feature data output by the pure radar feature extractor M3, and calculating the loss based on the first sleep stage and / or respiratory event information, the second sleep stage and / or respiratory event information, and the sleep stage label and / or respiratory event label.
2. The sleep monitoring model training method according to claim 1, characterized in that: The pre-trained teacher network model M2 is obtained as follows: An unsupervised learning method is adopted to train an autoencoder network model M1 using the data set S1. The autoencoder network model M1 includes an encoder, a fusion layer and a decoder. The training process includes the encoder encoding the radar data and pulse wave data in the data set S1 respectively, and converting them into radar feature vectors and pulse wave feature vectors of the same dimension; the fusion layer fuses the radar feature vectors and pulse wave feature vectors to obtain fused feature data; the decoder decodes the fused feature data to obtain reconstructed radar data and reconstructed pulse wave data; the radar data, the pulse wave data, the reconstructed radar data and the reconstructed pulse wave data are used to calculate the loss; the encoder and the fusion layer in the trained autoencoder network model M1 constitute a pre-trained teacher network model M2.
3. The sleep monitoring model training method according to claim 2, characterized in that: The loss is calculated using the radar data, the pulse wave data, the reconstructed radar data, and the reconstructed pulse wave data as follows: : ; in, Represents radar data, Indicates pulse wave data, Indicates that at the same time and The reconstructed radar data obtained by inputting the model, Indicates that at the same time and The reconstructed pulse wave data obtained by inputting the model, Indicates that only The reconstructed pulse wave data obtained by inputting the model, 、 、 is the weighting coefficient, express and The mean square error loss, express and The mean square error loss, express and The mean square error loss.
4. The sleep monitoring model training method according to claim 2, characterized in that: The encoder includes a radar data encoder and a pulse wave data encoder, and the decoder includes a radar data decoder and a pulse wave data decoder. The radar data encoder is used to encode the radar data in the data set S1 to obtain the radar feature vector, and the pulse wave data encoder is used to encode the pulse wave data in the data set S1 to obtain the pulse wave feature vector. The radar data decoder is used to decode the fused feature data to obtain the reconstructed radar data, and the pulse wave data decoder is used to decode the fused feature data to obtain the reconstructed pulse wave data.
5. The sleep monitoring model training method according to claim 1, characterized in that: The first recognition layer includes a first fully connected layer and a first probability activation layer. The first fully connected layer is used to perform linear transformation and nonlinear mapping on the fused feature data output by the teacher network model M2 to obtain a first classification feature vector that matches the number of sleep stages and / or respiratory event categories. The first probability activation layer is used to calculate the probability distribution of each sleep stage and / or respiratory event category based on the first classification feature vector to obtain the first sleep stage and / or respiratory event information.
6. The sleep monitoring model training method according to claim 1, characterized in that: The second recognition layer includes a second fully connected layer and a second probability activation layer. The second fully connected layer is used to perform linear transformation and nonlinear mapping on the radar feature data output by the pure radar feature extractor M3 to obtain a second classification feature vector that matches the number of sleep staging and / or respiratory event categories. The second probability activation layer is used to calculate the probability distribution of each sleep staging and / or respiratory event category based on the second classification feature vector to obtain second sleep staging and / or respiratory event information.
7. The method according to claim 1, characterized in that Also includes: The data sets S1 and S2 are preprocessed separately, wherein the preprocessed radar data is a two-dimensional range-time spectrogram, which includes three channels: low-frequency energy spectrum, high-frequency energy spectrum, and phase Doppler spectrum; the preprocessed pulse wave data is a two-dimensional frequency-time spectrogram, which includes one channel: energy spectrum.
8. A sleep monitoring method based on millimeter wave radar, characterized in that: include: Acquire radar data of target objects; The pure radar modality sleep monitoring model M5 trained by the method according to any one of claims 1 to 7 performs feature recognition on the radar data to obtain target sleep stage and / or respiratory event information.
9. A sleep monitoring model training device, characterized in that: include: A processor and a memory connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor so that the processor executes the sleep monitoring model training method as described in any one of claims 1-7.
10. A sleep monitoring device based on millimeter wave radar, characterized in that: include: A processor and a memory connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the processor to perform the millimeter-wave radar-based sleep monitoring method as described in claim 8.
Citation Information
Patent Citations
Self-supervised learning-based small-sample space target ISAR (inverse synthetic aperture radar) defocusing compensation method
CN115327544A
Automatic sleep staging method based on visual Transform
CN115374815A
Zero sample cross-modal retrieval method based on Transform network selective distillation
CN115563327A
Radar target identification method based on self-supervised model transfer learning
CN116778352A
Model training method and apparatus based on multi-modal data, and device and storage medium
WO2025140746A2
Cited By
Sleep hypopnea type automatic discrimination method and system based on multi-mode signal
CN121465519A
Sleep awakening detection and evaluation method and device based on radar and PPG signals
CN121533701A
Sleep-wake detection and assessment methods and equipment based on radar and PPG signals
CN121533701B
Multi-dimensional health data fusion analysis system and method
CN122376060A