Newborn sleep state monitoring method and system
By using non-contact monitoring technology that combines millimeter-wave radar and cameras, the system extracts the breathing, heartbeat, and facial features of newborns, solving the problems of allergies and discomfort caused by skin contact in existing technologies, and achieving efficient and accurate sleep state monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV
- Filing Date
- 2026-03-03
- Publication Date
- 2026-04-28
AI Technical Summary
Current neonatal sleep monitoring technology requires prolonged skin contact, which may cause skin allergies, restrict limb movement, and lead to discomfort or crying in infants.
A non-contact monitoring method combining millimeter-wave radar and cameras is adopted. Respiratory and heartbeat signal features are extracted from radar data, facial features are extracted by combining image recognition models, and feature fusion and classification are performed using self-attention mechanisms and machine learning models.
It achieves comfortable monitoring without skin contact or restriction of limb movement, improves the accuracy and robustness of sleep state recognition, adapts to different individual differences, and has good scalability.
Smart Images

Figure CN121926554A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for monitoring the sleep state of newborns. Background Technology
[0002] A newborn's sleep state is closely related to the maturity of their nervous system, the stability of their respiratory and circulatory systems, and their overall health. Therefore, accurate monitoring of sleep states not only provides objective reference for clinical practice but also plays a crucial role in newborn health management and disease prevention. On the one hand, abnormal behaviors during sleep (such as apnea, heart rate imbalance, or sleep cycle disturbances) are often early signs of potential health risks, and timely detection helps reduce the probability of Sudden Infant Death Syndrome (SIDS) and other serious complications. On the other hand, the characteristic distribution of different sleep states reflects the development of the central nervous system, which is particularly critical for assessing the growth and development of premature infants and high-risk newborns. Simultaneously, in clinical and home applications, reliable sleep monitoring technology can help healthcare professionals optimize nursing strategies and enable parents to better understand their infants' circadian rhythms, improving the scientific rigor and safety of daily care. Thus, monitoring newborn sleep states is not only a basic nursing task but also a vital link in promoting precision medicine and health management.
[0003] Currently, neonatal sleep monitoring primarily relies on precise monitoring of electroencephalograms (EEG), electrooculograms (EOG), electromyograms (EMG), electrocardiograms (ECG), respiratory airflow, and blood oxygen saturation. Additionally, wearable devices (chest straps, ankle bracelets, wristbands, etc.) are used to monitor sleep status based on heart rate, blood oxygen levels, and movement.
[0004] However, precise monitoring methods such as electroencephalography (EEG) require electrode patches to be attached to newborns, and wearable devices will also come into contact with the newborn's skin. Both precise monitoring methods like EEG and wearable devices will come into contact with the newborn's skin. Newborns have delicate skin, an incompletely developed stratum corneum, and poor barrier function. Prolonged contact between electrode patches or wearable devices and the skin may lead to skin allergies. At the same time, external foreign objects may restrict limb movement, causing discomfort to the newborn, potentially leading to crying or frequent startling. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a method for monitoring newborn sleep status, which can solve the technical problems of existing newborn sleep status monitoring technology that requires prolonged contact with the newborn's skin, which may lead to skin allergies, restrict limb movement, cause discomfort to the newborn, and cause the baby to cry or be easily startled.
[0006] A first aspect of this invention provides a method for monitoring the sleep state of newborns, comprising: Using millimeter-wave radar and cameras that do not contact the newborn, millimeter-wave radar data and facial video data corresponding to the newborn are collected; Based on the millimeter-wave radar data, respiratory signal characteristics and heartbeat signal characteristics are determined; Facial image features are extracted from the facial video data using an image recognition model; Based on the feature weights of the respiratory signal features, the heartbeat signal features, and the facial image features, feature fusion is performed on the respiratory signal features, the heartbeat signal features, and the facial image features to obtain fused features; Based on the fusion features, the sleep state of the newborn pair is determined, and the sleep state of the newborn pair is output as the sleep state monitoring result of the newborn.
[0007] A second aspect of the present invention provides a newborn sleep state monitoring system, comprising: a processor and a memory; The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the neonatal sleep state monitoring method as described in the first aspect.
[0008] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: (1) In this embodiment of the invention, a non-contact monitoring technology combining millimeter-wave radar and camera is used. It does not require contact with the newborn's skin, protects the newborn's skin, does not restrict the free movement of limbs, avoids discomfort to the newborn, and ensures the baby's comfort and natural sleep state while monitoring the newborn's sleep state.
[0009] (2) In this embodiment of the invention, physiological signals such as breathing and heartbeat, as well as external features such as facial expressions and movements, are acquired simultaneously. Weights are dynamically allocated using a self-attention mechanism to achieve complementary advantages between modalities, thereby improving the accuracy and robustness of sleep state recognition. Combining machine learning models to achieve automatic feature extraction and classification not only improves efficiency but also adapts to individual differences among newborns, possessing good scalability and practical application prospects. Attached Figure Description
[0010] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0011] Figure 1 This is a flowchart illustrating a method for monitoring the sleep state of newborns provided in an embodiment of the present invention.
[0012] Figure 2 This is a schematic diagram of a method for monitoring the sleep state of newborns provided in an embodiment of the present invention.
[0013] Figure 3 This is a schematic diagram of a newborn sleep state monitoring system provided in an embodiment of the present invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] The method for monitoring the sleep state of newborns provided by the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0016] Reference manual attached Figure 1 The diagram shows a flowchart of a method for monitoring the sleep state of newborns provided by an embodiment of the present invention.
[0017] Reference manual attached Figure 2 The diagram shows a structural schematic of a neonatal sleep state monitoring method provided by an embodiment of the present invention.
[0018] This invention provides a method for monitoring the sleep state of newborns, which may include the following steps: S1: Using millimeter-wave radar and camera that do not contact the newborn, collect millimeter-wave radar data and facial video data corresponding to the newborn.
[0019] S2: Determine the respiratory signal characteristics and heartbeat signal characteristics based on the millimeter-wave radar data.
[0020] In one possible implementation, S2 specifically includes sub-steps S201 to S204: S201: Extract minute displacement data caused by breathing and heartbeat from millimeter-wave radar data.
[0021] The micro-displacement microdata characterizes the displacement of the chest and abdomen caused by breathing and heartbeat.
[0022] In one possible implementation, S201 specifically includes sub-steps S2011 to S2014: S2011: Reassemble millimeter-wave radar data into a two-dimensional radar data matrix. The size of the two-dimensional radar data matrix is... M × N , M Indicates the total number of frames. N This indicates the total number of sampling points.
[0023] It should be noted that organizing the raw millimeter-wave radar data into an M×N two-dimensional matrix provides a more intuitive representation of the relationship between time frames and sampling points, facilitating subsequent signal processing and feature extraction. Furthermore, this structured representation promotes parallel computing and batch processing, improving the computational efficiency and scalability of the algorithm.
[0024] S2012: Using a moving target detection algorithm, human targets are distinguished from static backgrounds. Static backgrounds are removed, and clutter suppression is applied to millimeter-wave radar data to obtain a clutter-suppressed radar data matrix. in, S Represents a two-dimensional radar data matrix. S ( m , n ) represents the first element in the two-dimensional radar data matrix. m Line number n Column data, i.e., the first m Frame number n Data from each sampling point This represents the radar data matrix after clutter suppression. The first element in the clutter suppression radar data matrix represents the... m Line number n Column data.
[0025] It's important to note that static background noise is almost identical across different frames, and remains significant even after averaging. Subtracting it cancels out the static background. However, dynamic targets (such as the slight rise and fall of the chest with breathing) exhibit small variations in each frame, and the average cannot fully represent them. Subtracting the average still preserves these dynamic elements in the signal. By subtracting the average, static clutter is suppressed, while the motion information of dynamic targets (displacement caused by breathing / heartbeat) is highlighted.
[0026] Furthermore, by using moving target detection algorithms to distinguish newborn targets from static environmental backgrounds and removing irrelevant interference through clutter suppression, the signal-to-noise ratio (SNR) can be significantly improved. The radar data obtained after this processing is cleaner, ensuring that the subsequently extracted displacement signals truly reflect breathing and heartbeat, rather than environmental noise or reflections from fixed objects.
[0027] S2013: Based on the clutter suppression radar data matrix, summate each column to calculate the cumulative energy value for each sampling point. Based on the cumulative energy value, determine the location of the newborn's body. in, R h The value represents the range of cell indices of the newborn's body, and arg max represents the point where the maximum value is located. E n Indicates the first n The cumulative energy value of each sampling point n 1 indicates the range cell index corresponding to 0.5 meters. n 2 represents the range cell index corresponding to 2 meters.
[0028] It should be noted that, for a fixed distance unit, if a newborn human body exists at that distance, it will generate significant reflected energy in consecutive frames. The sum of the amplitudes of all frames is equivalent to the cumulative value of the target energy at that distance.
[0029] Furthermore, by summing the clutter-suppressed data column by column to calculate the cumulative energy value and finding the energy peak location within a defined distance range (0.5–2m), the location of the newborn within the radar observation range can be quickly and accurately pinpointed. This step avoids blind processing of the entire field of data, significantly reduces computational redundancy, and improves the accuracy and robustness of displacement extraction.
[0030] S2014: Based on millimeter-wave radar data within the range of the newborn's body, the MDACM algorithm is used to extract minute displacement data caused by respiration and heartbeat. in, x n Indicates the first n Displacement values corresponding to each sampling point l f The wavelength of the radar signal. π Represents pi (π). I k-1 Indicates the first k -1 sampling points of the I-channel sampled signal, Q k Indicates the first k Q-channel sampled signal at each sampling point I k Indicates the first k The I-channel sampled signal at each sampling point Q k-1 Indicates the first k -1 sampling points of Q channel sampling signal.
[0031] Among them, the MDACM (Modified Differential and Cross-Multiplication) algorithm is a displacement extraction method based on I / Q channel signal processing. By performing differential and cross-multiplication operations on the I and Q channel signals of adjacent sampling points, it can effectively eliminate phase ambiguity and noise interference, thereby extracting minute displacement changes in millimeter-wave radar echoes with high precision. It is particularly suitable for capturing sub-millimeter displacement signals such as breathing and heartbeat.
[0032] It should be noted that using the MDACM algorithm to demodulate the phase of the I / Q channel signals can effectively extract sub-millimeter-level thoracic displacement caused by respiration and heartbeat. This method is sensitive to small-amplitude periodic signals and can distinguish weak heartbeat and respiratory movements from background noise, laying a solid foundation for subsequent separation of respiratory / heartbeat components and feature extraction.
[0033] S202: Extract respiratory and heartbeat signal components from minute displacement data.
[0034] In one possible implementation, S202 specifically includes sub-steps S2021 to S2024: S2021: A variational optimization problem is constructed to extract respiratory and heartbeat signal components from minute displacement data, using VME decomposition as the objective. in, f This represents the input signal, which is essentially minute displacement data. f ( t )express t Input signal at time, u d This refers to the target signal that has been decomposed, which is either the respiratory signal component or the heartbeat signal component. u d ( t )express t The target signal at that moment, f r Indicates residual signal, f r ( t )express t The residual signal at time t.
[0035] Variational Mode Extraction (VME) is a signal decomposition method. Its core idea is to represent a signal as a combination of the target mode signal and the residual signal, and then extract components near the target frequency using a variational optimization framework. By setting a center frequency and constraints, it iteratively optimizes and solves for the target mode, effectively avoiding the noise and mode aliasing problems of EMD. Furthermore, it is more suitable than VMD for extracting signals with similar frequencies but large amplitude differences, such as separating respiratory and heartbeat signals.
[0036] It should be noted that by representing the minute displacement signal as "target mode + residual signal", the complex mixed signal decomposition problem is formalized into a variational optimization problem. This step clarifies the mathematical objective of separating the respiratory and heartbeat signals, enabling a stable and controllable decomposition process to be achieved within the optimization framework, without relying on empirical filtering or manually set bandwidth.
[0037] S2022: The objective function of the variational optimization problem is set to extract the target signal around a predefined center frequency and minimize the spectral overlap between the residual signal and the target signal. in, J Describe the objective function. oh d The center frequency of the target signal is represented by 'min', which means minimizing it. J 1 represents the first objective function, used to extract the target signal around a predefined center frequency. J 2 represents the second objective function, used to minimize the spectral overlap between the residual signal and the target signal. α Indicates the penalty coefficient. This represents the partial derivative operation. t This represents the operation of partial derivatives with respect to time. d ( t )express t The Dirac distribution at time t is used to represent instantaneous signal changes. j Represents the imaginary unit. This represents the frequency distribution of the signal, where e represents the natural constant. Represents the L2 norm operation. β Represents the frequency response function. β ( t () represents the time-domain representation of the frequency response function. β ( oh This represents the frequency domain representation of the frequency response function. The time domain representation and the frequency domain representation can be converted using Fourier transform. or Represents the regularization factor. oh Indicates frequency.
[0038] It should be noted that by designing the objective function, on the one hand, the target modal signal is made to converge around a preset center frequency, ensuring that the extracted components fall within the frequency band corresponding to respiration or heartbeat; on the other hand, by minimizing the spectral overlap between the residual and the target, modal aliasing is effectively avoided. This ensures that the separated respiratory and heartbeat signals are purer and do not interfere with each other.
[0039] S2023: Using the Lagrange multiplier method, the variational optimization problem is transformed by introducing Lagrange multipliers. in, L Represent the Lagrange function, l Represents the Lagrange multipliers. c Represents the regularization parameter. This represents the inner product operation. l ( t )express t The Lagrange multipliers at each moment will be automatically adjusted according to the constraints to ensure that the signal decomposition meets expectations.
[0040] The Lagrange multiplier method is a mathematical approach for solving constrained optimization problems. It transforms constrained optimization problems into unconstrained ones by introducing Lagrange multipliers into the original objective function, converting the constraints into additional terms. In signal decomposition, this method ensures that the target signal and the residual signal satisfy energy conservation and frequency constraints, facilitating the iterative solution to obtain stable signal components.
[0041] It should be noted that the Lagrange multiplier method can transform the original constrained optimization problem into an unconstrained optimization problem, while explicitly adding constraints to the objective function. This allows the decomposition process to automatically satisfy energy conservation and frequency constraints during optimization, avoiding signal distortion. The resulting Lagrange function provides the mathematical foundation for subsequent iterative solutions, ensuring the interpretability and stability of the decomposition results.
[0042] S024: Using the alternating direction method, an iterative solution to the Lagrange multiplier optimization problem is performed to extract respiratory and heartbeat signal components from small displacement data. in, Indicates the first n The target signal at +1 iteration, f ( oh This represents the frequency domain representation of the input signal. c Represents the regularization parameter. Indicates the first n The center frequency of the target signal at +1 iterations Indicates the first n The target signal at the next iteration. l ( oh ) represents the frequency domain representation of the Lagrange multiplier. Indicates the first n The center frequency of the target signal at the next iteration, where d represents the differential sign. Indicates the first n Lagrange multipliers at +1 iteration, Indicates the first n Lagrange multipliers in the next iteration t This represents the iteration step size parameter.
[0043] The Alternating Direction Method of Multipliers (ADMM) is a highly efficient iterative optimization algorithm. It decomposes a complex optimization problem into multiple subproblems, solves them alternately on different variables, and gradually approximates the global optimum by introducing multipliers. This method has a fast convergence speed and is particularly suitable for large-scale, non-convex optimization problems. In VME decomposition, it can be used to iteratively update the objective mode and Lagrange multipliers.
[0044] It should be noted that the alternating direction method can decompose complex optimization problems into parallelizable and alternating subproblems, gradually approaching the global optimum. Its iterative solution process ensures efficient convergence of the algorithm, while dynamically updating the target mode, center frequency, and Lagrange multipliers. This approach not only improves computational efficiency but also maintains good robustness in noisy environments, making the extraction of respiratory and heartbeat signals more accurate and reliable.
[0045] In this embodiment of the invention, the minute displacement signal is itself a mixed signal resulting from the combined effects of respiration and heartbeat. If not separated, the two rhythms will interfere with each other, making it difficult to accurately identify their respective characteristics. By decomposing the signals, core indicators such as respiratory rate, respiratory amplitude, heart rate, and heart rate variability can be obtained separately, providing accurate data for assessing the sleep state of newborns.
[0046] S203: Extract respiratory signal features of respiratory signal components and heartbeat signal features of heartbeat signal components through wavelet transform.
[0047] Wavelet transform is a time-frequency analysis tool that decomposes a signal into wavelet coefficients in different frequency bands, thereby simultaneously obtaining both time and frequency domain information. Compared to the traditional Fourier transform, wavelet transform can better handle non-stationary signals, making it particularly suitable for analyzing physiological signals that dynamically change over time, such as respiration and heart rate. Through wavelet decomposition, features such as respiratory rate and heart rate can be extracted, allowing for the differentiation of different physiological rhythms.
[0048] In one possible implementation, S203 specifically includes sub-steps S2031 and S2032: S2031: By using wavelet transform, respiratory or heartbeat signal components are decomposed into wavelet coefficients of different frequency bands, with different frequency bands corresponding to different physiological rhythms.
[0049] S2032: Determine respiratory signal characteristics or heartbeat signal characteristics based on wavelet coefficients of different frequency bands.
[0050] Specifically, wavelet coefficients can be used to calculate frequency characteristics, amplitude characteristics, energy characteristics, and variability characteristics.
[0051] In this embodiment of the invention, wavelet coefficients can be used to extract not only respiratory rate and heart rate, but also multi-dimensional features such as amplitude, energy distribution, and variability. These features can more comprehensively reflect the physiological state of newborns, providing reliable data support for subsequent sleep state classification.
[0052] S3: Extract facial image features from facial video data using an image recognition model.
[0053] In one possible implementation, S3 specifically includes sub-steps S301 and S302: S301: Using the MobileNetV3 model, 68 key feature points were extracted from the facial region of the facial video data.
[0054] MobileNetV3 is a lightweight convolutional neural network model that combines depthwise separable convolution, SE attention mechanism, and network architecture search (NAS), significantly reducing computational cost and model parameter count while maintaining high accuracy. It is particularly suitable for real-time image processing in mobile or embedded devices, such as facial landmark detection and feature extraction, and is highly efficient and practical in recognizing facial expressions and movements in newborns.
[0055] Furthermore, the 68 key feature points refer to a set of standard facial landmarks commonly used in face recognition and face analysis tasks. These are derived from the definitions in the dlib library and the iBUG 300-W dataset, and represent a very common annotation method in applications such as facial expression analysis, motion capture, and expression recognition. This annotation method is a very mature existing technology, and will not be elaborated upon further in this invention.
[0056] S302: Calculate facial image features based on the extracted key feature points.
[0057] Optionally, facial image features include: brow furrow degree, corner of mouth droop degree, and eyelid closure rate.
[0058] This invention innovatively identifies several important facial image features for recognizing the sleep state of newborns based on 68 key feature points.
[0059] Optionally, the degree of brow furrowing is specifically defined as the deviation between the actual distance between the inner feature points of the left and right eyebrows and the first reference distance. in, BF Indicates the degree to which the eyebrows are furrowed. b curr This represents the actual distance between the feature points on the inner sides of the left and right eyebrows. b neutral This represents the first baseline distance, the baseline distance between the inner feature points of the left and right eyebrows in a natural, quiet, and awake state. d Indicates distance calculation, P 21 This indicates the coordinates of the feature point on the inner side of the left eyebrow. P 22 This indicates the coordinates of the feature point on the inner side of the right eyebrow.
[0060] It should be noted that newborns often exhibit slight frowning when awake or stimulated by external stimuli, while their faces tend to relax when entering deep sleep. By calculating the change in the distance between the inner points of the left and right eyebrows, the degree of brow contraction can be quantified, thus reflecting whether the newborn is in a calm and relaxed state or experiencing tension, startle, or light sleep. Therefore, the degree of brow furrowing is one of the important external characteristics for judging sleep depth and stability.
[0061] Optionally, the drooping degree of the corners of the mouth is specifically defined as the deviation between the actual vertical distance between the feature points of the left and right corners of the mouth and the line connecting the center of the lips and the second reference distance. in, MCD Indicates the degree of drooping of the corners of the mouth. m curr Indicates the feature point of the left corner of the mouth P 48 Or the feature point at the right corner of the mouth P 54 The actual vertical distance of the line connecting the feature point of the lip center. m neutral This represents the second baseline distance, which is the baseline vertical distance between the feature points at the left and right corners of the mouth and the line connecting the center of the lips in a natural, quiet, and awake state. d y This indicates the calculation of vertical distance. P 48 This indicates the coordinates of the feature point at the left corner of the mouth. C m Indicates the coordinates of the center of the lip. P51 This indicates the coordinates of the center feature point of the upper lip. P 57 This indicates the coordinates of the central feature point of the lower lip.
[0062] It should be noted that the position and shape of the corners of the mouth directly reflect a newborn's facial expression and muscle tone. During sleep, especially in quiet sleep, the corners of the mouth usually droop naturally, while in light sleep or when there is a tendency to cry, the corners of the mouth may twitch noticeably. This change can be quantified by measuring the deviation of the vertical distance between the corner of the mouth and the center of the lips, which can help determine whether a newborn has entered deep sleep or whether there are signs of irritability or impending awakening.
[0063] Optionally, the eyelid closure rate is specifically defined as the degree of deviation between the actual vertical distance between the feature points of the upper and lower eyelids and the third reference distance. in, EAR Indicates eyelid closure rate. r curr This represents the actual vertical distance between feature points on the upper and lower eyelids. r neutral This represents the third reference distance, the baseline vertical distance between feature points of the upper and lower eyelids in a natural, quiet, and awake state. d y This indicates the calculation of vertical distance. P 37 This indicates the coordinates of the feature point on the upper eyelid of the left eye. P 41 This indicates the coordinates of the feature point on the lower eyelid of the left eye.
[0064] It's important to note that eyelid opening and closing is the most direct indicator of sleep state. Newborns have their eyelids open when awake, may experience eye movements and partial opening and closing during light sleep, and their eyelids are completely closed during deep sleep. The degree of eyelid closure can be accurately described by the ratio of the distance between key points on the upper and lower eyelids to a baseline distance, thus directly reflecting the transition between sleep and wakefulness. The eyelid closure rate is also a core indicator for judging sleep stages and sleep continuity.
[0065] Furthermore, the first, second, and third baseline distances are calibrated using facial data of newborns in a natural, quiet, and awake state. Specifically, the individual mean of newborns in a natural, quiet, and awake state can be used as the baseline distance.
[0066] S4: Based on the feature weights of the respiratory signal features, the heartbeat signal features, and the facial image features, feature fusion is performed on the respiratory signal features, the heartbeat signal features, and the facial image features to obtain fused features.
[0067] In one possible implementation, S4 specifically includes the following sub-steps S401 and S402: S401: Combining the Transformer self-attention mechanism, dynamically assigning feature weights to respiratory signal features, heartbeat signal features, and facial image features.
[0068] The Transformer self-attention mechanism dynamically assigns importance weights to different elements by calculating the dependency between any two positions in the input sequence. Compared to traditional recurrent neural networks, it can model global dependencies in parallel, capture long-range information, and has a multi-head mechanism to learn features from multiple subspaces. In multimodal fusion, the self-attention mechanism can be used to dynamically balance the contributions of breathing, heartbeat, and facial features, thereby improving the accuracy and robustness of the discrimination. It should be noted that the Transformer self-attention mechanism is also a very mature existing technology, and will not be elaborated upon in this invention. However, there is no precedent for applying the Transformer self-attention mechanism to neonatal sleep monitoring.
[0069] In one possible implementation, S401 specifically includes sub-steps S4011 to S4014: S4011: Convert the features in the respiratory signal features, heartbeat signal features, and facial image features into query vectors, key vectors, and value vectors through linear transformation.
[0070] S4012: Through each self-attention head, the query vector, key vector, and value vector are linearly transformed into different subspaces corresponding to the self-attention head.
[0071] S4013: Self-attention calculation is performed in each self-attention head.
[0072] S4014: Concatenate the self-attention calculated from each self-attention head, and calculate the feature weights through a linear transformation.
[0073] In this embodiment of the invention, a Transformer self-attention mechanism is introduced to dynamically allocate the weights of respiratory signal features, heartbeat signal features, and facial image features. The advantage of this is that it can adaptively adjust the feature contribution ratio based on the reliability and importance of each modality of information in different time periods and scenarios. This avoids the bias caused by interference from a single modality feature on the overall result; for example, video features may be unreliable in low light conditions, while radar signals are more valuable when the infant is making subtle movements. The self-attention mechanism can automatically capture these differences, assigning higher weights to key modalities, thereby achieving complementary advantages of multimodal information and improving the accuracy and robustness of sleep state recognition. Simultaneously, this dynamic fusion method has good interpretability; the weight distribution can intuitively reflect the system's decision-making basis under different conditions, providing a reference for subsequent optimization and clinical applications.
[0074] S402: Based on feature weights, feature fusion is performed on respiratory signal features, heartbeat signal features and facial image features to obtain fused features.
[0075] In one possible implementation, S402 specifically includes sub-steps S4021 and S4022: S4021: Multiply the respiratory signal features, heartbeat signal features, and facial image features by their respective feature weights.
[0076] S4022: Concatenate the features after multiplying them with the feature weights to form a fused feature vector.
[0077] In this embodiment of the invention, the weight multiplication operation strengthens the most representative features of different modalities, while the concatenation operation ensures that each type of feature is independently expressed in the final fused feature vector. This approach highlights the main information while preserving the complementarity between modalities, thus guaranteeing the integrity of the overall information.
[0078] S5: Based on the fusion features, determine the sleep state of the newborn pair, and output the sleep state monitoring results of the newborn.
[0079] In one possible implementation, S5 specifically involves monitoring the sleep status of newborns using a random forest model based on fused features.
[0080] Random forests are ensemble learning models composed of multiple decision trees that perform classification or regression tasks through voting or averaging of multiple sub-models. During training, they utilize feature subsampling and data subsampling to enhance the model's generalization ability and robustness. Compared to single decision trees, random forests effectively avoid overfitting, are suitable for handling nonlinear feature spaces, and can achieve efficient and stable multi-class discrimination in neonatal sleep state monitoring.
[0081] Specifically, the random forest model uses a fused feature formed by fusing respiratory signal features, heartbeat signal features, and facial image features as input. Each CART tree independently predicts a category (awake / light sleep / deep sleep), and the predictions from all trees are voted on by majority vote; the category with the most votes is the final output. CART trees are a supervised learning decision tree algorithm. For the specific prediction principle of each CART tree, it recursively divides the feature space into two sub-regions, selecting an optimal feature and optimal split point each time, making the resulting child nodes as "pure" as possible (i.e., from the same class of samples). Each internal node corresponds to a judgment condition (e.g., "heart rate <120?"), and each leaf node outputs a predicted category (e.g., "deep sleep"). During prediction, starting from the root node, the prediction proceeds layer by layer down according to the sample features, eventually landing on a leaf node; the category of that leaf node is the prediction result. CART trees are a very mature existing technology, and this invention will not elaborate further.
[0082] In traditional random forests, the overall model stability and prediction results typically improve with increasing decision trees. However, when the decision tree size becomes too large, the training overhead also increases significantly, manifesting as higher computational complexity and longer training time, leading to resource consumption issues. Furthermore, because the Bagging algorithm may generate a large proportion of overlap when sampling to construct sub-training sets, insufficient differentiation between different subsets can result in excessively strong correlations between some decision trees. Excessive correlation weakens the diversity of the random forest, thus limiting the classifier's advantage in improving accuracy. This invention innovatively proposes an improved random forest construction method.
[0083] By calculating the similarity between each CART tree, highly similar trees are identified, avoiding redundant decision trees. When the similarity between two CART trees exceeds a similarity threshold, the CART with higher classification accuracy is retained, while the CART with lower classification accuracy is marked as deletable. Otherwise, no deletion is performed, and both CART trees are retained. All CART trees marked as deletable are sorted in ascending order of classification accuracy, and CART trees are deleted sequentially until a predetermined number of CART trees remain.
[0084] In this embodiment of the invention, a strategy is adopted that first compares similarity and then retains high-quality trees based on classification accuracy. This ensures that only weak trees that contribute little to the final prediction are deleted, while trees that contribute significantly to the overall classification performance are retained. This reduces redundancy without sacrificing model accuracy.
[0085] Optionally, the similarity calculation method between CARTs is as follows: in, SimIndicating similarity, D u Indicates the first u Cart, D v Indicates the first v Cart, arccos Represents the inverse cosine function. W u Indicates the first u A subset of features of a CART tree W v Indicates the first v A subset of features of a CART tree w ua Indicates the first u The first feature subset of the CART tree a One characteristic, w vb Indicates the first v The first feature subset of the CART tree b One characteristic, p Indicates the first u The total number of features in the feature subset of a CART tree. q Indicates the first v The total number of features in the feature subset of a CART tree. IF Indicates an indicator function, only when hour, ,otherwise .
[0086] It's important to note that the similarity between two CART (Classification and Regression Trees) is measured by the degree of overlap in their feature subsets. Each CART decision tree selects a subset of features for splitting during training, denoted as a feature subset. The formula calculates the overlap in feature selection between the two trees using a dot product. An indicator function is used to determine if two features are identical; if identical, the contribution is 1, otherwise 0. The numerator represents the sum of the contributions from features shared by both trees, and the denominator represents the total contribution between all feature pairs. The result is then mapped to an angle using an Arccosine function (0° represents complete similarity, 90° represents orthogonal / completely different), providing an intuitive similarity angle. Lower similarity indicates closer similarity in feature selection between the two trees; higher similarity indicates stronger differences in feature selection between the two trees.
[0087] Optionally, the results of neonatal sleep status monitoring include: wakefulness, light sleep, and deep sleep.
[0088] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: (1) In this embodiment of the invention, a non-contact monitoring technology combining millimeter-wave radar and camera is used. It does not require contact with the newborn's skin, protects the newborn's skin, does not restrict the free movement of limbs, avoids discomfort to the newborn, and ensures the baby's comfort and natural sleep state while monitoring the newborn's sleep state.
[0089] (2) In this embodiment of the invention, physiological signals such as breathing and heartbeat, as well as external features such as facial expressions and movements, are acquired simultaneously. Weights are dynamically allocated using a self-attention mechanism to achieve complementary advantages between modalities, thereby improving the accuracy and robustness of sleep state recognition. Combining machine learning models to achieve automatic feature extraction and classification not only improves efficiency but also adapts to individual differences among newborns, possessing good scalability and practical application prospects.
[0090] Reference manual attached Figure 3 The diagram shows a schematic representation of a neonatal sleep monitoring system provided in an embodiment of the present invention.
[0091] This invention provides a newborn sleep state monitoring system 20, including a processor 201 and a memory 202.
[0092] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-described neonatal sleep state monitoring method and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.
Claims
1. A method for monitoring the sleep state of newborns, characterized in that, include: Using millimeter-wave radar and cameras that do not contact the newborn, millimeter-wave radar data and facial video data corresponding to the newborn are collected; Based on the millimeter-wave radar data, respiratory signal characteristics and heartbeat signal characteristics are determined; Facial image features are extracted from the facial video data using an image recognition model; Based on the feature weights of the respiratory signal features, the heartbeat signal features, and the facial image features, feature fusion is performed on the respiratory signal features, the heartbeat signal features, and the facial image features to obtain fused features; Based on the fusion features, the sleep state of the newborn pair is determined, and the sleep state monitoring results of the newborn are output.
2. The method for monitoring neonatal sleep status according to claim 1, characterized in that, The step of determining respiratory signal characteristics and heartbeat signal characteristics based on the millimeter-wave radar data includes: Extract minute displacement data from the millimeter-wave radar data, wherein the minute displacement data characterizes the displacement of the chest and abdomen due to breathing and heartbeat. Extract the respiratory and heartbeat signal components from the micro-displacement data; The respiratory signal features of the respiratory signal component and the heartbeat signal features of the heartbeat signal component are extracted by wavelet transform.
3. The method for monitoring neonatal sleep status according to claim 1, characterized in that, The step of fusing the respiratory signal features, heartbeat signal features, and facial image features according to their feature weights to obtain fused features includes: By combining the Transformer self-attention mechanism, feature weights are dynamically assigned to the respiratory signal features, the heartbeat signal features, and the facial image features; Based on the feature weights, the respiratory signal features, the heartbeat signal features, and the facial image features are fused to obtain fused features.
4. The method for monitoring neonatal sleep status according to claim 1, characterized in that, Determining the sleep state of the neonatal pair based on the fusion features includes: Based on the fusion features, the sleep status of the newborns is monitored using a random forest model.
5. The method for monitoring neonatal sleep status according to claim 2, characterized in that, The extraction of minute displacement data from the millimeter-wave radar data specifically includes: The millimeter-wave radar data is reassembled into a two-dimensional radar data matrix; By using a moving target detection algorithm, human targets are distinguished from static backgrounds, the static backgrounds are removed, and clutter suppression is applied to the millimeter-wave radar data to obtain a clutter-suppressed radar data matrix. Based on the clutter suppression radar data matrix, each column is summed to calculate the cumulative energy value of each sampling point, and the range of the newborn's body is determined based on the cumulative energy value. Based on millimeter-wave radar data within the range of the newborn's body, the MDACM algorithm is used to extract the minute displacement data caused by breathing and heartbeat.
6. The method for monitoring neonatal sleep status according to claim 2, characterized in that, The extraction of respiratory and heartbeat signal components from the minute displacement data specifically includes: With the goal of extracting the respiratory signal component and the heartbeat signal component from the small displacement data, a variational optimization problem of VME decomposition is constructed. The objective function of the variational optimization problem is set with the goal of extracting the target signal around a predefined center frequency and minimizing the spectral overlap between the residual signal and the target signal. The variational optimization problem is transformed by introducing Lagrange multipliers using the Lagrange multiplier method. The Lagrange multiplier optimization problem is solved iteratively using the alternating direction method to extract the respiratory signal component and the heartbeat signal component from the small displacement data.
7. The method for monitoring neonatal sleep status according to claim 2, characterized in that, The step of extracting respiratory signal features from the respiratory signal components and heartbeat signal features from the heartbeat signal components using wavelet transform specifically includes: By using wavelet transform, the respiratory signal component or the heartbeat signal component is decomposed into wavelet coefficients of different frequency bands, with different frequency bands corresponding to different physiological rhythms; The respiratory signal characteristics or the heartbeat signal characteristics are determined based on wavelet coefficients in different frequency bands.
8. The method for monitoring neonatal sleep status according to claim 1, characterized in that, The step of extracting facial image features from the facial video data using an image recognition model specifically includes: Using the MobileNetV3 model, 68 key feature points were extracted from the facial region of the facial video data; The facial image features are calculated based on the extracted key feature points.
9. The method for monitoring neonatal sleep status according to claim 8, characterized in that, The facial image features include: the degree of brow furrowing, the drooping of the corners of the mouth, and the eyelid closure rate; The degree of brow furrowing is specifically defined as the deviation between the actual distance between the feature points on the inner sides of the left and right eyebrows and the first reference distance. The drooping degree of the corners of the mouth is specifically defined as the degree of deviation between the actual vertical distance between the feature points of the left and right corners of the mouth and the line connecting the center of the lips and the second reference distance. The eyelid closure rate is specifically defined as the degree of deviation between the actual vertical distance between the feature points of the upper and lower eyelids and the third reference distance.
10. The method for monitoring neonatal sleep states according to claim 3, characterized in that, The dynamic allocation of feature weights for the respiratory signal features, the heartbeat signal features, and the facial image features, using the Transformer self-attention mechanism, specifically includes: The respiratory signal features, the heartbeat signal features, and the facial image features are transformed into query vectors, key vectors, and value vectors through linear transformation. By using each self-attention head, the query vector, the key vector, and the value vector are linearly transformed into different subspaces corresponding to the self-attention head; In each of the aforementioned self-attention heads, self-attention calculations are performed. The self-attention calculated from each of the self-attention heads is concatenated, and the feature weights are calculated through a linear transformation.