Non-contact sleep state and sleep quality detection method
Through contactless image detection technology, the target detection network is used to analyze the baby's eye status, which solves the problem of expensive and insufficient subjective reports of existing infant sleep detection equipment, and realizes high-precision sleep quality assessment without device dependence.
Patent Information
- Application Number
- CN202510368741.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-01
AI Technical Summary
The existing sleep detection and quality assessment methods for infants and young children have problems such as expensive equipment, complex operation or relying on subjective reports, resulting in insufficient accuracy and applicability.
Using a contactless method, the baby's images are acquired through the camera and the trained object detection network is used to detect the open and closed state of the baby's eyes, combining brightness enhancement and multi-scale feature extraction technology to automatically analyze the baby's sleep state and quality.
It realizes sleep status and quality detection of infants without professional equipment and personnel intervention, ensures the safety of infants' sleep, reduces the burden on guardians, and provides accurate sleep quality assessment.
Smart Images

Figure CN120412062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and statistical analysis, and particularly to a non-contact method for detecting sleep state and sleep quality. Background Art
[0002] Sleep plays an important role in the growth process of infants and young children. A good sleep state is not only crucial for the healthy development of the bones, organs and immune system of infants and young children, but also a key factor in brain development and the maturity of the central nervous system. Good sleep habits can promote the physical growth and cognitive development of infants and young children, laying a solid foundation for their overall health. However, current sleep quality assessment methods are mostly used for adults or patients, and there are many limitations when applied to infants and young children.
[0003] Currently, the commonly used sleep detection and quality assessment methods at home and abroad include polysomnography (PSG), sleep diary, sleep questionnaire, and activity recording. Among them, PSG is regarded as the "gold standard" for sleep quality assessment. However, this detection method requires professional equipment and technical personnel to operate, is expensive, and for infants and young children, wearing multiple sensors and electrodes may cause discomfort and affect their normal sleep state; sleep diary and sleep questionnaire are a cheap and easy-to-use assessment method, but they rely on long-term observation and subjective reports of guardians, which may not only interfere with the rest of the guardians themselves, but also their accuracy is limited by personal memory bias and subjective judgment; activity recording collects the activity data of infants and young children by wearing an actigraph on the hands of infants and young children and then analyzes the sleep quality. The contact design of this method may cause discomfort to infants and young children, and relying solely on body movement data cannot accurately reflect the sleep state of infants and young children.
[0004] In summary, the existing sleep detection and quality assessment methods either rely on subjective reports, or require complex professional equipment support, or are lacking in accuracy and applicability. Therefore, there is an urgent need to develop a new generation of infant sleep quality monitoring solutions that can not only ensure data accuracy but also be easily applied to the home environment. Summary of the Invention
[0005] In view of the above defects of the prior art, the present invention provides a non-contact method for detecting sleep state and sleep quality to solve the technical problems in the prior art that require the use of professional analysis equipment or have deficiencies in accuracy and applicability.
[0006] To achieve the above and other related objectives, the present invention provides a non-contact method for detecting sleep state and sleep quality, including: obtaining an image to be processed; inputting the image to be processed into a trained target detection network to obtain a baby eye detection result; obtaining the baby's sleep state according to the baby eye detection result; and obtaining the baby's sleep quality according to multiple baby sleep states within a preset time period.
[0007] In an embodiment of the present invention, inputting the image to be processed into a trained target detection network to obtain a baby eye detection result includes: obtaining the brightness and darkness of the image to be processed according to the image to be processed; determining whether the brightness and darkness of the image to be processed is greater than a preset brightness and darkness: if so, inputting the image to be processed into a trained target detection network to obtain a baby eye detection result; if not, performing brightness enhancement processing on the image to be processed and then inputting it into a trained target detection network to obtain a baby eye detection result.
[0008] In an embodiment of the present invention, performing brightness enhancement processing on the image to be processed includes: performing brightness enhancement processing on the image to be processed using a zero-reference depth curve estimation algorithm.
[0009] In an embodiment of the present invention, the target detection network includes a first network for detecting a baby's face and a second network for detecting the open / closed state of the baby's eyes; inputting the image to be processed into a trained target detection network to obtain a baby eye detection result includes: inputting the image to be processed into a trained first network to obtain a baby face detection result; obtaining a baby face image according to the baby face detection result; and inputting the baby face image into a trained second network to obtain a baby eye detection result.
[0010] In an embodiment of the present invention, the structure of the first network is the same as that of the second network, the training data of the first network is an image labeled with the baby's face area, and the training data of the second network is an image labeled with the baby's eye area; the first network includes a backbone feature extraction unit, a lightweight feature pyramid structure, and a detection head; inputting the image to be processed into a trained first network to obtain a baby face detection result includes: using the backbone feature extraction unit to perform multi-level feature extraction on the image to be processed to obtain a feature map with rich semantic information; using the lightweight feature pyramid structure to fuse the feature map output by the backbone feature extraction unit to obtain a multi-scale feature map; and using the detection head to process the multi-scale feature map to obtain the baby face detection result.
[0011] In an embodiment of the present invention, the backbone feature extraction unit includes five stages, and the expressions of the feature maps Y1~Y5 output by each stage are as follows: Y1 = ConvHead(X in ); Y2 = DWBlock(DWBlock(MaxPool(Y1))); Y i+1 = MaxPool(DWBlock(Y i )),i ∈ {2, 3, 4}; where X in is the input feature of the backbone feature extraction unit. ConvHead is composed of a cascaded convolutional layer, batch normalization, ReLU function, and depthwise separable convolution unit. DWBlock is composed of two cascaded depthwise separable convolution units. MaxPool is the max pooling layer. The depthwise separable convolution unit is composed of a cascaded first convolutional layer, second convolutional layer, and batch normalization and ReLU function. The convolutional kernel sizes of the first convolutional layer and the second convolutional layer are 1×1 and 3×3 respectively.
[0012] In an embodiment of the present invention, the lightweight feature pyramid structure is divided into three branches, which respectively process the feature maps output by the last three stages of the backbone feature extraction unit, and its expression is as follows: Y5’ = DWUnit(DWUnit(Y5)); Y4’ = DWUnit(DWUnit(Y4 + Upsample(DWUnit(Y5)))); Y3’ = DWUnit(DWUnit(Y3 + Upsample(DWUnit(Y4))));where Y3’, Y4’, Y5’ are the multi-scale feature maps, DWUnit is the depthwise separable convolution unit, and Upsample is upsampling.
[0013] In an embodiment of the present invention, according to the baby eye detection result, the baby sleep state is obtained, including: according to the baby eye detection result, determining whether the baby eyes are recognized: if so, determining whether both baby eyes are in the closed state: if so, the baby sleep state is asleep, otherwise the baby sleep state is awake; if not, calculating the frame difference between the image to be processed and its previous frame image, and determining whether the frame difference is lower than a preset threshold: if so, the baby sleep state is asleep, otherwise the baby sleep state is awake.
[0014] In one embodiment of the present invention, obtaining the infant sleep quality based on multiple infant sleep states within a preset duration includes: obtaining the infant sleep state corresponding to a time window based on the infant sleep states corresponding to multiple to-be-processed images within the preset time window; and obtaining the infant sleep quality based on the infant sleep states corresponding to multiple time windows within the preset duration.
[0015] In one embodiment of the present invention, obtaining the infant sleep quality based on the infant sleep states corresponding to multiple time windows within a preset duration includes: obtaining the sleep data of the infant within the preset duration based on the infant sleep states corresponding to multiple time windows within the preset duration; obtaining an infant sleep quality index based on the sleep data of the infant within the preset duration; and obtaining the infant sleep quality based on the infant sleep quality index.
[0016] In one embodiment of the present invention, the sleep data includes the total sleep duration, the maximum continuous sleep duration, the total wakefulness duration, and the total number of awakenings; obtaining an infant sleep quality index based on the sleep data of the infant within the preset duration includes: calculating the infant sleep quality index according to the following formula based on the sleep data of the infant within the preset duration: SQ = [(STT + MTSD * 0.5) - (TAT * 0.5 + TAS / 15)] * α p ; where SQ is the infant sleep quality index, STT is the total sleep duration, MTSD is the maximum continuous sleep duration, TAT is the total wakefulness duration, TAS is the total number of awakenings, and α p is a preset constant associated with the age range of infants and young children.
[0017] Advantages of the present invention: A non-contact sleep state and sleep quality detection method proposed by the present invention detects the eyes of an infant through a target detection network, and determines whether the infant is asleep or awake based on the detection result of the infant's eyes, so as to analyze the infant's sleep quality based on this state. Throughout the process, no professional equipment is required, no sensors need to be worn by the infant, and no manual recording by personnel is needed. Only by taking pictures of the infant can automatic identification and recording be carried out, which not only ensures the sleep safety and health of infants and young children, but also reduces the burden on guardians. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 Flow chart of the sleep quality detection method provided by an embodiment of the present invention; Figure 2 The first detailed flow chart of step S200 provided by an embodiment of the present invention; Figure 3 The second detailed flow chart of step S200 provided by an embodiment of the present invention; Figure 4 Detailed flow chart of step S210 provided by an embodiment of the present invention; Figure 5 Architecture diagram of the first network provided by an embodiment of the present invention; Figure 6 Detailed flow chart of step S400 provided by an embodiment of the present invention; Figure 7 Detailed flow chart of step S420 provided by an embodiment of the present invention. Detailed implementation manners
[0020] The following describes the implementation manners of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following examples and the features in the examples can be combined with each other. Except for the specific methods, devices, and materials used in the examples, according to the knowledge of those skilled in the art in the technical field and the description of the present invention, any methods, devices, and materials similar or equivalent to those described in the embodiments of the present invention in the prior art can also be used to implement the present invention.
[0021] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific specific implementation manners, rather than for limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field of the present invention.
[0022] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0023] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of methods and computer program products that can be implemented according to various embodiments disclosed in the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0024] Please refer to Figure 1 , Figure 1 A non-contact sleep state and sleep quality detection method provided by an embodiment of the present invention includes steps S100 to S400.
[0025] Step S100: Obtain the image to be processed. Specifically, the image to be processed can be obtained through a camera. When arranging the camera, the best viewing angle and light conditions of the camera should be manually adjusted to ensure the quality of the captured video or image. It can be understood that the camera can take an image at intervals as the image to be processed, or can directly capture a video, and then extract a specified frame from the video data as the image to be processed. In a specific embodiment of the present invention, the latter scheme is adopted, that is, video data is obtained, and then video frames are obtained by taking intervals of a specified number of frames, so as to obtain the image to be processed. The reason for adopting the method of taking intervals of a specified number of frames is to reduce the data volume and improve the processing speed. The specified number of frames can be, for example, 2 frames.
[0026] In a specific embodiment of the present invention, step S100 includes: obtaining the pressing state of the control button. When the pressing state of the control button is state one, the baby's sleep state is regarded as awake and step S400 is entered. When the pressing state of the control button is state two, the image to be processed is obtained. In this embodiment, setting the control button can prevent the problem of monitoring interruption caused by the baby being picked up after waking up.
[0027] The control button can be set to an active press type. When the guardian wants to pick up the infant, press the button (corresponding to state one), and it will be default that the infant is in a waking state. When the infant is placed down, press the button again (corresponding to state two), and it will continue to automatically obtain the image to be processed and perform subsequent processing. The control button can also be set to a passive type. For example, a pressure monitoring unit is set under the baby mat, so that it can automatically monitor whether the baby is picked up.
[0028] Step S200: Input the image to be processed into the trained target detection network to obtain the baby eye detection result. This step mainly uses the trained target detection network to detect and identify the baby's eyes based on the image to be processed, so as to obtain the baby eye detection result. The reason for detecting the baby's eyes is that the eyes directly reflect the baby's sleep state.
[0029] Please refer to Figure 2 , in a specific embodiment of the present invention, step S200 includes: S201: Obtain the brightness and darkness degree of the image to be processed according to the image to be processed; S202: Determine whether the brightness and darkness degree of the image to be processed is greater than the preset brightness and darkness degree: If so, input the image to be processed into the trained target detection network to obtain the baby eye detection result; If not, perform brightness enhancement processing on the image to be processed and then input it into the trained target detection network to obtain the baby eye detection result. In step S201, the brightness and darkness degree of the image to be processed can be obtained by using the average brightness calculation method. In this embodiment, by calculating the brightness and darkness degree of the image to be processed, it is to avoid the image to be processed being too dark, resulting in inaccurate detection of the baby's eyes.
[0030] In a specific embodiment of the present invention, performing brightness enhancement processing on the image to be processed includes: performing brightness enhancement processing on the image to be processed by using the zero-reference depth curve estimation algorithm. The zero-reference depth curve estimation algorithm is a calculation method that directly estimates depth information from a single image without relying on labeled data or reference information. Its core idea is to infer the scene depth through self-supervised or unsupervised learning, using low-level features of the image itself (such as texture, perspective, occlusion, etc.) or physical priors (such as shadows, blurs, etc.). This method can not only effectively improve the overall brightness of the image to be processed, but also maintain the details and contrast of the image to be processed. Its processing process is roughly as follows: First, extract the relative depth curve from the image to be processed through the zero-reference algorithm (such as texture gradient, shadow analysis) to distinguish the foreground (main body) from the background; Second, dynamically adjust the brightness according to the depth level, enhancing the details (brightening) of the nearby main body and moderately suppressing the distant background (to avoid overexposure); Finally, combine the original image with the adjusted brightness layer, retain the natural transition, and output the enhanced image. This method does not require labeled data and realizes brightness optimization in a lightweight manner.
[0031] Please refer toFigure 3 , in a specific embodiment of the present invention, the target detection network includes a first network for detecting the face of a baby and a second network for detecting the open / closed state of the baby's eyes. Step S200 includes the steps of: S210, inputting the image to be processed into the trained first network to obtain the baby face detection result; S220, obtaining the baby face image according to the baby face detection result; S230, inputting the baby face image into the trained second network to obtain the baby eye detection result. The reason for using two network models here is that the recognition result of the baby's eyes will be more accurate by adopting this two-stage method. Specifically, first use the first network to detect the baby's face, and then identify the eye state of the corresponding face image based on the baby face detection result.
[0032] It can be understood that the above steps S201 - S202 and steps S210 - S230 can coexist. That is, first judge whether to perform brightness enhancement processing on the image to be processed through steps S201 and S202, and then process according to steps S210 - S230.
[0033] In a specific embodiment of the present invention, although the above uses two network models, they are both used for target detection. Therefore, in this embodiment, the structure of the first network is the same as that of the second network. However, their training processes are different. The training data of the first network is an image marked with the baby face area. That is to say, the output result of the first network is at least the position parameters of the baby face detection box, so that the baby face area can be extracted from it. For the second network, its training data is an image marked with the baby eye area. When marking, for example, the area where the baby's eyes are closed can be marked as category one, and the area where the eyes are open can be marked as category two. In this way, the output result of the second network includes at least the category and the position parameters of the detection box of this category. Since the structures of the first network and the second network are the same, the first network will be used as an example for description below.
[0034] Please refer to Figure 4 and Figure 5, in a specific embodiment of the present invention, the first network includes a backbone feature extraction unit, a lightweight feature pyramid structure, and a detection head; step S210 includes: S211, using the backbone feature extraction unit to perform multi-level feature extraction on the image to be processed, obtaining a feature map with rich semantic information; S212, using the lightweight feature pyramid structure to fuse the feature maps output by the backbone feature extraction unit, obtaining a multi-scale feature map; S213, using the detection head to process the multi-scale feature map, obtaining the baby face detection result. In this embodiment, in order to enable the first network to be applied to edge computing devices, the present invention uses depthwise separable convolution instead of ordinary convolution to implement feature extraction. At the same time, in order to achieve accurate recognition of multi-scale targets, the present invention also introduces a feature pyramid structure, thereby obtaining a more accurate recognition result.
[0035] Please refer to Figure 5 , in a specific embodiment of the present invention, the backbone feature extraction unit includes five stages, and the expressions of the feature maps Y1~Y5 output by each stage are as follows: Y1 = ConvHead(X in ); Y2 = DWBlock(DWBlock(MaxPool(Y1))); Y i+1 = MaxPool(DWBlock(Y i , i ∈ {2, 3, 4}, this formula corresponds to three formulas, and a simplified representation is used here.
[0036] In the above formula, X in is the input feature of the backbone feature extraction unit. ConvHead is composed of cascaded convolutional layers (with a convolutional kernel size of 3×3), batch normalization, and a rectified linear unit (corresponding to Figure 5 BN&Relu in Figure 5 ), and a depthwise separable convolution unit (DWUnit). DWBlock is composed of two cascaded depthwise separable convolution units. MaxPool is a max pooling layer.
[0037] The depthwise separable convolution unit (DWUnit) consists of a cascaded first convolutional layer, a second convolutional layer, batch normalization, and a rectified linear unit. The convolutional kernels of the first convolutional layer and the second convolutional layer are 1×1 and 3×3 respectively. In the depthwise separable convolution unit, pointwise convolution is first performed using the first convolutional layer with a 1×1 convolutional kernel, then depthwise convolution is performed using the second convolutional layer with a 3×3 convolutional kernel. Finally, batch normalization (Batch Normalization, BN) and the Relu activation function are used to stabilize the training process and introduce non-linearity, thereby enhancing the expressive power and generalization performance of the model.
[0038] Please refer to Figure 5 , in a specific embodiment of the present invention, the lightweight feature pyramid structure is divided into three branches, which respectively process the feature maps output by the last three stages of the backbone feature extraction unit. The expression is as follows: Y5’ = DWUnit(DWUnit(Y5)); Y4’ = DWUnit(DWUnit(Y4 + Upsample(DWUnit(Y5)))); Y3’ = DWUnit(DWUnit(Y3 + Upsample(DWUnit(Y4)))); Among them, Y3’, Y4’, and Y5’ are multi-scale feature maps, DWUnit is the depthwise separable convolution unit, and Upsample is upsampling. In this embodiment, a lightweight feature pyramid structure is constructed using the last 3 stages of the backbone feature extraction unit. A series of new feature maps with the same resolution but different semantic intensities are formed through the top-down path, and the backbone network features at the same level are combined with the features generated by the top-down path using lateral connections to ensure that each layer can obtain sufficient spatial details and semantic information. At the same time, a smoothing layer is used to eliminate the upsampling aliasing effect and improve the feature smoothness.
[0039] In a specific embodiment of the present invention, three detection heads are provided, which respectively process the features output by three branches of the lightweight feature pyramid structure to obtain object detection results at different scales. In this embodiment, based on the multi-scale feature maps output by the feature pyramid, a task decoupling mechanism is adopted to construct three detection heads at different scales, so as to obtain object detection results at different scales. The "task decoupling mechanism" means that in the object detection task, different subtasks (such as classification and localization) are separated, so that each task can be optimized independently, thereby improving the detection performance. The outputs of different detection heads are only object detection results at different scales. For each detection head, its output generally includes object category, confidence, detection box position parameters, and IoU (Intersection over Union). With these recognition results, the next step of processing can be carried out.
[0040] In a specific embodiment of the present invention, since the first network or the second network outputs detection results at different scales, a post-processing step is introduced. The Non-Maximum Suppression (NMS) algorithm is a commonly used post-processing step in object detection for removing redundant detection boxes. Specifically, NMS filters the detection boxes according to the confidence and IoU of the detection boxes, retains the detection box with the highest confidence, and removes other detection boxes whose IoU exceeds a certain threshold.
[0041] The traditional NMS algorithm uses a fixed IoU threshold (such as 0.5 or 0.7) to filter the detection boxes. However, the fixed threshold may not be able to meet the requirements of different scenarios or different object scales. For example, for dense objects, the fixed IoU threshold may cause too many detection boxes to be retained, while for sparse objects, it may over-remove detection boxes. The adaptive threshold NMS optimizes the detection results by dynamically adjusting the IoU threshold. In the present invention, the IoU threshold is set to a range (0.5 - 0.7), which means that the NMS algorithm will dynamically select the most suitable IoU threshold according to the specific situation. For example: for dense objects, a higher IoU threshold (such as 0.7) can be selected to retain more detection boxes. For sparse objects, a lower IoU threshold (such as 0.5) can be selected to avoid over-removing detection boxes.
[0042] For the first network, after post - processing, the baby face detection result can be obtained. This result includes the position parameters of the detection box of the baby face image. According to these position parameters, the baby face image can be obtained, and then the baby face image is input into the second network. The second network will output the category (open eyes, closed eyes) and its confidence, the position parameters of the detection box, and the IoU. After post - processing, the baby eye detection result can be obtained. The baby eye detection result includes: no baby eyes detected, baby eyes detected, and the open / closed state of the eyes.
[0043] In a specific embodiment of the present invention, since there is less baby data, a large amount of adult face data can be used to train the pre - trained weights first, and then transfer learning training is carried out on the infant and toddler face dataset. The model is fine - tuned to adapt to the uniqueness of the infant and toddler facial features. This can effectively address the problems of large changes in infant and toddler facial features and limited sample size, and ensure the high accuracy and strong generalization ability of the model.
[0044] Step S300: Obtain the baby's sleep state according to the baby eye detection result.
[0045] In a specific embodiment of the present invention, step S300 includes: judging whether the baby's eyes are recognized according to the baby eye detection result: If so, judge whether both baby's eyes are in the closed - eye state: if so, the baby's sleep state is asleep; otherwise, the baby's sleep state is awake; If not, calculate the frame difference between the image to be processed and its previous frame, and judge whether the frame difference is lower than the preset threshold: if so, the baby's sleep state is asleep; otherwise, the baby's sleep state is awake.
[0046] Multiple judgments are required in this embodiment. First, it is necessary to judge whether the baby's eyes are detected. Due to the baby's sleeping position, it may not be possible to photograph the baby's eyes, or during the recognition of the baby's eyes, accurate recognition may not be achieved, both of which may result in no baby's eyes detected. If the baby's eyes are detected, it is necessary to further judge whether the baby is asleep according to the open / closed state of the baby's eyes. If no baby's eyes are detected, at this time, it is impossible to judge whether the baby is asleep through the baby's eyes, so the frame difference method is used to judge whether the baby has any movement, which can ensure that the baby's sleep state can be accurately judged even when the baby's eyes are not accurately recognized.
[0047] Step S400: Obtain the baby's sleep quality according to the baby sleep states within a preset duration. When the baby's sleep state can be judged, the baby sleep states within the preset duration can be statistically analyzed to obtain the baby's sleep quality. In this step, the preset duration can be, for example, 12 hours, or 24 hours, or a specified time, such as 8 pm to 9 am, etc.
[0048] Please refer toFigure 6 , in a specific embodiment of the present invention, step S400 includes: S410. Obtain the baby's sleep state corresponding to the time window according to the baby's sleep states corresponding to multiple images to be processed within a preset time window; S420. Obtain the baby's sleep quality according to the baby's sleep states corresponding to multiple time windows within a preset duration.
[0049] In steps S100 - S300, the baby's sleep state corresponding to a certain image to be processed can be obtained through processing, that is, each image corresponds to a baby's sleep state. Assume the frame rate of the baby video data is 15 FPS. If the images to be processed are obtained from the video data at an interval of 2 frames, then it is equivalent to extracting 5 frames per second. After being processed through steps S100 - S300, 5 baby's sleep states can be obtained per second. To improve the accuracy of the baby's sleep state, the concept of a preset time window is introduced in step S410. Assume the preset time window is 5 seconds, then there will be 25 baby's sleep states within these 5 seconds. The data volume is too large. Through the processing of step S410, these 25 baby's sleep states can be processed into 1 final baby's sleep state. When specifically processing, for example, the mode value method can be applied. It is a method for making decisions or filling based on the value with the highest frequency of occurrence in the data set. In step S410, for example, if the number of awake states among these 25 baby's sleep states is greater than or equal to 13, it means that within these 5 seconds, the baby is awake.
[0050] Please refer to Figure 7 , in a specific embodiment of the present invention, step S420 includes steps S421 - S423.
[0051] Step S421. Obtain the baby's sleep data within a preset duration according to the baby's sleep states corresponding to multiple time windows within the preset duration. The baby's sleep states corresponding to multiple time windows within the preset duration can be, for example, {awake, awake, awake, sleep, sleep, sleep,...}. The time corresponding to each baby's sleep state is the duration of the time window, which is 5 seconds in the above - mentioned embodiment. Therefore, the sleep data can be obtained based on this information.
[0052] In a specific embodiment of the present invention, the sleep data includes the total sleep duration, the maximum continuous sleep duration, the total awake duration, and the total number of awakenings. Among them, the total sleep duration = the number of sleep states × the time window duration, the maximum continuous sleep duration = the number of the most consecutive sleep states × the time window duration, the total awake duration = the number of awake states × the time window duration, and the total number of awakenings = the number of times changing from sleep to awake.
[0053] Step S422. Obtain the baby's sleep quality index according to the baby's sleep data within the preset duration.
[0054] In a specific embodiment of the present invention, for example, the infant sleep quality index can be calculated according to the following formula: SQ = [(STT + MTSD * 0.5) - (TAT * 0.5 + TAS / 15)] * α p ; where SQ is the infant sleep quality index, STT is the total sleep duration, MTSD is the maximum continuous sleep duration, TAT is the total wakefulness duration, TAS is the total number of awakenings, and α p is a preset constant associated with the age range of infants and young children.
[0055] In a specific embodiment of the present invention, α p can be preset as follows: for infants aged 0 - 6 weeks, it is 8.5; for infants aged 6 weeks - 3 months, it is 9.5; for infants aged 4 - 8 months, it is 10; for infants aged 9 - 24 months, it is 11; for infants aged 24 - 36 months, it is 10.5.
[0056] Step S423: Obtain the infant sleep quality based on the infant sleep quality index. The infant sleep quality index SQ is a specific calculated value, and the infant sleep quality can be obtained by processing it.
[0057] In a specific embodiment of the present invention, for example, it can be divided and evaluated as follows: when SQ ≥ 90, the infant sleep quality is excellent, indicating that the infant has sufficient sleep duration, high continuity of continuous sleep, and few awakenings; when 80 ≤ SQ < 90, the infant sleep quality is good, indicating that the infant's sleep duration is basically up to standard, with occasional short awakenings; when 60 ≤ SQ < 80, the infant sleep quality is medium, with insufficient sleep duration or frequent awakenings, and environmental factors need to be concerned; when SQ < 60, the infant sleep quality is poor, indicating that the infant's sleep is severely fragmented, and it is recommended to seek medical help.
[0058] In addition, a visual analysis chart can be established, using a pie chart to present the sleep stage distribution (the time proportion of the sleep and wakefulness stages), using a timeline to mark the start time and duration of the sleep stage, and a historical sleep quality trend comparison curve (the change of SQ values for consecutive days) can be presented.
[0059] It should be noted that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they contain the same logical relationship, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of the algorithm and process, are all within the protection scope of this patent.
[0060] Generally speaking, the present invention utilizes a camera to collect real-time video frames in the sleep environment of infants and young children, and through a series of optimized image processing and deep learning models, it ensures the accurate judgment of the eye opening and closing conditions of infants and young children under different lighting conditions, so as to distinguish their sleep or waking states. In addition, for cases where the deep learning algorithm cannot effectively detect the eyes, it will automatically switch to the frame difference method, and further confirm whether the infant or young child is in a sleep or waking state by analyzing the differences between two frames. During the real-time monitoring process, if it is detected that the infant or young child wakes up, a reminder can be immediately sent to the guardian, and they are allowed to view the state of the infant or young child in real time through remote monitoring, so as to make a timely response. In order to obtain a scientific and reasonable sleep quality assessment result, the present invention filters and statistically analyzes the state recognition results saved during the measurement time, and adopts the mode value method to ensure the accuracy of the final assessment. At the same time, the system will conduct statistical analysis on key indicators such as the total sleep time (STT), the maximum continuous sleep duration (MTSD), the total wake time (TAT), and the total number of wake times (TAS), and generate a detailed sleep quality report for the guardian's reference in combination with the sleep quality assessment algorithm.
[0061] In addition, the present invention particularly considers the special requirements in the actual application scenario, such as preventing the interruption of monitoring caused by picking up the infant or young child. There is a special control button that allows the guardian to manually / automatically trigger the system to enter or exit the corresponding monitoring state when necessary. Through these innovative designs and technical means, the present invention effectively solves the limitations existing in the existing sleep detection methods, brings an innovative solution to the field of infant and young child sleep quality monitoring, not only guarantees the sleep safety and health of infants and young children, but also reduces the burden on guardians, and can provide new ideas and technical support for the research and development of intelligent guardianship systems.
[0062] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A non-contact sleep state and sleep quality detection method, characterized in that, Including: Obtain the image to be processed; Input the image to be processed into the trained target detection network to obtain the baby eye detection result; Obtain the baby's sleep state according to the baby eye detection result; Obtain the baby's sleep quality according to multiple baby sleep states within a preset duration.
2. The non-contact sleep state and sleep quality detection method according to claim 1, characterized in that, Inputting the image to be processed into the trained target detection network to obtain the baby eye detection result includes: Obtain the brightness of the image to be processed according to the image to be processed; Determine whether the brightness of the image to be processed is greater than the preset brightness: If so, input the image to be processed into the trained target detection network to obtain the baby eye detection result; If not, perform brightness enhancement processing on the image to be processed and then input it into the trained target detection network to obtain the baby eye detection result.
3. The non-contact sleep state and sleep quality detection method according to claim 1, characterized in that, The target detection network includes a first network for detecting the baby's face and a second network for detecting the open / closed state of the baby's eyes; Inputting the image to be processed into the trained target detection network to obtain the baby eye detection result includes: Input the image to be processed into the trained first network to obtain the baby face detection result; Obtain the baby face image according to the baby face detection result; Input the baby face image into the trained second network to obtain the baby eye detection result.
4. The non-contact sleep state and sleep quality detection method according to claim 3, characterized in that, The structure of the first network is the same as that of the second network. The training data of the first network is the image marked with the baby face area, and the training data of the second network is the image marked with the baby eye area; The first network includes a backbone feature extraction unit, a lightweight feature pyramid structure, and a detection head; Inputting the image to be processed into the trained first network to obtain the baby face detection result includes: Use the backbone feature extraction unit to perform multi-level feature extraction on the image to be processed to obtain a feature map with rich semantic information; Use the lightweight feature pyramid structure to fuse the feature maps output by the backbone feature extraction unit to obtain a multi-scale feature map; Use the detection head to process the multi-scale feature map to obtain the baby face detection result.
5. The non-contact sleep state and sleep quality detection method according to claim 4, wherein The backbone feature extraction unit includes five stages, and the expressions of the feature maps Y1~Y5 output by each stage are as follows: Y1 = ConvHead(X in ); Y2 = DWBlock(DWBlock(MaxPool(Y1))); Y i+1 =MaxPool(DWBlock(Y i )),i∈{2,3,4}; Among them, X in is the input feature of the backbone feature extraction unit. ConvHead is composed of cascaded convolutional layers, batch normalization and ReLU functions, and depthwise separable convolution units. DWBlock is formed by cascading two depthwise separable convolution units. MaxPool is the max pooling layer. The depthwise separable convolution unit is composed of cascaded first convolutional layer, second convolutional layer, batch normalization and ReLU functions. The kernel sizes of the first convolutional layer and the second convolutional layer are 1×1 and 3×3 respectively.
6. The non-contact sleep state and sleep quality detection method according to claim 5, characterized in that The lightweight feature pyramid structure is divided into three branches, which respectively process the feature maps output by the last three stages of the backbone feature extraction unit, and its expression is as follows: Y5’ = DWUnit(DWUnit(Y5)); Y4’ = DWUnit(DWUnit(Y4 + Upsample(DWUnit(Y5)))); Y3’ = DWUnit(DWUnit(Y3 + Upsample(DWUnit(Y4)))); Wherein, Y3’, Y4’, Y5’ are the multi-scale feature maps, DWUnit is the depthwise separable convolution unit, and Upsample is upsampling.
7. The non-contact sleep state and sleep quality detection method according to claim 1, wherein Based on the baby eye detection results, obtain the baby's sleep state, including: Based on the baby eye detection results, determine whether the baby's eyes are recognized: If so, determine whether both of the baby's eyes are in a closed state: If so, the baby's sleep state is asleep; otherwise, the baby's sleep state is awake; If not, calculate the frame difference between the image to be processed and its previous frame, and determine whether the frame difference is lower than a preset threshold: If so, the baby's sleep state is asleep; otherwise, the baby's sleep state is awake.
8. The non-contact sleep state and sleep quality detection method according to claim 1, wherein Based on multiple baby sleep states within a preset duration, obtain the baby's sleep quality, including: Based on the baby sleep states corresponding to multiple images to be processed within a preset time window, obtain the baby sleep state corresponding to the time window; Based on multiple baby sleep states corresponding to the time windows within a preset duration, obtain the baby's sleep quality.
9. The non-contact sleep state and sleep quality detection method according to claim 8, characterized in that Based on multiple baby sleep states corresponding to the time windows within a preset duration, obtain the baby's sleep quality, including: Based on multiple baby sleep states corresponding to the time windows within a preset duration, obtain the sleep data of the baby within the preset duration; Based on the sleep data of the baby within the preset duration, obtain the baby sleep quality index; Based on the baby sleep quality index, obtain the baby's sleep quality.
10. The non-contact sleep state and sleep quality detection method according to claim 9, characterized in that, The sleep data includes total sleep duration, maximum continuous sleep duration, total awake duration, and total number of awakenings; Based on the sleep data of the baby within the preset duration, obtain the baby sleep quality index, including: Based on the sleep data of the baby within the preset duration, calculate the baby sleep quality index according to the following formula: SQ = [(STT + MTSD * 0.5) - (TAT * 0.5 + TAS / 15)] * α p ; Wherein, SQ is the baby sleep quality index, STT is the total sleep duration, MTSD is the maximum continuous sleep duration, TAT is the total wake duration, TAS is the total number of wake times, and α p is a preset constant associated with the age range of infants and toddlers.