Sleep awakening state identification method and system

The video stream of premature babies is obtained through non-contact vision sensors, behavioral and face biometrics are extracted, and state judgment is used using multimodal decision model, which solves the risk problems of low accuracy and contact sensors in traditional methods, and achieves higher recognition accuracy and reliability.

CN120203524AInactive Publication Date: 2025-06-27PEKING UNIV FIRST HOSPITAL NINGXIA WOMENS & CHILDRENS HOSPITAL (NINGXIA HUI AUTONOMOUS REGION MATERNAL & CHILD HEALTH HOSPITAL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476833.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has a problem of low accuracy in recognition of sleep awakening status in premature infants, especially because traditional methods rely on contact sensors, which poses a risk of skin irritation and infection, and it is difficult for a single modal data to fully capture the correlation between physiological and behavioral characteristics.

Method used

The whole-body motion video stream and facial area video stream of premature babies were synchronized by using non-contact vision sensors, and the state judgment was made by extracting behavioral biometric features (such as respiratory interval variation coefficient and motion energy histogram) and face biometric features (such as eyelid tremor frequency and facial blood oxygen change rate), and using a multimodal decision model.

Benefits of technology

It improves the accuracy and reliability of sleep awakening status recognition in premature infants, avoids the stimulation problem of contact sensors, and enhances the flexibility and judgment accuracy of the system through multimodal fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120203524A_ABST
    Figure CN120203524A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of feature recognition, and discloses a sleep awakening state recognition method, which comprises the following steps: synchronously acquiring a whole-body movement video stream and a face region video stream of a premature infant through a non-contact visual sensor; behavior biological characteristics of the premature infant are extracted from the whole body motion video stream, and the behavior biological characteristics comprise a breathing interval variable coefficient and a motion energy histogram; face biological characteristics of the premature infant are extracted from the face area video stream, and the face biological characteristics comprise the eyelid tremor frequency and the face blood oxygen change rate; and based on the behavior biological characteristics and the face biological characteristics, performing state judgment on the premature infant through a multi-modal decision model. The invention further provides a sleep awakening state recognition system. According to the invention, the accuracy of recognizing the sleep awakening state of the premature infant can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of feature recognition, and particularly to a method and system for recognizing sleep-wake states. Background Art

[0002] In the field of premature infant care, the accurate recognition of sleep-wake states is crucial for clinical care. Traditional methods mainly rely on contact sensors (such as electrode patches, piezoelectric sensors) to collect physiological signals (such as respiration, heart rate). Such devices need to directly contact the skin, which easily leads to the risk of skin irritation or infection for premature infants, and long-term wearing may interfere with their natural sleep behavior.

[0003] In addition, existing technologies mostly make state judgments based on single-modal data (such as only respiratory rate or limb movement amplitude), and have the following limitations: a single sensor cannot comprehensively capture the correlation between physiological and behavioral characteristics (such as respiratory disorders being asynchronous with facial micro-expressions), resulting in a high misjudgment rate; traditional models use static weight allocation and cannot dynamically adjust the judgment logic according to real-time physiological events (such as sudden drops in blood oxygen), affecting the timeliness of early warnings. Therefore, there is an urgent need for a non-contact, multi-modal fusion monitoring solution to improve the accuracy and reliability of sleep-wake state recognition for premature infants. Summary of the Invention

[0004] The present invention provides a method and system for recognizing sleep-wake states, and its main purpose is to solve the problem of low accuracy in recognizing sleep-wake states of premature infants.

[0005] To achieve the above object, a method for recognizing sleep-wake states provided by the present invention includes:

[0006] Simultaneously obtaining a full-body motion video stream and a facial region video stream of a premature infant through a non-contact visual sensor;

[0007] Extracting the behavioral biometric features of the premature infant from the full-body motion video stream, wherein the behavioral biometric features include: coefficient of variation of respiratory intervals and motion energy histogram;

[0008] Extracting the facial biometric features of the premature infant from the facial region video stream, wherein the facial biometric features include: eyelid tremor frequency and facial blood oxygen change rate;

[0009] Based on the behavioral biometric features and the facial biometric features, making a state judgment on the premature infant through a multi-modal decision model, and determining that the premature infant is in a wake state when the following conditions are met:

[0010] The entropy value of the motion energy histogram exceeds a preset entropy threshold;

[0011] The correlation coefficient between the eyelid tremor frequency and the facial blood oxygen change rate is less than a preset correlation coefficient threshold;

[0012] The growth rate of the coefficient of variation of the respiratory interval exceeds a preset growth threshold.

[0013] Optionally, the non-contact visual sensor includes: a near-infrared camera and an RGB camera with a narrow-band filter.

[0014] Optionally, extracting the motion energy histogram of the premature infant from the whole-body motion video stream includes:

[0015] Determining the limb motion vector field of the premature infant based on the whole-body motion video stream;

[0016] Statistically analyzing the motion speed of each pixel point in the limb motion vector field, dividing it according to the speed interval, and counting the number of pixels in each interval, so as to construct a motion energy histogram.

[0017] Optionally, the calculation formula of the coefficient of variation of the respiratory interval is as follows:

[0018]

[0019] Wherein, RIV is the coefficient of variation of the respiratory interval, RR j is the duration of consecutive respiratory cycles, σ(RR j ) is the standard deviation of the duration of consecutive respiratory cycles RR j and μ(RR j ) is the mean value of the duration of consecutive respiratory cycles RR j .

[0020] Optionally, extracting the eyelid tremor frequency of the premature infant from the facial area video stream includes:

[0021] Enhancing the micro-expression signal in the facial area video stream through a phase amplification algorithm;

[0022] Performing time-domain analysis on the enhanced micro-expression signal, and counting the number of eyelid tremors of the premature infant per unit time, so as to obtain the eyelid tremor frequency.

[0023] Optionally, the calculation formula of the facial blood oxygen change rate is as follows:

[0024]

[0025] Wherein, ΔSpO2 is the facial blood oxygen change rate within the time interval Δt, Δt is the set time interval, i is the sampling serial number, N is the number of sampling points, and S i is the blood oxygen sampling value of a continuous 0.5-second window, It is the mean of all blood oxygen sampling values.

[0026] Optionally, the multi-modal decision model adopts a two-stream neural network architecture, including:

[0027] A first processing channel for performing spatio-temporal feature encoding on the motion energy histogram using an Inception-v3 network to obtain high-order behavior features;

[0028] A second processing channel for performing temporal modeling on the facial biometric features using a bidirectional LSTM network to obtain high-order facial features.

[0029] Optionally, the multi-modal decision model adopts a dynamic weight allocation mechanism during the fusion process, where the dynamic weight allocation mechanism is: when the growth rate of the coefficient of variation of the respiratory interval exceeds the growth threshold, the weight of the behavioral biometric feature is increased to 0.7 - 0.9;

[0030] When it is detected that the sudden drop in the facial blood oxygen change rate exceeds 15%, the priority calculation of the facial biometric feature is forcibly triggered, and the state determination is preferably based on the facial biometric feature.

[0031] Optionally, after determining that the premature infant is in a waking state, a prompt message can be sent to the medical staff according to a preset rule.

[0032] To solve the above problems, the present invention also provides a sleep-wake state recognition system, which includes:

[0033] A video stream acquisition module for synchronously acquiring the whole-body motion video stream and the facial area video stream of the premature infant through a non-contact visual sensor;

[0034] A behavioral biometric feature extraction module for extracting the behavioral biometric features of the premature infant from the whole-body motion video stream, where the behavioral biometric features include: the coefficient of variation of the respiratory interval and the motion energy histogram;

[0035] A facial biometric feature extraction module for extracting the facial biometric features of the premature infant from the facial area video stream, where the facial biometric features include: the eyelid tremor frequency and the facial blood oxygen change rate;

[0036] A premature infant state determination module for determining the state of the premature infant through a multi-modal decision model based on the behavioral biometric features and the facial biometric features, and determining that the premature infant is in a waking state when the following conditions are met:

[0037] The entropy value of the motion energy histogram exceeds a preset entropy value threshold;

[0038] The correlation coefficient between the eyelid tremor frequency and the facial blood oxygen change rate is less than a preset correlation coefficient threshold;

[0039] The growth rate of the coefficient of variation of the respiratory interval exceeds a preset growth threshold.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. Dynamically adjust the weights of behaviors and facial biometrics according to the real-time data of the coefficient of variation of the respiratory interval and the facial blood oxygen change rate, avoid misjudgment of a single feature, and enhance the flexibility and determination accuracy of the system;

[0042] 2. Adopt a two-stream neural network architecture to process the motion energy histogram and model the temporal features, so as to extract the spatio-temporal features and the physiological temporal dependence relationship, and solve the problem that the traditional model ignores spatio-temporal heterogeneity;

[0043] 3. Analyze the blood oxygen change through the facial video stream, replace the traditional fingertip contact oximeter, reduce interference and achieve accurate capture of local microcirculation changes;

[0044] 4. Capture the whole-body motion video stream through a near-infrared camera, use the optical flow method to extract the limb motion vector field, construct a motion energy histogram, non-contact quantify the motion intensity and distribution, solve the problem of skin irritation caused by traditional contact sensors, and at the same time determine the motion complexity through the entropy threshold, significantly improving the characterization ability of motion features. Brief Description of the Drawings

[0045] Figure 1 It is a schematic flowchart of a sleep-wake state recognition method provided by an embodiment of the present invention;

[0046] Figure 2 It is a functional module diagram of a sleep-wake state recognition system provided by an embodiment of the present invention;

[0047] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0048] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] An embodiment of the present application provides a method for identifying sleep-wake states. The execution subject of the sleep-wake state identification method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the sleep-wake state identification method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0050] Referring to Figure 1 As shown, it is a schematic flowchart of the sleep-wake state identification method provided by an embodiment of the present invention. In this embodiment, the sleep-wake state identification method includes:

[0051] S1. Synchronously obtain the full-body motion video stream and the facial area video stream of a premature infant through a non-contact visual sensor.

[0052] In the embodiment of the present invention, the non-contact visual sensor includes: a near-infrared camera and an RGB camera with a narrow-band filter.

[0053] Specifically, a near-infrared camera (NIR) and an RGB camera with a narrow-band filter are used to solve the two major pain points of the adaptability to low-light environments and the facial blood oxygen specificity in premature infant care. Among them, the near-infrared camera (NIR) has strong penetrability and can work in low-light or no visible light environments, capturing weak limb movements (such as finger tremors, chest ups and downs), avoiding the interference of traditional infrared supplementary light on the sleep of premature infants, and supporting nighttime monitoring at the same time; the filter in the RGB camera with a narrow-band filter filters ambient light (such as retaining the 520-570nm band), enhancing the absorption characteristics of hemoglobin and accurately extracting facial blood oxygen (SpO2). For example, in an incubator, the narrow-band filter can eliminate the interference of the incubator light and only retain the blood oxygen-sensitive band, improving the signal-to-noise ratio of the blood oxygen signal, while traditional RGB cameras are easily polluted by ambient light.

[0054] Specifically, the full-body motion video stream is obtained by monitoring limb movements (such as kicking legs, waving arms) and chest ups and downs caused by breathing; the facial area video stream (RGB) is used to capture eyelid tremors, facial micro-expressions and blood oxygen changes. Among them, the dual-view data is spatially and temporally aligned, laying a foundation for the correlation analysis of behavior and physiological characteristics.

[0055] Example: In an incubator, the NIR camera captures the periodic leg kicking (sleep cycle) of a premature infant, and the RGB camera synchronously records the slight eyelid tremors (signs of awakening).

[0056] Generally speaking, non-contact data acquisition avoids the irritation of the skin of premature infants caused by contact sensors (such as electrode patches), and the dual-view acquisition based on the whole body and face provides the possibility for cross-modal association of behavioral and physiological characteristics, while traditional single cameras cannot simultaneously take into account both whole body movements and facial micro-expressions.

[0057] S2. Extract the behavioral biometric features of the premature infant from the whole body movement video stream, where the behavioral biometric features include: coefficient of variation of respiratory interval and histogram of motion energy.

[0058] In the embodiment of the present invention, the coefficient of variation of respiratory interval (RIV) represents capturing the chest undulation through the NIR camera to generate a respiratory waveform.

[0059] Specifically, a peak detection algorithm is used to identify the peaks (inhalation) and valleys (exhalation) of the chest undulation, calculate the time difference between adjacent peaks, and obtain a sequence of continuous respiratory cycle durations.

[0060] Specifically, when the premature infant is in the waking state, the excitement of the sympathetic nerve causes the respiratory rate to increase and become irregular, and the RIV value increases significantly (such as from 15% to 30%); when the premature infant is in the deep sleep state, the breathing is stable, and the RIV value is lower than the threshold (such as 10%).

[0061] In the embodiment of the present invention, extracting the histogram of motion energy of the premature infant from the whole body movement video stream includes:

[0062] Determining the limb motion vector field of the premature infant based on the whole body movement video stream;

[0063] Statistically analyze the motion speed of each pixel point in the limb motion vector field, divide and count the number of pixels in each interval according to the speed interval, so as to construct a histogram of motion energy.

[0064] Specifically, analyze the pixel displacement between video frames by the optical flow method, calculate the motion vector (direction and speed) of each pixel between adjacent video frames, generate a limb motion vector field, count the number of pixels in each speed interval (such as 0 - 10, 10 - 20, 20 - 30 pixels / frame), and form a histogram. For example: if the leg of the premature infant moves 10 pixels in two consecutive frames, the motion speed of the corresponding pixel is 10 pixels / frame.

[0065] Specifically, when awake, the motion intensity is high and distributed dispersedly (such as rapid kicking), and the entropy value of the histogram increases; when sleeping, the motion is concentrated in the low-speed interval (such as 0 - 10 pixels / frame).

[0066] Example: During the waking state, the number of pixels in the high-speed range (such as 30-50 pixels / frame) of the motion energy histogram increases significantly, and the entropy value exceeds the preset threshold.

[0067] In the embodiment of the present invention, the calculation formula of the respiratory interval coefficient of variation is as follows:

[0068]

[0069] Wherein, RIV is the respiratory interval coefficient of variation, and RR j is the duration of consecutive respiratory cycles (analyzing the chest undulation interval through NIR video), and σ(RR j ) is the standard deviation of the duration of consecutive respiratory cycles RR j , reflecting the volatility of the respiratory rhythm, and μ(RR j ) is the mean of the duration of consecutive respiratory cycles RR j , which is used to normalize the degree of variation.

[0070] Specifically, premature infants have underdeveloped central nervous systems. When awake, their respiratory frequencies increase and are irregular. Therefore, σ(RR j ) will increase, while during deep sleep, the breathing is stable ( low ratio), and the RIV value increases significantly.

[0071] Specifically, compared with a single respiratory frequency (such as 30 breaths per minute), RIV introduces the "coefficient of variation" to quantify the dynamic fluctuations of breathing (for example: RR j changes from 3s → 2s → 4s, σ(RR j ) = 0.816, μ(RR j ) = 3, and RIV = 27.2%, indicating a waking trend.

[0072] Generally speaking, by observing the chest undulation in the whole-body video stream (captured by the NIR camera), the respiratory cycle is calculated non-contact, avoiding the shift error of traditional piezoelectric sensors.

[0073] S3. Extract the face biometrics of the premature infant from the video stream of the facial area, where the face biometrics include: the frequency of eyelid flutter and the rate of facial blood oxygen change.

[0074] In the embodiment of the present invention, extracting the frequency of eyelid flutter of the premature infant from the video stream of the facial area includes:

[0075] Enhancing the micro-expression signal in the video stream of the facial area through the phase amplification algorithm;

[0076] Performing time-domain analysis on the enhanced micro-expression signal, and counting the number of eyelid flutters of the premature infant per unit time to obtain the frequency of eyelid flutter.

[0077] Specifically, the phase amplification algorithm performs spatio-temporal filtering on video frames to amplify micro-expression signals in specific frequency bands (such as eyelid fluttering), which is used to enhance micro-expression signals in facial videos and highlight the tiny movements of the eyelids. The phase amplification algorithm can achieve non-contact high-frequency monitoring (such as sampling 10 times per second); time-domain analysis refers to counting the number of eyelid closures per unit time (such as per minute).

[0078] Specifically, the phase amplification algorithm includes: decomposing the video into sub-bands of different frequencies and selectively amplifying the frequency band related to eyelid fluttering (such as 1 - 5 Hz).

[0079] Specifically, time-domain analysis includes: in the enhanced video, locating the eyelid contour through an edge detection algorithm and counting the number of eyelid closures per unit time (such as per minute).

[0080] For example: The frequency of eyelid fluttering may increase from 1 time per minute during sleep to 5 times per minute during wakefulness.

[0081] Generally speaking, traditional electrooculogram (EOG) requires electrodes to be attached, and the sampling rate is limited (usually ≤100 Hz), while phase amplification can achieve an equivalent sampling rate of 1000 Hz.

[0082] In the embodiments of the present invention, the calculation formula of the facial blood oxygen change rate is as follows:

[0083]

[0084] Wherein, ΔSpO2 is the facial blood oxygen change rate within the time interval Δt, which reflects the fluctuation rate of blood oxygen concentration, Δt is the set time interval, i is the sampling serial number, N is the number of sampling points, and S i is the blood oxygen sampling value of a continuous 0.5 - second window (analyzing the facial blood flow change through an RGB camera, (such as sampling once every 0.1 second, N = 5)), is the average value of all blood oxygen sampling values.

[0085] First, capture the facial blood flow change through an RGB camera and calculate the blood oxygen saturation using the Lambert - Beer law; secondly, generate the facial blood oxygen change rate based on the blood oxygen saturation.

[0086] Specifically, in the waking state, the activation of the sympathetic nerve causes peripheral vasoconstriction, and the blood oxygen may suddenly drop (such as a 10% drop within 10 seconds)

[0087] Specifically, when premature infants wake up and cry, and their limb activities increase, it causes fluctuations in facial blood flow while the blood oxygen is stable during sleep

[0088] Example: When a premature infant suddenly wakes up, the blood oxygen level changes from 98% → 95% → 97% (within 3 seconds), calculate exceeding the threshold by 0.5, triggering the face feature priority.

[0089] Generally speaking, different from fingertip blood oxygen (contact type), facial blood oxygen is obtained non - contact through video stream, reflecting local microcirculation changes (such as facial congestion during crying), and forming a "physiological - behavior" closed - loop verification with whole - body movement.

[0090] S4. Based on the behavioral biometric feature and the face biometric feature, use a multi - modal decision model to determine the state of the premature infant.

[0091] In the embodiment of the present invention, the multi - modal decision model adopts a two - stream neural network architecture, including:

[0092] The first processing channel is used to perform spatio - temporal feature encoding on the motion energy histogram by using the Inception - v3 network to obtain high - order behavioral features;

[0093] The second processing channel is used to perform temporal sequence modeling on the face biometric feature by using a bidirectional LSTM network to obtain high - order face features.

[0094] Specifically, the first processing channel is used to extract spatio - temporal features of the motion energy histogram (such as periodic motion patterns, action intensity distributions). Among them, the multi - scale convolutional kernels (1x1, 3x3, 5x5) of Inception - v3 process the motion energy histogram, which can capture motion features in different speed intervals. Inception - v3 (spatial dimension) captures motion patterns (such as "limbs moving together" vs "unilateral limb movement").

[0095] Specifically, the second processing channel is used to model the temporal dependence relationship between eyelid flutter and blood oxygen change (such as blood oxygen decline lagging behind eyelid flutter); the bidirectional LSTM considers both past and future data simultaneously, enhancing the ability to analyze temporal sequence signals (such as: the flutter frequency increases → the blood oxygen fluctuation lags by 2 seconds, which conforms to the physiological logic of awakening).

[0096] Generally speaking, traditional fusion models (such as early splicing) ignore the spatio - temporal heterogeneity of features. The two - stream architecture respectively retains the spatial distribution of behaviors (histogram entropy value) and the temporal dynamics of the face (changing trend of flutter frequency).

[0097] In the embodiment of the present invention, the multi - modal decision model adopts a dynamic weight allocation mechanism during the fusion process. Among them, the dynamic weight allocation mechanism is: when the growth rate of the coefficient of variation of the respiratory interval exceeds the growth threshold, the weight of the behavioral biometric feature is increased to 0.7 - 0.9;

[0098] When it is detected that the sudden drop in the facial blood oxygen change rate exceeds 15%, the priority calculation of the face biometric features is forcibly triggered, and the status determination is preferentially based on the face biometric features.

[0099] Specifically, at 2:00 am, the RIV changed from 12% to 22% (within 1 minute). The model determined that the main cause was unstable breathing and gave priority to trusting the motion energy histogram (for example, the high entropy value corresponding to the kicking-off-the-quilt action).

[0100] Specifically, the trigger condition for the weight increase of the behavioral features from 0.7 to 0.9 is that the growth rate of the respiratory interval variation coefficient (RIV) exceeds the threshold (such as 20%). The sudden increase in RIV indicates respiratory disorder. At this time, the motion and respiratory features are emphasized to avoid misjudgment. For example, if the RIV suddenly increases from 15% to 35%, the system will set the weight of the behavioral features to 0.8 and weaken the influence of the facial features.

[0101] Specifically, the trigger condition for the forced trigger of the face features is that the facial blood oxygen suddenly drops by more than 15%. Because the sudden drop in blood oxygen may endanger life, the face features are forcibly preferred for determination to ensure timely response. For example, when the blood oxygen suddenly drops from 95% to 80%, even if the motion energy does not exceed the threshold, the system still determines it as awakening and issues an alarm.

[0102] In the embodiment of the present invention, when the following conditions are met, it is determined that the premature infant is in the awakening state:

[0103] The entropy value of the motion energy histogram exceeds the preset entropy value threshold, reflecting an increase in motion complexity;

[0104] The correlation coefficient between the eyelid flutter frequency and the facial blood oxygen change rate is less than the preset correlation coefficient threshold, indicating that the changes of the two are asynchronous during awakening (such as frequent eyelid fluttering but stable blood oxygen);

[0105] The growth rate of the respiratory interval variation coefficient exceeds the preset growth threshold, indicating a significant disorder in the respiratory rhythm.

[0106] For example: when it is detected that the entropy value of the motion energy suddenly increases (awakening period activity), the eyelid flutter frequency is 5 times per minute (correlation coefficient < preset correlation coefficient threshold 0.3), and the RIV increases by 25% (> preset growth threshold 20%), it is comprehensively determined that the infant is in the awakening state, and the medical staff is notified to adjust the nursing plan.

[0107] For example: a premature infant had an increase in RIV (18% → 25%) at 3:15, the histogram entropy value was 1.3 (> preset entropy value threshold 1.2), but the eyelid flutter - blood oxygen correlation coefficient was 0.6 (> preset correlation coefficient threshold 0.4). The model determined it as light sleep rather than awakening to avoid false alarms.

[0108] Specifically, a high entropy value indicates a complex motion pattern, and rapid and irregular movements are common during the awakening period;

[0109] Specifically, the Pearson correlation coefficient between the eyelid tremor frequency and the blood oxygen change rate is calculated. During awakening, the changes of the two are asynchronous (such as frequent eyelid tremors but stable blood oxygen), and the correlation coefficient approaches 0.

[0110] Generally speaking, three conditions need to be met simultaneously to avoid misjudgment by a single feature.

[0111] In an embodiment of the present invention, after determining that the premature infant is in the awakening state, a prompt message can be sent to medical staff according to a preset rule.

[0112] The present invention dynamically adjusts the weights of behavior and face biometrics according to the real-time data of the respiratory interval coefficient of variation and the blood oxygen change rate of the face, avoiding misjudgment by a single feature, and enhancing the flexibility and determination accuracy of the system; adopts a two-stream neural network architecture to process the motion energy histogram and model the temporal sequence features, so as to extract the spatio-temporal features and the physiological temporal sequence dependence relationship, and solves the problem that the traditional model ignores spatio-temporal heterogeneity; analyzes the blood oxygen change through the face video stream, replaces the traditional fingertip contact blood oxygen meter, reduces interference and realizes the accurate capture of local microcirculation changes; captures the whole-body motion video stream through a near-infrared camera, and uses the optical flow method to extract the limb motion vector field, constructs the motion energy histogram, and non-contact quantifies the motion intensity and distribution, solving the problem of skin irritation caused by traditional contact sensors. At the same time, the motion complexity is determined by the entropy value threshold, significantly improving the representation ability of motion features. Therefore, the accuracy of sleep-wake state recognition of premature infants is improved.

[0113] As Figure 2 shown, it is a functional module diagram of a sleep-wake state recognition system provided by an embodiment of the present invention.

[0114] The sleep-wake state recognition system 100 of the present invention can be installed in an electronic device. According to the functions achieved, the sleep-wake state recognition system 100 can include a video stream acquisition module 101, a behavior biometric extraction module 102, a face biometric extraction module 103, and a premature infant state determination module 104. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0115] In this embodiment, the functions of each module / unit are as follows:

[0116] The video stream acquisition module 101 is used to synchronously acquire the whole-body motion video stream and the facial area video stream of a premature infant through a non-contact visual sensor;

[0117] The behavioral biometric extraction module 102 is configured to extract the behavioral biometrics of the premature infant from the full-body motion video stream, where the behavioral biometrics include: the coefficient of variation of respiratory intervals and the motion energy histogram;

[0118] The facial biometric extraction module 103 is configured to extract the facial biometrics of the premature infant from the facial area video stream, where the facial biometrics include: the eyelid tremor frequency and the facial blood oxygen change rate;

[0119] The premature infant state determination module 104 is configured to determine the state of the premature infant based on the behavioral biometrics and the facial biometrics through a multi-modal decision model. When the following conditions are met, it is determined that the premature infant is in the waking state:

[0120] The entropy value of the motion energy histogram exceeds a preset entropy threshold;

[0121] The correlation coefficient between the eyelid tremor frequency and the facial blood oxygen change rate is less than a preset correlation coefficient threshold;

[0122] The growth rate of the coefficient of variation of respiratory intervals exceeds a preset growth threshold.

[0123] In several embodiments provided by the present invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0124] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0125] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0126] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0127] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for identifying sleep-wake states, characterized in that: The method comprises: The whole body motion video stream and the facial area video stream of premature infants are synchronously acquired through non-contact visual sensors; Extracting behavioral biometric features of the premature infant from the whole-body motion video stream, wherein the behavioral biometric features include: coefficient of variation of breathing intervals and motion energy histogram; Extracting facial biometric features of the premature infant from the facial region video stream, wherein the facial biometric features include: eyelid tremor frequency and facial blood oxygen change rate; Based on the behavioral biometrics and the facial biometrics, the state of the premature infant is determined by a multimodal decision model, and the premature infant is determined to be in an awake state when the following conditions are met: The entropy value of the motion energy histogram exceeds a preset entropy value threshold; The correlation coefficient between the eyelid tremor frequency and the facial blood oxygen change rate is less than a preset correlation coefficient threshold; The growth rate of the respiratory interval variation coefficient exceeds a preset growth threshold.

2. The sleep-wake state recognition method according to claim 1, characterized in that: The non-contact visual sensor includes a near-infrared camera and an RGB camera with a narrow-band filter.

3. The sleep-wake state recognition method according to claim 1, characterized in that: Extracting a motion energy histogram of the premature infant from the whole-body motion video stream comprises: Determine a limb motion vector field of the premature infant based on the whole body motion video stream; The movement speed of each pixel in the limb movement vector field is counted, and the field is divided into speed intervals and the number of pixels in each interval is counted, so as to construct a movement energy histogram.

4. The sleep-wake state recognition method according to claim 1, characterized in that: The calculation formula of the respiratory interval variation coefficient is as follows: Where RIV is the coefficient of variation of the respiratory interval, RR j is the duration of the continuous breathing cycle, σ(RR j ) is the duration of the continuous breathing cycle RR j The standard deviation, μ(RR j ) is the duration of the continuous breathing cycle RR j The mean of .

5. The sleep-wake state recognition method according to claim 1, characterized in that: Extracting the eyelid twitching frequency of the premature infant from the facial region video stream comprises: The micro-expression signal in the facial region video stream is enhanced by a phase amplification algorithm; The enhanced micro-expression signal is subjected to time domain analysis, and the number of eyelid tremors of the premature infant per unit time is counted to obtain the eyelid tremor frequency.

6. The sleep-wake state recognition method according to claim 1, characterized in that: The calculation formula of the facial blood oxygen change rate is as follows: Wherein, ΔSpO2 is the facial blood oxygen change rate within the time interval Δt, Δt is the set time interval, i is the sampling sequence number, N is the number of sampling points, S i is the blood oxygen sampling value of a continuous 0.5 second window, It is the mean of all blood oxygen sampling values.

7. The sleep-wake state recognition method according to claim 1, characterized in that: The multimodal decision model adopts a two-stream neural network architecture, including: A first processing channel is used to encode the spatiotemporal features of the motion energy histogram using an Inception-v3 network to obtain high-order behavioral features; The second processing channel is used to use a bidirectional LSTM network to perform time series modeling on the facial biometric features to obtain high-order facial features.

8. The sleep-wake state recognition method according to claim 1, characterized in that: The multimodal decision model adopts a dynamic weight allocation mechanism in the fusion process, wherein the dynamic weight allocation mechanism is: when the growth rate of the coefficient of variation of the breathing interval exceeds the growth threshold, the weight of the behavioral biological feature is increased to 0.7-0.9; When it is detected that the facial blood oxygen change rate drops suddenly by more than 15%, the priority calculation of the facial biometric feature is forcibly triggered, and the status is determined based on the facial biometric feature first.

9. The sleep-wake state recognition method according to any one of claims 1 to 8, characterized in that: After determining that the premature infant is in an awake state, a prompt message can be sent to medical staff according to preset rules.

10. A sleep-wake state recognition system, characterized in that: The system comprises: A video stream acquisition module, used for synchronously acquiring a whole body motion video stream and a facial region video stream of a premature infant through a non-contact visual sensor; A behavioral biometric feature extraction module, used to extract the behavioral biometric features of the premature infant from the whole-body motion video stream, wherein the behavioral biometric features include: a coefficient of variation of breathing intervals and a motion energy histogram; A facial biometric feature extraction module, used to extract facial biometric features of the premature infant from the facial region video stream, wherein the facial biometric features include: eyelid tremor frequency and facial blood oxygen change rate; The premature infant state determination module is used to determine the state of the premature infant through a multimodal decision model based on the behavioral biological features and the facial biological features, and determine that the premature infant is in an awake state when the following conditions are met: The entropy value of the motion energy histogram exceeds a preset entropy value threshold; The correlation coefficient between the eyelid tremor frequency and the facial blood oxygen change rate is less than a preset correlation coefficient threshold; The growth rate of the respiratory interval variation coefficient exceeds a preset growth threshold.