Visual determination of sleep state

JP2024530535A5Active Publication Date: 2025-07-01JACKSON LAB THE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023580368
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-27
Filing Date
2022-06-27
Publication Date
2025-07-01
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

Existing methods for sleep state analysis in rodents, such as EEG/EMG recordings, are invasive, require surgery, and have low throughput, while non-invasive methods like beam break systems and imaging techniques cannot distinguish between wakeful, REM, and NREM states accurately.

Method used

A computer-implemented method using machine learning models to process video data, segmenting objects from background, and analyzing features like kurtosis, mean, and power spectral density to determine sleep states, including REM and NREM, without invasive electrodes.

Benefits of technology

Enables accurate, non-invasive, and high-throughput sleep state determination in rodents, facilitating large-scale mechanistic studies for therapeutic discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The systems and methods described herein provide techniques for determining sleep state data by processing video data of a subject. The systems and methods may determine a plurality of features from the video data and may use the plurality of features to determine the sleep state data of the subject. In some embodiments, the sleep state data may be based on frequency domain features and / or time domain features corresponding to the plurality of features.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (Related Applications) This application claims the benefit under 35 U.S.C. Section 119(e) of U.S. Provisional Application No. 63 / 215,511, filed June 27, 2021, the disclosure of which is incorporated herein by reference in its entirety.

[0002] The invention relates in some aspects to determining a subject's sleep state by processing video data using machine learning models.

[0003] (Government support) This invention was made with government support under grants DA041668 (NIDA), DA048634 (NIDA), and HL094307 (NHLBI) awarded by the National Institutes of Health. The Government has certain rights in this invention. [Background technology]

[0004] Sleep is a complex behavior regulated by homeostatic processes and whose function is critical for survival. Sleep and circadian disorders are found in many diseases, including neuropsychiatric, neurodevelopmental, neurodegenerative, physiological, and metabolic disorders. Sleep and circadian functions have a bidirectional relationship with these diseases, and alterations in sleep and circadian patterns can lead to or be the cause of disease states. Although the bidirectional relationship between sleep and many diseases has been well described, their genetic etiology has not been fully elucidated. Indeed, treatments for sleep disorders are limited due to a lack of knowledge about sleep mechanisms. Rodents serve as a readily available model of human sleep due to similarities in sleep biology, and mice in particular are a genetically traceable model for mechanistic studies of sleep and potential therapeutics. One reason for this significant gap in treatment is due to technical barriers that prevent reliable phenotyping of large numbers of mice for assessment of sleep status. The gold standard for sleep analysis in rodents utilizes electroencephalogram / electromyogram (EEG / EMG) recordings. This method requires surgery for electrode implantation and often manual scoring of recordings, resulting in low throughput. New methods utilizing machine learning models are beginning to automate EEG / EMG scoring, but data generation remains low throughput. In addition, the use of tethered electrodes restricts the animal's movement and may alter the animal's behavior.

[0005] To overcome the low throughput limitations of some existing systems, several non-invasive approaches for sleep analysis have been investigated. These include activity assessment by beam-breaking systems, or videography, where a certain amount of inactivity is interpreted as sleep. Piezo pressure sensors have also been used as a simpler and more sensitive method of accessing activity. However, these methods only assess wakefulness relative to sleep states and cannot distinguish between wakefulness, rapid eye movement (REM), and non-REM states. This is important because activity determination of sleep states can be inaccurate in humans as well as rodents, which have low general activity. Other methods of assessing sleep states include pulsed Doppler-based methods that access movement and respiration, as well as whole-body plethysmography, which directly measures breathing patterns. Both of these approaches require specialized equipment. Electric field sensors, which detect respiration and other movements, have also been used to assess sleep states. Summary of the Invention

[0006] According to an embodiment of the present invention, a computer-implemented method is provided, the method including receiving video data representing a video of an object, using the video data to determine a plurality of features corresponding to the object, and using the plurality of features to determine sleep state data for the object. In some embodiments, the method also includes processing the video data using a machine learning model to determine segmentation data indicative of a first set of pixels corresponding to the object and a second set of pixels corresponding to a background. In some embodiments, the method also includes processing the segmentation data to determine ellipse fit data corresponding to the object. In some embodiments, determining the plurality of features includes processing the segmentation data to determine the plurality of features. In some embodiments, the plurality of features includes a plurality of visual features for each video frame of the video data. In some embodiments, the method also includes determining a time domain feature for each visual feature of the plurality of visual features, the plurality of features including a time domain feature. In some embodiments, determining the time domain feature includes determining one of kurtosis data, mean data, median data, standard deviation data, maximum data, and minimum data. In some embodiments, the method also includes determining a frequency domain feature for each visual feature of the plurality of visual features, the plurality of features including a frequency domain feature. In some embodiments, determining the frequency domain features includes determining one of a kurtosis of the power spectral density, a skewness of the power spectral density, a mean power spectral density, a sum power spectral density, a maximum data, a minimum data, a mean data, and a standard deviation of the power spectral density. In some embodiments, the method also includes determining a time domain feature for each of the plurality of features, determining a frequency domain feature for each of the plurality of features, and processing the time domain features and the frequency domain features using a machine learning classifier to determine sleep state data. In some embodiments, the method also includes processing the plurality of features using a machine learning classifier to determine a sleep state for the video frames of the video data, the sleep state being one of a wakefulness state, a REM sleep state, and a non-REM sleep state.In some embodiments, the sleep state data indicates one or more of a duration of a sleep state, a duration and / or a frequency interval of one or more of a wakefulness state, a REM state, and a non-REM state, and a change in one or more sleep states. In some embodiments, the method also includes determining a plurality of body regions of the subject using the plurality of features, where each body region of the plurality of body regions corresponds to a video frame of the video data, and determining the sleep state data based on a change in the plurality of body regions in the video. In some embodiments, the method also includes determining a plurality of width-length ratios using the plurality of features, where each width-length ratio of the plurality of width-length ratios corresponds to a video frame of the video data, and determining the sleep state data based on a change in the plurality of width-length ratios in the video. In some embodiments, determining the sleep state data includes detecting a transition from a non-REM state to a REM state based on a change in a body region or shape of the subject, where the change in the body region or shape is a result of muscle weakness. In some embodiments, the method also includes determining a plurality of width-length ratios for the subject, the width-length ratios of the plurality of width-length ratios corresponding to a video frame of the video data; determining a time domain feature using the plurality of width-length ratios; determining a frequency domain feature using the plurality of width-length ratios, the time domain feature and the frequency domain feature representing abdominal movement of the subject; and determining sleep state data using the time domain feature and the frequency domain feature. In some embodiments, the video captures the subject in its natural state. In some embodiments, the natural state of the subject includes no invasive detection means in or on the subject. In some embodiments, the invasive detection means includes one or both of an electrode attached to the subject and an electrode inserted into the subject. In some embodiments, the video is a high-resolution video. In some embodiments, the method also includes: processing the plurality of features using a machine learning classifier to determine a plurality of sleep state predictions for each video frame of the video data; and processing the plurality of sleep state predictions using a transition model to determine a transition between a first sleep state and a second sleep state.In some embodiments, the transition model is a hidden Markov model. In some embodiments, the subject is a rodent, optionally a mouse. In some embodiments, the subject is a genetically engineered subject.

[0007] According to another aspect of the present invention, there is provided a method of determining a sleep state of a subject, the method comprising monitoring a response of the subject, the monitoring means comprising any of the embodiments of the computer-implemented methods described above. In some embodiments, the sleep state comprises one or more of a sleep stage, a duration of a sleep interval, a change in a sleep stage, and a duration of a non-sleep interval. In some embodiments, the subject has a sleep disorder or condition. In some embodiments, the sleep disorder or condition comprises one or more of sleep apnea, insomnia, and narcolepsy. In some embodiments, the sleep disorder or condition is a result of brain injury, depression, a psychiatric disorder, a neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, obesity, overweight, the effect of an administered drug, and / or the effect of alcohol intake, a neurological condition that can alter the sleep state, or a metabolic disorder or condition that can alter the sleep state. In some embodiments, the method also comprises administering a therapeutic agent to the subject prior to receiving the video data. In some embodiments, the therapeutic agent comprises one or more of a sleep-promoting agent, a sleep inhibitor, and a drug that can alter one or more sleep stages of the subject. In some embodiments, the method also includes administering a behavioral treatment to the subject. In some embodiments, the behavioral treatment includes a sensory therapy. In some embodiments, the sensory therapy is light exposure therapy. In some embodiments, the subject is a genetically engineered subject. In some embodiments, the subject is a rodent, optionally a mouse. In some embodiments, the mouse is a genetically engineered mouse. In some embodiments, the subject is an animal model of a sleep pathology. In some embodiments, the sleep state data determined for the subject is compared to control sleep state data. In some embodiments, the control sleep state data is sleep state data from a control subject determined by a computer-implemented method. In some embodiments, the control subject does not have the subject's sleep disorder or pathology. In some embodiments, the control subject does not receive the therapeutic agent or behavioral treatment administered to the subject. In some embodiments, the control subject is administered a dose of a therapeutic agent that is different from the dose of a therapeutic agent administered to the subject.

[0008] According to another aspect of the present invention, there is provided a method of identifying the effectiveness of a candidate therapeutic agent and / or behavioral treatment for treating a sleep disorder or condition in a subject, the method comprising administering the candidate therapeutic agent and / or behavioral treatment to a test subject and determining sleep state data for the test subject, the determining means comprising any embodiment of any of the aforementioned computer-implemented methods, wherein the determining indicative of a change in the sleep state data in the test subject identifies the effect of the candidate therapeutic agent or behavioral treatment, respectively, on the subject's sleep disorder or condition. In some embodiments, the sleep state data comprises one or more of data on sleep stages, duration of sleep intervals, changes in sleep stages, and duration of non-sleep intervals. In some embodiments, the test subject has a sleep disorder or condition. In some embodiments, the sleep disorder or condition comprises one or more of sleep apnea, insomnia, and narcolepsy. In some embodiments, the sleep disorder or condition is the result of brain injury, depression, psychiatric disease, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, obesity, overweight, the effect of administered medication, and / or the effect of alcohol intake, a neurological condition that can alter sleep state, or a metabolic disorder or condition that can alter sleep state. In some embodiments, the candidate therapeutic and / or candidate behavioral treatment is administered to the test subject one or more of before or during receipt of the video data. In some embodiments, the candidate therapeutic comprises one or more of a sleep promoter, a sleep inhibitor, and an agent that can alter one or more sleep stages of the test subject. In some embodiments, the behavioral treatment comprises a sensory therapy. In some embodiments, the sensory therapy is light exposure therapy. In some embodiments, the subject is a genetically engineered subject. In some embodiments, the test subject is a rodent, optionally a mouse. In some embodiments, the mouse is a genetically engineered mouse. In some embodiments, the test subject is an animal model of a sleep pathology. In some embodiments, the sleep state data determined for the test subject is compared to control sleep state data. In some embodiments, the control sleep state data is sleep state data from a control subject determined by a computer-implemented method.In some embodiments, the control subject does not have the sleep disorder or condition of the test subject. In some embodiments, the control subject does not receive the candidate therapeutic agent administered to the test subject. In some embodiments, the control subject receives a dose of the candidate therapeutic agent that is different from the dose of the candidate therapeutic agent administered to the test subject. In some embodiments, the control subject receives a candidate behavioral treatment regimen that is different from the candidate therapeutic agent regimen administered to the test subject. In some embodiments, the behavioral treatment regimen includes one or more of the following treatment characteristics: the length of the behavioral treatment, the intensity of the behavioral treatment, the light intensity of the behavioral treatment, and the frequency of the behavioral treatment. [Brief description of the drawings]

[0009] For a more complete understanding of the present disclosure, reference is now made to the following descriptions taken in conjunction with the accompanying drawings, in which:

[0010] [Figure 1] FIG. 1 is a conceptual diagram of a system for determining sleep state data of a subject using video data, according to an embodiment of the present disclosure. [Figure 2A] FIG. 2A is a flow chart illustrating a process for determining sleep state data according to an embodiment of the present disclosure. [Figure 2B] FIG. 2B is a flow chart illustrating a process for determining sleep state data for multiple objects depicted in a video according to an embodiment of the present disclosure. [Diagram 3] FIG. 3 is a conceptual diagram of a system for training components that determine sleep state data, according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a block diagram conceptually illustrating example components of a device in accordance with an embodiment of the present disclosure. [Diagram 5] FIG. 5 is a block diagram conceptually illustrating example components of a server in accordance with an embodiment of the present disclosure. [Figure 6A] FIG. 6A shows a schematic diagram illustrating the organization of data collection, annotation, feature generation, and classifier training according to an embodiment of the present disclosure. [Figure 6B]FIG. 6B shows a schematic diagram of the frame-level information used for visual features, according to an embodiment of the present disclosure, where the trained neural network was used to generate a segmentation mask of mouse-related pixels for use in downstream classification. [Figure 6C] FIG. 6C shows a schematic diagram of multiple frames of a video containing multiple objects, where, according to an embodiment of the present disclosure, an instance segmentation technique is used to generate segmentation masks for individual objects, even when the objects are close to each other. [Figure 7A] FIG. 7A shows example graphs of selected signals in the time and frequency domain within one epoch, showing m00 (region of the segmentation mask) for wake, NREM, and REM states (leftmost column), the FFT of the corresponding signal (middle column), and the autocorrelation of the signal (rightmost column). [Figure 7B] FIG. 7B shows an example graph of a selected signal in the time and frequency domain within one epoch, similar to FIG. 6A, showing the wl_ratio in the time and frequency domains. [Figure 8A] Figures 8A-B show plots illustrating extraction of respiratory signals from video footage. Figure 8A shows exemplary spectral analysis plots for REM and non-REM epochs. Continuous wavelet transform spectral response (top panel) and associated dominant signal (respective bottom left panel), as well as histogram of the dominant signal (respective bottom right panel). Non-REM epochs typically showed lower means and standard deviations than REM epochs. [Figure 8B] Figure 8B shows a plot examining the larger time scale of the epoch showing that the NREM signal was stable up until the onset of REM, with the dominant frequency being the typical mouse respiratory rate frequency. [Figure 9A]Figures 9A-C show graphs illustrating validation of respiration signal in video data for wl_ratio measurements. Figure 9A shows the mobility cutoff used to select for sleep epochs in C57BL / 6J vs. C3H / HeJ respiration rate analysis. Below the 10% quantile cutoff threshold (black vertical line), epochs consisted of 90.2% NREM (red line), 8.1% REM (green line), and 1.7% wake (blue line). [Figure 9B] Figure 9B shows a cross-strain comparison of the dominance frequencies observed during sleep epochs (blue, males; orange, females). [Figure 9C] Figure 9C shows that using the C57BL / 6J annotated epochs, a higher standard deviation in the dominant frequency was observed in the REM condition (blue line) than in the NREM condition (orange line). [Figure 9D] FIG. 9D shows that the increase in standard deviation was consistent across all animals. [Figure 10A] Figures 10A-D show graphs and tables showing the performance metrics of the classifiers. Figure 10A shows the classifier performance compared at different stages, starting with an XgBoost classifier, adding an HMM model, increasing features to include 7 Hu moments, and integrating SINDLE annotations to improve epoch quality. It was observed that by adding each of these steps, the overall accuracy rate improved. [Figure 10B] FIG. 10B shows the top 20 features that are most important to the classifier. [Figure 10C] FIG. 10C shows the confusion matrix obtained from the 10-fold cross-validation. [Figure 10D] FIG. 10D shows the precision-recall table. [Figure 11A] 11A-D show graphs illustrating the validation of visual scoring: Fig. 11A shows a hypnogram of visual scoring and EEG / EMG scoring. [Figure 11B] FIG. 11B shows plots of 24-hour visually scored sleep stages (top) and predicted stages (bottom) for a mouse (B6J_7). [Figure 11C] Figures 11C-D show a comparison of human and visual scoring across C57BL / 6J mice, demonstrating high concordance between the two methods. Data are plotted in 1-hour bins over a 24-hour period (Figure 11C). [Figure 11D] Plots were made over a 24- or 12-h period (FIG. 11D). [Figure 12] FIG. 12 shows a bar graph illustrating the results of additional data augmentation on the classifier model. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] The present disclosure relates to determining a sleep state of a subject by processing video data of the subject using one or more machine learning models. The breathing, movement, or posture of a subject is useful in itself for distinguishing sleep states. In some embodiments of the present disclosure, a combination of breathing, movement, and posture features is used to determine the sleep state of a subject. Using a combination of these features increases the accuracy rate of predicting sleep states. The term "sleep state" is used in reference to rapid eye movement (REM) sleep states and non-rapid eye movement (NREM) sleep states. The method and system of the present invention can be used to evaluate and distinguish between REM sleep states, NREM sleep states, and wakefulness (non-sleep) states of a subject.

[0012] To distinguish between wakefulness, non-REM, and REM states, in some embodiments, a video-based method using high-resolution video is used based on determining that information about sleep states is encoded in the video data. When moving from non-REM to REM states, slight changes in the area and shape of the object are observed, likely due to the attenuation of the REM state. Over the past few years, significant improvements have been made in the field of computer vision, mainly due to advances in the fields of machine learning, particularly deep learning. Some embodiments use advanced machine vision methods to significantly improve visual sleep state classification. Some embodiments involve extracting features from the video data related to the subject's breathing, movement, and / or posture. Some embodiments combine these features to determine sleep states in a subject, such as a mouse. Embodiments of the present disclosure involve non-invasive video-based methods that can be implemented with low hardware investment and result in high quality sleep state data. The ability to access sleep states reliably, non-invasively, and in a high-throughput manner enables the large-scale mechanistic studies necessary for therapeutic discovery.

[0013] FIG. 1 conceptually illustrates a system 100 (e.g., automatic sleep state system 100) for determining sleep state data of a subject using video data. The automatic sleep state system 100 may operate using various components shown in FIG. 1. The automatic sleep state system 100 may include an image capture device 101, a device 102, and one or more systems 105 connected via one or more networks 199. The image capture device 101 may be part of, included within, or connected to another device (e.g., device 400 shown in FIG. 4), and may further be a camera, or a high-speed video camera, or other type of device that may capture images or video. The device 101 may include motion detection sensors, infrared sensors, temperature sensors, ambient condition detection sensors, and other sensors configured to detect various characteristics / environmental conditions in addition to or instead of the image capture device. Device 102 may be a laptop, desktop, tablet, smartphone, or other type of computing device capable of displaying data, and may include one or more of the components described in connection with device 400 below.

[0014] The image capture device 101 may capture a video (or one or more images) of an object and may transmit video data 104 representing the video to the system 105 for processing as described herein. The video may be of an object in an open field arena. In some cases, the video data 104 may correspond to images (image data) captured by the device 101 at a particular time interval, such that the images capture the object over a period of time. In some embodiments, the video data 104 may be a high-resolution video of the object.

[0015] 1 and may be configured to process the video data 104 to determine sleep state data for the subject. The system 105 may generate sleep state data 152 corresponding to the subject, which may indicate one or more sleep states (e.g., awake / non-asleep states, non-REM states, and REM states) of the subject observed in the video. The system 105 may transmit the sleep state data 152 to the device 102 for output to a user to view the results of processing the video data 104.

[0016] In some embodiments, the video data 104 may include video of more than one subject, and the system 105 may process the video data 104 to determine sleep state data for each subject represented in the video data 104.

[0017] The system 105 may be configured to determine various data from the video data 104 about the subject. To determine these data and to determine the sleep state data 152, the system 105 may include a number of different components. As shown in FIG. 1, the system 105 may include a segmentation component 110, a feature extraction component 120, a spectrum analysis component 130, a sleep state classification component 140, and a post-classification component 150. The system 105 may include fewer or more components than those shown in FIG. 1. In some embodiments, these various components may be located on the same physical system 105. In other embodiments, one or more of the various components may be located on different / separate physical systems 105. Communication between the various components may occur directly or via a network 199. Communication between the device 101, the system 105, and the device 102 may occur directly or via a network 199.

[0018] In some embodiments, one or more components illustrated as part of system 105 may be located on device 102 or a computing device connected to image capture device 101 (eg, device 400).

[0019] At a high level, the system 105 may be configured to process the video data 104 to determine a number of features corresponding to a subject, and to use the number of features to determine sleep state data 152 of the subject.

[0020] 2A is a flow chart illustrating a process 200 for determining sleep state data 152 of a subject according to an embodiment of the present disclosure. One or more of the steps of process 200 may be performed in another order / sequence than that shown in FIG. 2A. One or more steps of process 200 may be performed by components of system 105 shown in FIG.

[0021] At step 202 of process 200 shown in FIG. 2A, system 105 may receive video data 104 representing a video of an object. In some embodiments, video data 104 may be received by segmentation component 110 or provided by system 105 to segmentation component 110 for processing. In some embodiments, video data 104 may be video capturing an object in its natural state. The object may be in its natural state when there are no invasive methods applied to the object (e.g., no electrodes are inserted or attached to the object, no dye / color markings are applied to the object, no surgical methods are performed on the object, no invasive detection means are in or on the object, etc.). Video data 104 may be high resolution video of the object.

[0022] At step 204 of process 200 shown in FIG. 2A, segmentation component 110 may perform a segmentation process using video data 104 to determine ellipse data 112 (shown in FIG. 1). Segmentation component 110 may process video data 104 to generate a segmentation mask that identifies objects in video data 104 and then employ techniques to generate an ellipse fit / representation for the object. Segmentation component 110 may use one or more techniques (e.g., one or more ML models) for object tracking of the video / image data and may be configured to identify the object. Segmentation component 110 may generate a segmentation mask for each video frame of video data 104. The segmentation mask may indicate which pixels of the video frame correspond to the object and / or which pixels of the video frame correspond to the background / non-object. The segmentation component 110 may use a machine learning model to process the video data 104 to determine segmentation data indicative of a first set of pixels corresponding to the object and a second set of pixels corresponding to the background.

[0023] A video frame, as used herein, may be a portion of the video data 104. The video data 104 may be divided into multiple portions / frames of the same length / time. For example, a video frame may be 1 millisecond of the video data 104. In determining data such as a segmentation mask for a video frame of the video data 104, a component of the system 105, such as the segmentation component 110, may process a set of video frames (a window of video frames). For example, to determine a segmentation mask for an instant video frame, the segmentation component 110 may process (i) a set of video frames that occur before (in terms of time) the instant video frame (e.g., three video frames before the instant video frame), (ii) the instant video frame, and (iii) a set of video frames that occur after (in terms of time) the instant video frame (e.g., three video frames after the instant video frame). Thus, in this example, the segmentation component 110 may process seven video frames to determine a segmentation mask for one video frame. Such processing may be referred to herein as window-based processing of the video frames.

[0024] Using the segmentation mask of the video data 104, the segmentation component 110 may determine ellipse data 112. The ellipse data 112 may be an ellipse fit (an ellipse drawn around the body of the subject) to the object. For different types of objects, the system 105 may be configured to determine different shape fits / representations (e.g., a circle fit, a rectangle fit, a square fit, etc.). The segmentation component 110 may determine the ellipse data 112 as a subset of pixels of the segmentation mask that correspond to the object. The ellipse data 112 may include the subset of pixels. The segmentation component 110 may determine an ellipse fit of the object for each video frame of the video data 104. The segmentation component 110 may determine the ellipse fit for the video frames using window-based processing of the video frames as described above. The ellipse data 112 may be a vector or a matrix of pixels that represents an ellipse fit for all video frames of the video data 104. The segmentation component 110 may process the segmentation data to determine ellipse fit data 112 that corresponds to the object.

[0025] In some embodiments, the ellipse fit 112 to the object may define parameters of a portion of the object. For example, the ellipse fit may correspond to a position of the object and may include coordinates (e.g., x and y) that represent a pixel location (e.g., center of the ellipse) of the object in a video frame of the video data 104. The ellipse fit may correspond to a major axis length and a minor axis length of the object. The ellipse fit may include a sine and cosine of a vector angle of the major axis. The angle may be defined relative to the direction of the major axis. The major axis may extend from the tip of the subject's head or nose to the end of the subject's body, such as the base of the tail. The ellipse fit may also correspond to a ratio between the length of the major axis and the length of the minor axis of the object. In some embodiments, the ellipse data 112 may include the aforementioned measurements for all video frames of the video data 104.

[0026] In some embodiments, the segmentation component 110 may use one or more neural networks to process the video data 104 to determine the segmentation mask and / or ellipse data 112. In other embodiments, the segmentation component 110 may use other ML models, such as an encoder-decoder architecture, to determine the segmentation mask and / or ellipse data 112.

[0027] The ellipse data 112 may also include a confidence score of the segmentation component 110 in determining an ellipse fit for a video frame. The ellipse data 112 may alternatively include a probability or likelihood of an ellipse fit corresponding to an object.

[0028] In embodiments in which the video data 104 captures more than one object, the segmentation component 110 may identify each of the captured objects and may determine ellipse data 112 for each of the captured objects. The ellipse data 112 for each of the objects may be provided separately (in parallel or serially) to the feature extraction component 120 for processing.

[0029] At step 206 of process 200 shown in FIG. 2A, feature extraction component 120 may determine a plurality of features using ellipse data 112. Feature extraction component 120 may determine a plurality of features for each video frame of video data 104. In some exemplary embodiments, feature extraction component 120 may determine 16 features for each video frame of video data 104. The determined features may be stored as frame feature data 122 shown in FIG. 1. Frame feature data 122 may be a vector or matrix that includes values ​​of the plurality of features corresponding to each video frame of video data 104. Feature extraction component 120 may determine the plurality of features by processing segmentation data (determined by segmentation component 110) and / or ellipse data 112.

[0030] The feature extraction 120 may determine a number of features to include a number of visual characteristics of the object for each video frame of the video data 104. The following are example features that may be determined by the feature extraction component 120 and included in the frame feature data 122:

[0031] The feature extraction component 120 may process pixel information contained in the ellipse data 112. In some embodiments, the feature extraction component 120 may determine the major axis length, the minor axis length, and the ratio of the major axis length to the minor axis length for each video frame of the video data 104. These features may already be contained in the ellipse data 112, or the feature extraction component 120 may determine these features using pixel information contained in the ellipse data 112. The feature extraction component 120 may also determine the area (e.g., surface area) of the object using the ellipse fit information contained in the ellipse data 112. The feature extraction component 120 may determine the location of the object, represented as the center pixel of the ellipse fit. The feature extraction component 120 may also determine the change in the location of the object based on the change in the center pixel of the ellipse fit from one video frame to another (later occurring) video frame of the video data 104. The feature extraction component 120 may also determine the perimeter (e.g., circumference) of the ellipse fit.

[0032] The feature extraction component 120 may determine one or more (e.g., seven) Hu moments. Hu moments (also known as Hu moment invariants) may be a set of seven numbers calculated using the central moments of an image / video frame that are invariant to image transformations. The first six moments are proven to be invariant to translation, scale, rotation, and inversion, while the sign of the seventh moment changes with image inversion. In image processing, computer vision, and related fields, image moments are certain weighted averages (moments) of the intensities of image pixels, or functions of such moments, usually chosen to have some attractive properties or interpretations. Image moments are useful for describing objects after segmentation. The feature extraction component 120 may determine Hu image moments, which are a numerical description of the segmentation mask of an object, through integration and linear combination of the central image moments.

[0033] At step 208 of process 200 shown in FIG. 2A, the spectral analysis component 130 may perform a spectral analysis using the multiple features to determine frequency domain features 132 and time domain features 134. The spectral analysis component 130 may use signal processing techniques to determine the frequency domain features 132 and the time domain features 134 from the frame feature data 122. In some embodiments, the spectral analysis component 130 may determine a set of time domain features and a set of frequency domain features for each feature (from the feature data 122) of each video frame of the video data 104 in an epoch. In an exemplary embodiment, the spectral analysis component 130 may determine six time domain features for each feature of each video frame of the epoch. In some embodiments, the spectral analysis component 130 may determine fourteen frequency domain features for each feature of each video frame of the epoch. An epoch may be a duration of the video data 104, e.g., 10 seconds, 5 seconds, etc. The frequency domain features 132 may be a vector or a matrix representing the frequency domain features determined for each feature of the feature data 122 and for each epoch of the video frame. The time domain features 134 may be a vector or a matrix representing the time domain features determined for each feature of the feature data 122 and for each epoch of the video frame. The frequency domain features 132 may be graph data, for example, as shown in Figures 7A-7B and 8A-8B.

[0034] In an exemplary embodiment, the frequency domain features 132 may be the kurtosis of the power spectral density, the skewness of the power spectral density, the average power spectral density from 0.1 to 1 Hz, the average power spectral density from 1 to 3 Hz, the average power spectral density from 3 to 5 Hz, the average power spectral density from 5 to 8 Hz, the average power spectral density from 8 to 15 Hz, the total power spectral density, the maximum power spectral density, the minimum power spectral density, the mean of the power spectral density, and the standard deviation of the power spectral density.

[0035] In an exemplary embodiment, the time domain features 134 may be kurtosis, the mean of the feature signal, the median of the feature signal, the standard deviation of the feature signal, the maximum value of the feature signal, and the minimum value of the feature signal.

[0036] At step 210 of process 200 shown in FIG. 2A, sleep state classification component 140 may process frequency domain features 132 and time domain features 134 to determine a sleep prediction for a video frame of video data 104. Sleep state classification component 140 may determine a label representing a sleep state for each video frame of video data 104. Sleep state classification component 140 may classify each video frame into one of three sleep states: awake, non-REM, and REM. The awake state may be or may be similar to a non-sleep state. Sleep state classification component 140 may use frequency domain features 132 and time domain features 134 to determine a sleep state label for a video frame. In some embodiments, sleep state classification component 140 may use window-based processing for the above-mentioned video frames. For example, to determine a sleep state indicator for an instant video frame, the sleep state classification component 140 may process data (time and frequency domain features 132, 134) for a set of video frames occurring before the instant video frame and data for a set of video frames occurring after the instant video frame. The sleep state classification component 140 may output frame prediction data 142, which may be a vector or matrix of sleep state indicators for each video frame of the video data 104. The sleep state classification component 140 may also determine a confidence score associated with the sleep state indicator, which may represent the likelihood of the video frame corresponding to the indicated sleep state or the reliability of the sleep state classification component 140 in determining a sleep state indicator for the video frame. The confidence score may be included in the frame prediction data 142.

[0037] The sleep state classification component 140 may determine the frame prediction data 142 from the frequency domain features 132 and the time domain features 134 using one or more ML models. In some embodiments, the sleep state classification component 140 may use a gradient boosting ML technique (e.g., XGBoost technique). In other embodiments, the sleep state classification component 140 may use a random forest ML technique. In still other embodiments, the sleep state classification component 140 may use a neural network ML technique (e.g., multi-layer perceptron (MLP)). In still other embodiments, the sleep state classification component 140 may use a logistic regression technique. In still other embodiments, the sleep state classification component 140 may use a singular value decomposition (SVD) technique. In some embodiments, the sleep state classification component 140 may use a combination of one or more of the aforementioned ML techniques. The ML techniques may be trained to classify video frames of video data for a subject into sleep states, as described in connection with FIG. 3 below.

[0038] In some embodiments, the sleep state classification component 140 may use additional or alternative data / features (e.g., the video data 104, the ellipse data 112, the frame feature data 122, etc.) to determine the frame prediction data 142.

[0039] The sleep state classification component 140 may be configured to recognize transitions between one sleep state and another sleep state based on variations between the frequency domain features 132 and the time domain features 134. For example, the frequency domain and time domain signals for the region of the object differ in time and frequency for the awake, non-REM, and REM states. As another example, the frequency domain and time domain signals for the width-to-length ratio (ratio of major axis length to minor axis length) of the object differ in time and frequency for the awake, non-REM, and REM states. In some embodiments, the sleep state classification component 140 may determine the frame prediction data 142 using one of a plurality of features (e.g., the body region of the object or the width-to-length ratio). In other embodiments, the sleep state classification component 140 may determine the frame prediction data 142 using a combination of features from a plurality of features (e.g., the body region of the object and the width-to-length ratio).

[0040] 2A, post-classification component 150 may perform post-classification processing to determine sleep state data 152 representing the subject's sleep states and transitions between sleep states during the duration of the video. Post-classification component 150 may process frame prediction data 142, which includes sleep state indicators (and corresponding confidence scores) for each video frame, to determine sleep state data 152. Post-classification component 150 may use a transition model to determine transitions from a first sleep state to a second sleep state.

[0041] The transitions between the awake, non-REM, and REM states are not random and generally follow expected patterns. For example, typically, a subject transitions from an awake state to a non-REM state and then from a non-REM state to a REM state. The post-classification component 150 may be configured to recognize these transition patterns and use a transition probability matrix and occurrence probability for a given state. The post-classification component 150 may serve as a validation component for the frame prediction data 142 determined by the sleep state classification component 140. For example, in some cases, the sleep state classification component 140 may determine that a first video frame corresponds to a awake state and a subsequent second video frame corresponds to a REM state. In such a case, the post-classification component 150 may update the sleep state of the first or second video frame based on knowing that a transition from a awake state to a REM state is unlikely, especially in the short period covered by the video frames. The post-classification component 150 may use window-based processing of the video frames to determine the sleep state of the video frames. In some embodiments, post-classification component 150 may also take into account the duration of a sleep state before transitioning to another sleep state. For example, post-classification component 150 may determine whether the sleep state of a video frame determined by sleep state classification component 140 is accurate based on how long a non-REM state lasts for a subject in video data 104 before transitioning to a REM state. In some embodiments, post-classification component 150 may use various techniques, such as, for example, statistical models (e.g., Markov models, hidden Markov models, etc.), probabilistic models, etc. The statistical or probabilistic models may model dependencies between the sleep states (wake, non-REM, and REM states).

[0042] Post-classification component 150 may process frame prediction data 142 to determine the duration of one or more sleep states (awake, non-REM, REM) of a subject depicted in video data 104. Post-classification component 150 may process frame prediction data 142 to determine the frequency (number of times a sleep state occurs in video data 104) of one or more sleep states (awake, non-REM, REM) of a subject depicted in video data 104. Post-classification component 150 may process frame prediction data 142 to determine changes in one or more sleep states of a subject. Sleep state data 152 may include the duration of one or more sleep states of a subject, the frequency of one or more sleep states of a subject, and / or changes in one or more sleep states of a subject.

[0043] Post-classification component 150 can output sleep state data 152, which may be a vector or matrix including a sleep state indicator for each video frame of video data 104. For example, sleep state data 152 may include a first indicator "Awake" corresponding to a first video frame, a second indicator "Awake" corresponding to a second video frame, a third indicator "NREM" corresponding to a third video frame, a fourth indicator "REM" corresponding to a fourth video frame, etc.

[0044] The system 105 may transmit the sleep state data 152 to the device 102 for display. The sleep state data 152 may be presented as graph data, for example, as shown in Figures 11A-D.

[0045] As described herein, in some embodiments, the automatic sleep state system 100 may use multiple features (determined by the feature extraction component 120) to determine multiple body regions of the subject, each body region corresponding to a video frame of the video data 104, and the automatic sleep state system 100 may determine the sleep state data 152 based on changes in the multiple body regions in the video.

[0046] As described herein, in some embodiments, the automatic sleep state system 100 may use the features (determined by the feature extraction component 120) to determine a plurality of width-to-length ratios, each width-to-length ratio of the plurality of width-to-length ratios corresponding to a video frame of the video data 104, and the automatic sleep state system 100 may determine the sleep state data 152 based on changes in the plurality of width-to-length ratios during the video.

[0047] In some embodiments, the automatic sleep state system 100 may detect a transition from a NREM state to a REM state based on a change in a subject's body region or shape, which may be the result of muscle relaxation. Such transition information may be included in the sleep state data 152.

[0048] Other correlations between features derived from the video data 104 and a subject's sleep state that the automatic sleep state system 100 may use are described below in the Examples section.

[0049] In some embodiments, the automatic sleep state system 100 may be configured to determine the subject's respiration / respiration rate by processing the video data 104. The automatic sleep state system 100 may determine the subject's respiration rate by processing a plurality of features (determined by the feature extraction component 120). In some embodiments, the automatic sleep state system 100 may use the respiration rate to determine the subject's sleep state data 152. In some embodiments, the automatic sleep state system 100 may determine the respiration rate based on frequency domain and / or time domain features determined by the spectral analysis component 130.

[0050] The subject's breathing rate may vary between sleep states and may be detected using features derived from the video data 104. For example, the subject's body area and / or width-to-length ratio may change over a period of time such that a signal representation (time or frequency) of the body area and / or width-to-length ratio may be a consistent signal between 2.5-3 Hz. Such a signal representation may look like a ventilation waveform. The automatic sleep state system 100 may process the video data 104 to extract features representative of changes in body shape and / or changes in chest size that correlate / correspond to breathing by the subject. Such changes may be visible in the video and may be extracted as time-domain and frequency-domain features.

[0051] During a NREM state, the subject may have a particular respiration rate, for example, 2.5-3 Hz. The automatic sleep state system 100 may be configured to recognize a particular correlation between respiration rate and sleep state. For example, the width-to-length ratio signal may be more prominent / evident in a NREM state than in a REM state. As a further example, the width-to-length ratio signal may change more during a REM state. The aforementioned exemplary correlation may be a result of the subject's respiration rate changing more during a REM state than during a NREM state. Another exemplary correlation may be low frequency noise captured in the width-to-length ratio signal during a NREM state. Such a correlation may be due to the subject's motion / movement to adjust its sleep position during a NREM state, and the subject may not move during a REM state due to muscle relaxation.

[0052] At least the width-to-length ratio signal (and other signals for other features) derived from the video data 104 illustrates that the video data 104 captures visual movement of the subject's abdomen and / or chest, which can be used to determine the subject's respiratory rate.

[0053] 2B is a flow chart illustrating a process 250 for determining sleep state data 152 for multiple objects depicted in a video according to an embodiment of the present disclosure. One or more of the steps of process 250 may be performed in a different order / sequence than that shown in FIG. 2B. One or more steps of process 250 may be performed by components of system 105 shown in FIG.

[0054] At step 252 of process 250 shown in FIG. 2B, system 105 may receive video data 104 representing an image of multiple objects (eg, as shown in FIG. 6C).

[0055] At step 254, the segmentation component 110 may perform an instance segmentation process using the video data 104 to identify individual objects depicted in the video. The segmentation component 110 may process the video data 104 using an instance segmentation technique to generate segmentation masks that identify individual objects in the video data 104. The segmentation component 110 may generate a first segmentation mask for a first object, a second segmentation mask for a second object, etc., where the individual segmentation masks may indicate which pixels in the video frames correspond to the respective objects. The segmentation component 110 may also determine which pixels in the video frames correspond to background / non-objects. The segmentation component 110 may process the video data 104 using one or more machine learning models to determine first segmentation data indicative of a first set of pixels in the video frames that correspond to the first object, second segmentation data indicative of a second set of pixels in the video frames that correspond to the second object, etc.

[0056] The segmentation component 110 may track respective segmentation masks for individual objects using indicators (e.g., text indicators, numeric indicators, or other data) such as “Object 1”, “Object 2”, etc. The segmentation component 110 may assign respective indicators to segmentation masks determined from various video frames of the video data 104, thus tracking sets of pixels corresponding to individual objects through multiple video frames. The segmentation component 110 may be configured to track individual objects across multiple video frames, even as the objects move, change position, or change location. The segmentation component 110 may also be configured to identify individual objects when the objects are in close proximity to one another, for example, as shown in FIG. 6C. In some cases, subjects may prefer to sleep in close proximity or near one another, and the instance segmentation technique can identify individual objects even when this occurs.

[0057] Instance segmentation techniques may involve the use of computer vision techniques, algorithms, models, etc. Instance segmentation may involve identifying each object instance in an image / video frame and assigning a label to each pixel of the video frame. Instance segmentation may use object detection techniques to identify all objects in a video frame, classify individual objects, and localize each object instance using a segmentation mask.

[0058] In some embodiments, the system 105 may identify and track individual subjects from multiple subjects based on some metrics about the subject, such as body size, body shape, body / hair color, etc.

[0059] At step 256 of process 250, segmentation component 110 may use the segmentation masks for each object to determine ellipse data 112 for each object. For example, segmentation component 110 may determine first ellipse data 112 using a first segmentation mask for a first object, second ellipse data 112 using a second segmentation mask for a second object, etc. Segmentation component 110 may determine ellipse data 112 in a manner similar to that described above in connection with process 200 shown in FIG. 2A.

[0060] At step 258 of process 250, feature extraction component 120 may determine a number of features for each individual object using each ellipse data 112. The number of features may be frame-based features, i.e., a number of features may be for each individual video frame of video data 104 and provided as frame feature data 122. Feature extraction component 120 may determine a first frame feature data 122 using the first ellipse data 112 and corresponding to the first object, a second frame feature data 122 using the second ellipse data 112 and corresponding to the second object, and so on. Feature extraction component 120 may determine frame feature data 122 in a manner similar to that described above in connection with process 200 shown in FIG. 2A.

[0061] At step 260 of process 250, spectral analysis component 130 may use the multiple features to perform spectral analysis (in a manner similar to that described above in connection with process 200 shown in FIG. 2A) to determine frequency domain features 132 and time domain features 134 for the individual objects. Spectral analysis component 130 may determine a first frequency domain feature 132 for the first object, a second frequency domain feature 132 for the second object, a first time domain feature 134 for the first object, a second time domain feature 134 for the second object, etc.

[0062] At step 262 of process 250, sleep state classification component 140 may process each frequency domain feature 132 and time domain feature 134 for each subject to determine a sleep prediction for the individual subject (in a manner similar to that described above in connection with process 200 shown in FIG. 2A ) for the video frames of video data 104. For example, sleep state classification component 140 may determine a first frame prediction data 142 for a first subject, a second frame prediction data 142 for a second subject, etc.

[0063] At step 264 of process 250, post-classification component 150 may perform post-classification processing (in a manner similar to that described above in connection with process 200 shown in FIG. 2A) to determine sleep state data 152 representative of the sleep states of individual subjects during the duration of the footage and transitions between sleep states. For example, post-classification component 150 may determine first sleep state data 152 for a first subject, second sleep state data 152 for a second subject, etc.

[0064] In this manner, using instance segmentation techniques, the system 105 may identify multiple objects in a video and use feature data (and other data) corresponding to each object to determine sleep state data for each individual object. By being able to identify each object even when they are close to each other, the system 105 may determine the sleep state of multiple objects housed together (i.e., multiple objects contained in the same enclosure). One advantage of this is that the object may be observed in a natural environment under natural conditions, which may involve cohabitation with another object. In some cases, the behavior of other objects may also be identified / study based on the cohabitation of the objects (e.g., the effect of cohabitation on sleep state, whether cohabitation causes objects to follow the same / similar sleep patterns, etc.). Another advantage is that the sleep state data for multiple objects may be determined by processing the same / single video, which may reduce resources used (e.g., time, computational resources, etc.) compared to resources used to process multiple separate videos, each of which represents one object.

[0065] Figure 3 conceptually illustrates components and data that may be used to configure the sleep state classification component 140 shown in Figure 1. As described herein, the sleep state classification component 140 may include one or more ML models for processing features derived from the video data 104. The ML models may be trained / configured using various types of training data and training techniques.

[0066] In some embodiments, the spectral training data 302 may be processed by the model building component 310 to train / configure the trained classifier 315. In some embodiments, the model building component 310 may also process the EEG / EMG training data to train / configure the trained classifier 315. The trained classifier 315 may be configured to determine a sleep state indicator for a video frame based on one or more features corresponding to the video frame.

[0067] The spectral training data 302 may include frequency and / or time domain signals for one or more features of the objects depicted in the video data used for training. Such features may correspond to features determined by the feature extraction component 120. For example, the spectral training data 302 may include frequency and / or time domain signals corresponding to body regions of the objects in the video. The frequency and / or time domain signals may be annotated / labeled with corresponding sleep states. The spectral training data 302 may also include frequency and / or time domain signals for other features such as object width-length ratio, object width, object length, object location, Hu image moments, and other features.

[0068] The EEG / EMG training data 304 may be electroencephalogram (EEG) data and / or electromyogram (EMG) data corresponding to subjects used to train / configure the sleep state classification component 140. The EEG and / or EMG data may be annotated / labeled with corresponding sleep states.

[0069] The spectral training data 302 and the EEG / EMG training data 304 may correspond to sleep of the same subject. The model building component 310 may correlate the spectral training data 302 and the EEG / EMG training data 304 to train / construct a trained classifier 315 to identify sleep states from the spectral data (frequency and time domain features).

[0070] There may be an imbalance in the training dataset due to subjects experiencing more non-REM than REM states during sleep. A balanced training dataset may be generated to include the same / similar number of REM, non-REM, and wakefulness states to train / configure the trained classifier 315.

[0071] subject Some aspects of the invention include determining sleep state data of a subject. As used herein, the term "subject" may refer to a human, a non-human primate, a cow, a horse, a pig, a sheep, a goat, a dog, a cat, a pig, a bird, a rodent, or other suitable vertebrate or invertebrate. In certain embodiments of the invention, the subject is a mammal, and in certain embodiments of the invention, the subject is a human. In some embodiments, the subject used in the methods of the invention is a rodent, including but not limited to mice, rats, gerbils, hamsters, and the like. In some embodiments of the invention, the subject is a normal, healthy subject, and in some embodiments, the subject is known to have a disease or condition, or is at risk of having a disease or condition, or is suspected of having a disease or condition. In certain embodiments of the invention, the subject is an animal model for a disease or condition. For example, and not intended to be limiting, in some embodiments of the invention, the subject is a mouse, which is an animal model for sleep apnea.

[0072] As a non-limiting example, the subject evaluated by the method and system of the present invention may be a subject having, suspected to have, and / or animal model for a condition such as sleep apnea, insomnia, narcolepsy, brain injury, depression, psychiatric disease, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, neurological condition that can change the status of sleep state, and metabolic disorder or condition that can change sleep state. A non-limiting example of a metabolic disorder or condition that can change sleep state is a high-fat diet. Additional physical conditions may also be evaluated using the method of the present invention, non-limiting examples of which are obesity, overweight, the effect of administered medication, and / or the effect of alcohol intake. Additional diseases and conditions can also be evaluated using the method of the present invention, including but not limited to sleep pathologies resulting from chronic disease, drug abuse, injury, etc.

[0073] The method and system of the present invention may also be used to evaluate subjects or test subjects that do not have one or more of sleep apnea, insomnia, narcolepsy, brain injury, depression, psychiatric disease, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, neurological conditions that can change the status of sleep state, and metabolic disorders or pathologies that can change sleep state.In some embodiments, the method of the present invention is used to evaluate the sleep state of subjects that are not obese, overweight, or alcoholic.These subjects can serve as control subjects, and the results of evaluation using the method of the present invention can be used as control data.

[0074] In some embodiments of the method of the present invention, the subject is a wild-type subject. As used herein, the term "wild-type" refers to the phenotype and / or genotype of a species that is typical of the form in which it occurs in nature. In certain embodiments of the present invention, the subject is a non-wild-type subject, e.g., a subject that has one or more genetic modifications compared to the wild-type genotype and / or phenotype of the species of the subject. In some instances, the difference in the subject's genotype / phenotype compared to the wild-type is due to inherited (germline) mutation or acquired (somatic) mutation. Factors that can cause a subject to exhibit one or more somatic mutations include, but are not limited to, environmental factors, toxins, ultraviolet radiation, spontaneous errors occurring in cell division, radiation, maternal infection, chemicals, and other teratogenic events, but are not limited to these.

[0075] In certain embodiments of the method of the present invention, the subject is a genetically modified organism, also referred to as a genetically modified subject. A genetically modified subject may include preselected and / or deliberate genetic modifications, and therefore exhibit one or more genotypic and / or phenotypic traits that are different from those in non-modified subjects. In some embodiments of the present invention, by using conventional genetic engineering techniques, a genetically modified subject can be produced that exhibits genotypic and / or phenotypic differences compared to non-modified subjects of that kind. As a non-limiting example, a genetically modified mouse in which a functional gene product is absent or present at reduced levels in the mouse, and the method or system of the present invention can be used to evaluate the phenotype of the genetically modified mouse, and the results can be compared to those obtained from a control (control results).

[0076] In some embodiments of the present invention, a subject may be monitored using the automatic sleep state determination method or system of the present invention, and the presence or absence of a sleep disorder or condition may be detected. In certain embodiments of the present invention, a test subject that is an animal model of a sleep condition may be used to evaluate the test subject's response to the sleep condition. In addition, a test subject that is an animal model of a sleep and / or activity condition may be administered a candidate therapeutic agent or method that is monitored using the automatic sleep state determination method and / or system of the present invention, and the results may be used to determine the effectiveness of the candidate therapeutic agent for treating the condition.

[0077] As described elsewhere herein, the methods and systems of the present invention may be configured to determine a sleep state of a subject regardless of the subject's physical characteristics. In some embodiments of the present invention, one or more physical characteristics of the subject may be pre-specified characteristics. For example, and not intended to be limiting, the pre-specified physical characteristics may be one or more of body type, build, coat color, sex, age, and disease or condition phenotype.

[0078] Testing and Screening of Control and Candidate Compounds The results obtained for a subject using the method or system of the present invention can be compared with the results for a control.The method of the present invention can also be used to evaluate the phenotypic difference between a subject and a control.Thus, some aspects of the present invention provide a method for determining the presence or absence of changes in one or more sleep states in a subject compared to a control.Some embodiments of the present invention include using the method of the present invention to identify phenotypic characteristics of a disease or condition, and in certain embodiments of the present invention, automated phenotyping is used to evaluate the effect of a candidate therapeutic compound on a subject.

[0079] The results obtained using the method or system of the present invention can be advantageously compared to a control. In some embodiments of the present invention, one or more subjects can be evaluated using the method of the present invention, and then the subjects can be re-tested after administering a candidate therapeutic compound to the subjects. In connection with a subject evaluated using the method or system of the present invention, the terms "subject" and "test subject" may be used herein, and the terms "subject" and "test subject" are used interchangeably herein. In certain embodiments of the present invention, the results obtained using the method of the present invention to evaluate a test subject are compared to results obtained from a method performed on another test subject. In some embodiments of the present invention, the results of a test subject are compared to results of sleep state evaluations performed on the test subject at different times. In some embodiments of the present invention, the results obtained using the method of the present invention to evaluate a subject are compared to control results.

[0080] As used herein, a control result may be a pre-determined value that can take a variety of forms. It may be a single cut-off value, such as a median or mean value. It may be established based on a comparison group, such as subjects evaluated using the system or method of the present invention under similar conditions as the test subjects, where the test subjects are administered the candidate therapeutic agent and the comparison group is not administered the candidate therapeutic agent. Another example of a comparison group may include subjects known to have a disease or condition and a group without the disease or condition. Another comparison group may be subjects with a family history of the disease or condition and subjects from a group without such a family history. The pre-determined value may be set, for example, so that the tested population is grouped evenly (or unevenly) based on the test results. Those skilled in the art will be able to select appropriate control groups and values ​​for use in the comparison method of the present invention.

[0081] A subject evaluated using the method or system of the present invention may be monitored for the presence or absence of changes in one or more sleep state characteristics occurring in a test condition compared to a control condition. As a non-limiting example, the changes occurring in a subject may include, but are not limited to, one or more sleep state characteristics, such as the duration of a sleep state, the time interval between two sleep states, the number of one or more sleep states during a sleep period, the ratio of RM sleep states to NRM sleep states, and the period before entering a sleep state. The method and system of the present invention may be used on a test subject to evaluate the effect of a disease or condition on the test subject, and may also be used to evaluate the effectiveness of a candidate therapeutic agent. As a non-limiting example of a method of the present invention for evaluating the presence or absence of changes in one or more characteristics of a subject's sleep state as a means of identifying the effectiveness of a candidate therapeutic agent, a subject known to have a disease or condition that affects the subject's sleep state is evaluated using the method of the present invention. The subject is then administered the candidate therapeutic agent and evaluated again using the method. The presence or absence of changes in the test subject's results indicates the presence or absence of the effect of the candidate therapeutic agent on the disease or condition that affects the sleep state, respectively.

[0082] It will be appreciated that in some embodiments of the invention, a subject may act as its own control, for example, by being assessed more than once using a method of the invention and comparing results obtained at two or more of the different assessments. The methods and systems of the invention may be used to assess the progression or regression of a subject's disease or condition using embodiments of the method or system of the invention, by identifying and comparing changes in phenotypic characteristics, such as sleep state characteristics, of a subject over time using two or more assessments of the subject.

[0083] Exemplary Devices and Systems One or more components of the automatic sleep state system 100 may implement an ML model, which may take many forms, including an XgBoost model, a random forest model, a neural network, a support vector machine, or other model, or any combination of these models.

[0084] Various machine learning techniques may be used to train and operate the model to perform various steps described herein, such as determining segmentation masks, determining ellipse data, determining feature data, and determining sleep state data. The model may be trained and operated according to various machine learning techniques. Such techniques may include, for example, neural networks (deep neural networks and / or recurrent neural networks, etc.), inference engines, trained classifiers, etc. Examples of trained classifiers include support vector machines (SVMs), neural networks, decision trees, AdaBoost (short for "adaptive boost") combined with decision trees, and random forests. Focusing on SVMs as an example, SVMs are supervised learning models with associated learning algorithms that analyze data to recognize patterns in the data, and are commonly used for classification and regression analysis. Given a set of training examples, each marked as belonging to one of two categories, the SVM training algorithm builds a model that assigns new examples to one category or the other, making it a non-probabilistic binary linear classifier. More complex SVM models may be built with a training set that identifies three or more categories, and the SVM determines which category is most similar to the input data. The SVM model may be mapped such that examples of distinct categories are separated by a clear gap. New examples are then mapped into the same space and predicted to belong to a category based on which side of the gap they fall on. The classifier may issue a "score" indicating which category the data most closely matches. The score may provide an index of how closely the data matches the category.

[0085] A neural network may include several layers, ranging from an input layer to an output layer. Each layer is configured to receive a particular type of data as input and to output another type of data. The output from one layer is received as input to the next layer. Although the values ​​of the input / output data at a particular layer are unknown until the neural network actually operates at run time, the data describing the neural network describes the structure, parameters, and operation of the layers of the neural network.

[0086] One or more of the intermediate layers of a neural network may also be known as a hidden layer. Each node of the hidden layer is connected to each node of the input layer and to each node of the output layer. If the neural network includes multiple intermediate networks, each node of the hidden layer will be connected to each node in the next higher layer and in the next lower layer. Each node of the input layer represents a potential input to the neural network, and each node of the output layer represents a potential output from the neural network. Each connection from one node to another node in the next layer may be associated with a weight or score. The neural network may output a single output or a weighted set of possible outputs. Different types of neural networks may be used, such as recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep neural networks (DNNs), long short-term memories (LSTMs), and / or others.

[0087] The processing by a neural network is determined by the learned weights of each node input and the structure of the network: given a particular input, the neural network determines the output, one layer at a time, until the output layer of the entire network has been calculated.

[0088] The connection weights may be initially learned by the neural network during training, where a given input is associated with a known output. In a set of training data, various training examples are fed into the network. In each example, the weight of the correct connection from the input to the output is typically set to 1, and all connections are given a weight of 0. When the training data examples have been processed by the neural network, the inputs may be sent to the network and compared with the associated outputs to determine how the network performance compares to a target performance. Training techniques such as backpropagation may be used to update the neural network weights to reduce errors made by the neural network when processing the training data.

[0089] In order to apply machine learning techniques, the machine learning process itself needs to be trained. In this case, in order to train a machine learning component, such as either the first model or the second model, a "ground truth" for the training examples needs to be established. In machine learning, the term "ground truth" refers to the accuracy rate of classification of a training set for a supervised learning technique. Various techniques may be used to train the model, including backpropagation, statistical learning, supervised learning, semi-supervised learning, probability learning, or other known techniques.

[0090] FIG. 4 is a block diagram conceptually illustrating a device 400 that may be used with the system. FIG. 5 is a block diagram conceptually illustrating exemplary components of a remote device, such as the system 105, that may assist in processing video data, identifying subject behavior, and the like. The system 105 may include one or more servers. As used herein, "server" may refer to a traditional server as understood in a server / client computing architecture, but may also refer to a number of different computing components that may assist in the operations described herein. For example, a server may include one or more physical computing components (such as a rack server) that are connected physically and / or via a network to other devices / components and may perform computing operations. A server may also include one or more virtual machines that emulate a computer system and run on one device or across multiple devices. A server may also include other combinations of hardware, software, firmware, or the like, for performing the operations described herein. The server may be configured to operate using one or more of a client-server model, a computer bureau model, grid computing technology, fog computing technology, mainframe technology, utility computing technology, a peer-to-peer model, sandbox technology, or other computing technologies.

[0091] Multiple systems 105 may be included in the overall system of the present disclosure, such as one or more systems 105 for determining ellipse data, one or more systems 105 for determining frame features, one or more systems 105 for determining frequency domain features, one or more systems 105 for determining time domain features, one or more systems 105 for determining frame-based sleep marker predictions, one or more systems 105 for determining sleep state data, etc. In operation, each of these systems may include computer readable and computer executable instructions present on the respective device 105, as described further below.

[0092] Each of these devices (400 / 105) may include one or more controller / processors (404 / 504), which may include a central processing unit (CPU) for processing data and computer-readable instructions, and a memory (406 / 506) for storing the respective device's data and instructions. The memory (406 / 506) may individually include volatile random access memory (RAM), non-volatile read-only memory (ROM), non-volatile magnetoresistive memory (MRAM), and / or other types of memory. Each device (400 / 105) may also include a data storage component (408 / 508) for storing data and controller / processor executable instructions. Each data storage component (408 / 508) may individually include one or more non-volatile storage types, such as magnetic storage, optical storage, solid-state storage, etc. Each device (400 / 105) may also be connected to removable or external non-volatile memory and / or storage (removable memory cards, memory key drives, network storage, etc.) via a respective input / output device interface (402 / 502).

[0093] Computer instructions for operating each device (400 / 105) and its various components may be executed by the respective device's controller / processor (404 / 504), using the memory (406 / 506) as temporary "working" storage during execution. The device's computer instructions may be stored in a non-transitory manner in non-volatile memory (406 / 506), in storage (408 / 508), or in an external device. Alternatively, some or all of the executable instructions may be embedded in hardware or firmware on the respective device in addition to or instead of software.

[0094] Each device (400 / 105) includes an input / output device interface (402 / 502). As described further below, various components may be connected via the input / output device interfaces (402 / 502). In addition, each device (400 / 105) may include an address / data bus (424 / 524) for transmitting data between the components of each device. Each component within the device (400 / 105) may also be directly connected to other components in addition to (or instead of) being connected to other components via the bus (424 / 524).

[0095] 4, device 400 may include an input / output device interface 402 for connecting to various components, such as an audio output component, such as a speaker 412, a wired or wireless headset (not shown), or other component that may output audio. Device 400 may additionally include a display 416 for displaying content. Device 400 may further include a camera 418.

[0096] Via antenna 414, input / output device interface 402 may be connected to one or more networks 199 via a wireless local area network (WLAN) (such as WiFi) radio, Bluetooth, and / or a wireless network radio, e.g., a radio capable of communicating with wireless communication networks such as a long-term evolution (LTE) network, a WiMAX network, a 3G network, a 4G network, a 5G network, etc. Wired connections such as Ethernet may also be supported. Via network 199, the system may be distributed across a network environment. I / O device interface (402 / 502) may also include communication components that allow data to be exchanged between devices, such as different physical servers, in a collection of servers or other components.

[0097] The components of device 400 or system 105 may include their own dedicated processors, memory, and / or storage. Alternatively, one or more of the components of device 400 or system 105 may utilize the I / O interfaces (402 / 502), processors (404 / 504), memory (406 / 506), and / or storage (408 / 508) of device 400 or system 105, respectively.

[0098] As mentioned above, multiple devices may be employed within a single system. In such a multi-device system, each of the devices may include different components for performing different aspects of the system's processing. Multiple devices may include overlapping components. The components of device 400 and system 105 described herein are exemplary and may be deployed as stand-alone devices or may be included in whole or in part as components of a larger device or system.

[0099] The concepts disclosed herein may be applied within many different devices and computer systems, including, for example, general purpose computing systems, video / image processing systems, and distributed computing environments.

[0100] The above-described aspects of the present disclosure are meant to be illustrative. They are selected to illustrate the principles and applications of the present disclosure, and are not intended to be exhaustive or to limit the present disclosure. Many modifications and variations of the disclosed aspects may be apparent to those skilled in the art. Those skilled in the art of computer and voice processing will recognize that the components and process steps described herein may be interchangeable with other components or steps, or with combinations of components or steps, and still achieve the benefits and advantages of the present disclosure. Moreover, it will be apparent to those skilled in the art that the present disclosure may be practiced without some or all of the specific details and steps disclosed herein.

[0101] Aspects of the disclosed system may be implemented as a computer method or as an article of manufacture, such as a memory device or a non-transitory computer-readable storage medium. The computer-readable storage medium may be readable by a computer and may contain instructions for causing a computer or other device to perform the processes described in this disclosure. The computer-readable storage medium may be implemented by volatile computer memory, non-volatile computer memory, hard drives, solid-state memory, flash drives, removable disks, and / or other media. In addition, components of the system may be implemented as firmware or hardware. EXAMPLES

[0102] Example 1. Development of a mouse sleep state classifier model method Animal husbandry, surgery, and experimental set-up Sleep studies were performed in 17 C57BL / 6J (The Jackson Laboratory, Bar Harbor, MD, USA) male mice. C3H / HeJ (The Jackson Laboratory, Bar Harbor, MD, USA) mice were also imaged without surgery for characterization. All mice were obtained at 10–12 weeks of age. All animal studies were performed in accordance with the guidelines published by the National Institutes of Health for the care and use of laboratory animals and were approved by the University of Pennsylvania Animal Care and Use Committee. Test methods were performed as previously described [Pack, AI et al. Physiol. Genomics. 28(2):232-238 (2007);McShane, BB et al., Sleep. 35(3):433-442 (2012)].

[0103] Briefly, mice were housed individually in standard mouse cages (6 x 6 inches) without lids. The height of each cage was extended to 12 inches to prevent mice from jumping out of the cage. This design allowed for simultaneous assessment of mouse behavior by video and sleep / wake stages by EG / EMG recordings. Animals were provided with water and food ad libitum and were kept on a 12-h light / dark cycle. During the light phase, the lux level at the bottom of the cage was 80 lux. For EEG recordings, four silver ball electrodes were placed on the skull (two frontal and two parieto-temporal). For EMG recordings, two silver wires were sutured to the dorsal nuchal muscle. All leads were placed subcutaneously in the center of the skull and connected to a plastic socket pedestal (Plastics One, Torrington, CT) that was fixed to the skull with dental cement. Electrodes were implanted under general anesthesia. After surgery, animals were allowed a 10-day recovery period before recordings.

[0104] EEG / EMG acquisition For EEG / EMG recordings, raw signals were read and amplified (20,000×) using Grass Gamma Software (Astro-Med, West Warwick, RI). Signal filter settings for EEG were a low cutoff frequency of 0.1 Hz and a high cutoff frequency of 100 Hz. Signal filter settings for EMG were a low cutoff frequency of 10 Hz and a high cutoff frequency of 100 Hz. Recordings were digitized at 256 Hz samples / sec / channel.

[0105] Acquiring footage A night vision setup on a Raspberry Pi 3 Model B (Raspberry Pi Foundation, Cambridge, UK) was used to record high quality video data in both day and night conditions. A SainSmart (SainSmart, Las Vegas, NV) infrared night vision security camera was used, accompanied by infrared LEDs to illuminate the scene in the absence of visible light. The camera was mounted 18 inches above the home cage floor, providing a top-down view of the mouse for observation. During the day, the video data was in color. At night, the video data was in black and white. Video was recorded at 1920x1080 pixel resolution and 30 frames per second using v412-ctl capture software. For information regarding aspects of the V412-CTL software, see, for example, www.kernel.org / doc / html / latest / userspace-api / media / v4l / v4l2.html or alternatively, an abridged version: www.kernel.org / .

[0106] Synchronization of video and EEG / EMG data Computer clock time was used to synchronize the video and EEG / EMG data. The EEG / EMG data collection computer was used as the source clock. Visual cues were added to the video at known time points on the EEG / EMG computer. Visual cues typically lasted 2-3 frames in the video, suggesting that possible errors in synchronization could be up to 100 ms. Because EEG / EMG data were analyzed at ten second (10s) intervals, possible errors in temporal alignment would be negligible.

[0107] EEG / EMG annotation of training data 24-hour synchronized video and EEG / EMG data were collected for 17 C57BL / 6J male mice aged 10–12 weeks from Jackson Laboratory. Both EEG / EMG data and video were divided into 10-second epochs, and each epoch was scored by trained scorers and labeled as REM, NREM, or awake phase based on the EEG and EMG signals. A total of 17,700 EEG / EMG epochs were scored by human experts. Among them, 48.3%+ / -6.9% of the epochs were annotated as awake, 47.6%+ / -6.7% as NREM, and 4.1%+ / -1.2% as REM phase. In addition, the SPINDLE method was applied for the second annotation [Miladinovic, D. et al., PLoS Comput Biol. 15, e1006968 (2019)]. Similar to the human experts, 52% of epochs were annotated as awake, 44% as NREM, and 4% as REM. Because SPINDLE annotated 4-second (4s) epochs, we combined three consecutive epochs and compared them to a 10-s epoch, and only compared epochs when the three 4-s epochs did not change. When a given epoch correlated, the agreement between the human annotations and SPINDLE was 92% (89% awake, 95% NREM, 80% REM).

[0108] Data Preprocessing Starting from the video data, we applied the segmentation neural network architecture described above to generate a mask of the mouse. [Webb JM and Fu YH., Curr. Opin. Neurobiol. 69:19-24 (2021)]. 313 frames were annotated to train the segmentation network. A 4x4 diamond dilation followed by a 5x5 diamond shrinkage filter was applied to the raw predicted segmentation. These stationary operations were used to improve the segmentation quality. Using the predicted segmentation and the resulting ellipse fit, various per-frame image measurement signals were extracted from each frame, as described in Table 1.

[0109] All these measurements (Table 1) were calculated by applying OpenCV contour functions to the neural network predicted segmentation mask. OpenCV functions used included fitEllipse, contourArea, arcLength, moments, and getHuMoments. For more information on the OpenCV software, see, for example, http: / / opencv.org. All measured signal values ​​within an epoch were used to derive a set of 20 frequency and time domain features (Table 2). These were calculated using standard signal processing approaches and can be found in the example code [github.com / KumarLabJax / MouseSleep].

[0110] Training the classifier Due to the inherent dataset imbalance, i.e., more NREM epochs compared to REM sleep, an equal number of REM, NREM, and awake epochs were randomly selected to generate a balanced dataset. A cross-validation approach was used to evaluate the performance of the classifier. All epochs from 13 animals from the balanced dataset for training were randomly selected, and the unbalanced data from the remaining 4 animals were used for testing. The process was repeated 10 times to generate a range of accuracy measures. This approach allowed us to take advantage of training the classifier on balanced data while observing the performance on real imbalanced data.

[0111] Post-processing predictions To improve the prediction quality by integrating larger-scale temporal information, a hidden Markov model (HMM) approach was applied. The HMM model corrects the erroneous predictions made by the classifier by integrating the probability of sleep state transitions, and thus can obtain more accurate prediction results. The hidden states of the HMM model are the sleep stages, whereas the observables are obtained from the probability vector results from the XgBoost algorithm. The transition matrix was empirically calculated from the training set sequence of sleep states, followed by applying the Viterbi algorithm [Viterbi AJ (April 1967) IEEE Transactions on Information Theory vol. 13(2): 260-269] to estimate the sequence of stages, given the sequence of out-of-bag class votes of XgBoost. In this study, the transition matrix was a 3 × 3 matrix T = {S_ij}, where S_ij represented the probability of transition from state S_i to state S_j. T (Table 2).

[0112] Classifier performance analysis Performance was evaluated using measures of accuracy as well as several measures of classification performance: precision, recall, and F1 score. Precision was defined as the ratio of epochs classified by both the classifier and the human scorer for a given sleep stage to all epochs assigned by the classifier as that sleep stage. Recall was defined as the ratio of epochs classified by both the classifier and the human scorer for a given sleep stage to all epochs classified by the human scorer as that given sleep stage. F1 combined precision and recall to measure the harmonic mean of recall and precision. Means and standard deviations of accuracy and performance matrices were calculated from 10-fold cross-validation.

[0113] result Experimental design As shown in the schematic diagram in Figure 6A, the goal of the study described herein was to quantify the validity of using only video data to classify mouse sleep states. The experimental paradigm was designed to utilize the current gold standard for sleep state classification, EEG / EMG recordings, as labels to train and evaluate visual classifiers. In total, synchronized EEG / EMG and video data were recorded for 17 mice (24 h per animal). The data were divided into 10-s epochs. Each epoch was manually scored by a human expert. Concurrently, features from the video data that could be used in machine learning classifiers were designed. These features were constructed based on frame-by-frame measurements that described the visual appearance of the animals in individual video frames (Table 1). Signal processing techniques were then applied to the per-frame measurements to integrate the temporal information and generate a set of features for use in the machine learning classifier (Table 2). Finally, the human-labeled dataset was split by contributing individual animals to the training and validation datasets (80:20, respectively). The training dataset was used to train a machine learning classifier to classify 10 second epochs of video into three states: wake, sleep NREM, and sleep REM. The set of animals provided was used in the validation dataset to quantify the performance of the classifier. When separating the validation set from the training set, the entire animal data was provided to ensure that the classifier generalized well across animals, instead of learning to predict well only for the animals shown. [Table 1] [Table 2]

[0114] Features per frame Computer vision techniques were applied to extract detailed visual measurements of the mouse for each frame. The first computer vision technique used was segmentation of pixels related to mouse pixels and background pixels (Figure 6B). A segmentation neural network was trained as an approach that works well in light and dark conditions, as well as dynamic and challenging environments such as the moving bedding found in mouse arenas [Webb, JM and Fu, YH., Curr. Opin. Neurobiol. 69:19-24 (2021)]. Segmentation also allowed for the removal of EEG / EMG cables exiting the instrument on each mouse's head, thus not affecting visual measurements with information about head movement. The segmentation network predicted pixels that were only mouse, and therefore measurements were based only on mouse movement and not on movement of wires connected to the mouse's skull. Randomly sampled frames from all footage were annotated to achieve this high-quality segmentation and ellipse fit using the previously described network [Geuther, BQ et al., Commun. Biol. 2:124 (2019)] (Figure 6B). The neural network only required 313 annotated frames to achieve good performance in segmenting the mouse. An example of the performance of the segmentation network (not shown) on the original footage by coloring pixels predicted to be non-mouse in red and pixels predicted as mouse in blue. Following segmentation, 16 measurements from the neural network predicted segmentation that described the mouse's shape and position were calculated (Table 1). These included the mouse's major length, minor length, and ratio from the ellipse fit that described the mouse's shape. The mouse's position (x, y) and change in x, y (dx, dy) were extracted about the center of the ellipse fit. We also calculated area (m00), perimeter, and seven rotation-invariant Hu image moments (HU0-6) [Scammell, TE et al., Neuron. 93(4):747-765(2017)] of the segmented mice.Hu image moments are a numerical description of mouse segmentation through integrals and linear combinations of median image moments [Allada, R. and Siegel, JM Curr. Biol. 18(15):R670-R679 (2008)].

[0115] Time-frequency features Next, time- and frequency-based analysis was performed on each 10-s epoch using the per-frame features. This analysis allowed integration of the temporal information by applying signal processing techniques. As shown in Table 3, six time-domain features (kurtosis, mean, median, standard deviation, maximum, and minimum of each signal) and 14 frequency-domain features (kurtosis of power spectral density, skewness of power spectral density, average power spectral density at 0.1-1 Hz, 1-3 Hz, 3-5 Hz, 5-8 Hz, and 8-15 Hz, total power spectral density, maximum, minimum, mean, and standard deviation of power spectral density) were extracted for each per-frame feature of the epoch, resulting in 320 total features (16 measurements × 20 time-frequency features) for each 10-s epoch. [Table 3]

[0116] These spectral window features were visually inspected to determine whether they differed between the awake, REM, and NREM states. Figures 7A-B show representative epoch examples of m00 (area, Figure 7A) and wl_ratio (ratio of the width to length of the major and minor axes of the ellipse, Figure 7B) features, with different time and frequency domains for the awake, NREM, and REM states. The raw signals of m00 and wl_ratio showed clear oscillations in the NREM and REM states (left panels, Figures 7A and 7B), which can be seen in the FFT (middle panels, Figures 7A and 7B) and autocorrelation (right panels, Figures 7A and 7B). There was a single dominant frequency in the NREM epochs, and a broader peak in the REM. In addition, the FFT peak frequency changed slightly between NREM (2.6 Hz) and REM (2.9 Hz), and generally more regular and consistent oscillations were observed in the NREM epochs than in the REM epochs. Thus, initial investigation of the features revealed differences between sleep states, providing confidence that the features encoded useful metrics for use in a visual sleep classifier.

[0117] breathing rate Previous studies in both humans and rodents have demonstrated that respiration and movement differ between sleep stages [Stradling, JR et al., Thorax. 40(5):364-370 (1985); Gould, GA et al., Am. Rev Respir. Dis. 138(4):874-877 (1988); Douglas, NJ et al., Thorax. 37(11):840-844 (1982); Kirjavainen, T. et al., J. Sleep. Res. 5(3):186-194 (1996); Friedman, L. et al., J. Appl. Physiol. 97(5):1787-1795 (2004)]. When examining the characteristics of m00 and wl_ratio, we found a consistent signal between 2.5 and 3 Hz that appeared as a ventilation waveform (Figure 7A-7B). Inspection of the footage revealed changes in body shape and chest size due to breathing, which could be captured by the time-frequency features. To visualize this signal, a continuous wavelet transform (CWT) spectrogram was performed on the wl_ratio feature (Figure 8A, top panel). To summarize the data from these CWT spectrograms, the dominant signal of the CWT was identified (Figure 8A, respective bottom left panel) and a histogram of the dominant frequency of the signal was plotted (Figure 8A, respective bottom right panel). The mean and variance of the frequencies contained in the dominant signal were calculated from the corresponding histograms.

[0118] Previous studies have demonstrated that C57BL / 6J mice have a breathing rate of 2.5-3 Hz during the NREM state [Friedman, L. et al., J. Appl. Physiol. 97(5):1787-1795(2004); Fleury Curado, T. et al., Sleep. 41(8):zsy089 (2018)]. Investigations of long bouts of sleep (10 min) including both REM and NREM showed that the wl_ratio signal was more prominent in NREM than in REM, but was clearly present in both (Figure 8B). In addition, the REM state evoked a higher and more variable breathing rate than the NREM state, so the signal varied more within the 2.5-3.0 Hz range during the REM state. Low-frequency noise in this signal in the NREM state due to greater movement of the mouse, such as adjusting its sleep posture, was also observed. This suggested that the wl_ratio signal captured the visual movement of the mouse abdomen.

[0119] Breathing rate verification Genetic validation studies were performed to confirm that the signals observed in REM and NREM epochs for m00 and wl_ratio features were abdominal movements and correlated with respiratory rate. It has been previously demonstrated that C3H / HeJ mice have awake respiratory frequencies approximately 30% less than those of C57BL / 6J mice, ranging from 4.5 Hz vs. 3.18 Hz [Berndt, A. et al., Physiol. Genomics. 43(1):1-11 (2011)], 3.01 vs. 2.27 Hz [Groeben, H. et al., Br. J. Anaesth. 91(4):541-545 (2003)], and 2.68 Hz vs. 1.88 Hz [Vium, Inc., Breathing Rate Changes Monitored Non-Invasively 24 / 7. (2019)] for C57BL / 6J and C3H / HeJ, respectively. Uninstrumented C3H / HeJ mice (5 males, 5 females) were videotaped and the classic sleep / wake movement rule of thumb (distance traveled) [Pack, AI et al. Physiol. Genomics. 28(2):232-238(2007)] was applied to identify sleep epochs. Epochs were conservatively selected within the lowest 10% quantile for movement. Annotated C57BL / 6J EEG / EMG data were used to confirm that the movement-based cutoff could accurately identify sleep periods. Using the EEG / EMG annotation data of C57BL / 6J mice, this cutoff was found to identify primarily NREM and REM epochs (Figure 9A). The epochs selected in the annotated data consisted of 90.2% NREM, 8.1% REM, and 1.7% awake epochs. Thus, as expected, this mobility-based cutoff method correctly distinguished sleep / wake, but not REM / NREM. From these low-motion sleep epochs, we calculated the mean value of the dominant frequency in the wl_ratio signal. This measure was chosen because of its sensitivity to thoracic region movement. The distribution of the mean dominant frequency for each mouse was plotted, and a consistent distribution was observed across animals.For example, C57BL / 6J animals had a vibration range of mean frequency from 2.2 to 2.8 Hz, while C3H / HeJ animals ranged from 1.5 to 2.0 Hz, with the respiration rate of C3H / HeJ being approximately 30% lower than that of C57BL / 6J. This is a statistically significant difference between the two strains, i.e., C57BL / 6J and C3H / HeJ (p<0.001, Figure 9B), and is in a similar range to that previously reported [Berndt, A. et al., Physiol. Genomics. 43(1):1-11(2011); Groeben,H.et al.,Br. J. Anaesth 91(4):541-545 (2003); Vium, Inc., Breathing Rate Changes Monitored Non-Invasively 24 / 7. (2019)]. Therefore, using this genetic validation method, it was concluded that the observed signal strongly correlates with respiratory rate.

[0120] In addition to overall changes in breathing frequency due to genetics, breathing during sleep has been shown to be more organized and less variable during NREM than REM in both humans and rodents [Mang, GM et al., Sleep. 37(8):1383-1392 (2014); Terzano, MG et al., Sleep. 8(2):137-145 (1985)]. It was hypothesized that the detected breathing signal would show greater variation in REM epochs than in NREM epochs. The EEG / EEG annotated C57BL / 6J data was examined to determine if there was a change in the variability of the CWT peak signal in epochs across REM and NREM states. Using only the C57BL / 6J data, epochs were split by NREM and REM states to observe the variability of the CWT peak signal (Figure 9C). The NREM state showed a smaller standard deviation of this signal, while the REM state had a broader and higher peak. The NREM states appeared to contain multiple distributions, possibly indicating a subdivision of NREM sleep states [Katsageorgiou, VM. et al., PLoS Biol. 16(5):e2003663 (2018)]. To confirm that this odd shape of the NREM distribution was not an artifact of combining data from multiple animals, we plotted the data for each animal, and each mouse showed an increase in standard deviation from the NREM to the REM state (Figure 9D). Individual animals also showed this long-tailed NREM distribution. Both of these experiments indicated that the observed signal was a respiratory rate signal. These results suggested good classifier performance.

[0121] classification Finally, a machine learning classifier was trained to predict sleep states using the 320 visual features. For validation, all data from the animals were provided to avoid any bias that may be introduced by correlated data in the footage. For calculation of training and test accuracy rates, 10-fold cross-validation was performed by shuffling which animals were provided. A balanced dataset was created as described in Materials and Methods above herein, and multiple classification algorithms, including XgBoost, Random Forest, MLP, Logistic Regression, and SVD, were compared. It was observed that performance varied widely between classifiers (Table 4). Both XgBoost and Random Forest achieved good accuracy on the provided test data. However, the Random Forest algorithm achieved a 100% training accuracy rate, indicating that it overfitted the training data. Overall, the best-performing algorithm was the XgBoost classifier. [Table 4]

[0122] The transitions between awake, NREM, and REM states are not random but generally follow expected patterns. For example, awake typically transitions to NREM, which then transitions to REM sleep. Hidden Markov models are ideal candidates for modeling dependencies between sleep states. The transition probability matrix and occurrence probability for a given state are learned using training data. By adding the HMM model, we observed a 7% improvement in the overall classifier accuracy rate, from 0.839 + / - 0.022 to 0.906 + / - 0.021 (Figure 10A, +HMM).

[0123] To improve classifier performance, Hu moment measurements were adopted from the segmentation to include in the input features for classification [Hu, MK. IRE Trans Inf Theory. 8(2):179-187 (1962)]. These image moments were a numerical description of the mouse segmentation through integrals and linear combinations of median image moments. The addition of Hu moment features achieved a small increase in overall accuracy and increased classifier robustness by reducing the variability of cross-validation performance from 0.906+ / -0.021 to 0.913+ / -0.019 (Figure 10A, +Hu moments).

[0124] EEG / EMG scoring was performed by trained human experts, but there is often disagreement between trained annotators [Pack, AI et al. Physiol. Genomics. 28(2):232-238 (2007)]. In fact, the two experts generally agreed only 88-94% of the time for REM and NREM [Pack, AI et al. Physiol. Genomics. 28(2):232-238(2007)]. We used a recently published machine learning method to score the EEG / EMG data to complement the data from the human scorers [Miladinovic, D. et al., PLoS Comput. Biol. 15(4):e1006968 (2019)]. We compared the SPINDLE annotations with the human annotations and found agreement for 92% of all epochs. Only epochs on which both the human and machine-based methods agreed were then used as indicators for visual classifier training. Training the classifier using only SPINDLE and human-matched epochs increased the accuracy rate by an additional 1% (Figure 10A, + filter annotation). Thus, the final classifier achieved a three-state classification accuracy rate of 0.92 + / - 0.05.

[0125] The classification features used were investigated to determine which were most important. Mouse region and movement measures were identified as the most important features (Figure 10B). While not intended to be limiting, it is believed that this result was observed because movement is the only feature used in the binary sleep-wake classification algorithm. In addition, three of the top five features were low frequency (0.1-1.0 Hz) power spectral density (Figure 7A and Figure 7B, FFT columns). Furthermore, it was also observed that wake epochs had the most power in low frequencies, REM had low power in low frequencies, and NREM had the least power in low frequency signals.

[0126] Good performance was observed using the best performing classifier (Figure 10C). The rows of the matrix shown in Figure 10C represent the sleep states assigned by the human scorer, and the columns represent the stages assigned by the classifier. Wake had the highest accuracy rate among the classes with an accuracy rate of 96.1%. By looking off the diagonal of the matrix, it can be seen that the classifier performed better in distinguishing wake from any sleep state than between sleep states, indicating that distinguishing REM from NREM was a difficult task.

[0127] The final classifier achieved an overall accuracy rate of 0.92 + / - 0.05 on average. The prediction accuracy rate for the waking phase was 0.97 + / - 0.01, with a mean matching recall of 0.98. The prediction accuracy rate for the NREM phase was 0.92 + / - 0.04, with a mean matching recall of 0.93. The prediction accuracy rate for the REM phase was 0.88 + / - 0.05, with a mean matching recall of 0.535. The lower matching recall rate for REM was due to a very small proportion of epochs labeled as REM phases (4%).

[0128] In addition to prediction accuracy, performance metrics including precision, recall, and F1 score were measured to evaluate the models from 10-fold cross-validation (Figure 10D). Considering the imbalanced data, precision-recall was a better indicator for classifier performance [Powers, DMW, J. Mach. Learn. Technol. 2(1):37-63 (2011) ArXiv:2010.16061;Saito, T. and Rehmsmeier, M. PLoS ONE. 10(3):e0118432 (2015)]. Precision measured the proportion of correctly predicted positive items, while recall measured the proportion of correctly identified actual positives. F1 score was the weighted average of precision and recall.

number

number

number

[0129] The final classifier performed well for both awake and NREM states. However, the poorest performance was noted for the REM phase, with a precision of 0.535 and an F1 of 0.664. The majority of misclassified phases were between NREM and REM. Because the REM state was a minority class (only 4% of the dataset), even a relatively small false positive rate would result in a large number of false positives that would overwhelm the rare true positives. For example, 9.7% of REM periods were incorrectly identified as NREM by the visual classifier, and 7.1% of predicted REM periods were in fact NREM (Figure 10C). Although these misclassification errors seem small, the imbalance between REM and NREM may disproportionately affect the precision of the classifier. Despite this, the classifier was also able to correctly identify 89.7% of the REM epochs present in the validation dataset.

[0130] Within the context of other existing alternatives for EEG / EMG recording, the model outperformed. Table 5 compares the performance of each of the previously reported models with that of the classifier model described herein. Note that each of the previously reported models used a different dataset with different characteristics. In particular, the Piezo system was evaluated on a balanced dataset that may show a higher accuracy rate due to a lower chance of false positives. The classifier approach developed herein outperformed all approaches for wakefulness and NREM state prediction. REM prediction was a more challenging task for all approaches. Of the machine learning approaches, the model described herein achieved the highest accuracy rate. Figure 11A and Figure 11B show a visual performance comparison of our classifier and manual scoring by a human expert (hypnograms). The x-axis is time, consisting of consecutive epochs, and the y-axis corresponds to the three stages. For each subfigure, the top panel represents the human scoring results, and the bottom panel represents the classifier scoring results. The hypnogram shows accurate transitions between stages along with the frequency of isolated false positives (Figure 11A). We also plot visual and human scoring for a single animal over 24 hours (Figure 11B). Raster plots show better overall correlation between state classifications (Figure 11B). All C57BL / 6J animals are then compared between human EEG / EMG scoring and visual scoring (Figure 11C,D). High correlation is observed across all states, concluding that our visual classifier scoring results are consistent with the human scores. [Table 5]

[0131] We also attempted various data augmentation approaches to improve the classifier performance. The proportions of different sleep states in a 24-hour period were significantly unbalanced (awake 48%, NREM 48%, and REM 4%). Typical augmentation techniques used for time series data include jittering, scaling, rotation, reordering, and cropping. These methods can be applied in combination with each other. It has been previously shown that the classification accuracy rate can be increased by augmenting the training data by combining four data augmentation techniques [Rashid, KM and Louis, J. Adv Eng Inform. 42:100944 (2019)]. However, since the features extracted from the time series were dependent on the spectral composition, it was decided to augment the size of the training dataset using a dynamic time warping-based approach [Fawaz, HI, et al., arXiv :1808:02455] to improve the classifier. After data augmentation, the size of the dataset increased by about 25% (from 14K epochs to 17K epochs). It was observed that adding more data through the augmentation algorithm resulted in a decrease in prediction accuracy. The average predicted averages for awake, NREM, and REM states were 77%, 34%, and 31%. Without wishing to be bound by any particular theory, the performance after data augmentation may be due to the introduction of more noise from the REM state data and a decrease in the performance of the classifier. Performance was shown with 10-fold cross-validation. The results of this data augmentation application are shown in Figure 12 and Table 6 (Table 6 shows the numerical results of data augmentation). This data augmentation approach did not improve the classifier performance and was not pursued further. Overall, the visual sleep state classifier was able to accurately identify sleep states using only visual data. The inclusion of HMM, Hu moments, and highly accurate labels improved performance, while data augmentation using dynamic time warping and motion amplification did not improve performance. [Table 6]

[0132] Consideration Sleep disorders are a hallmark of many diseases, and high-throughput studies in model organisms are important for the discovery of new therapeutics [Webb, JM and Fu, YH., Curr. Opin. Neurobiol. 69:19-24 (2021); Scammell,TE et al.,Neuron. 93(4):747-765 (2017); Allada, R. and Siegel, JM Curr. Biol. 18(15):R670-R679 (2008)]. Sleep studies in mice are difficult to conduct at scale due to the time investment for performing surgery, recovery time, and scoring the recorded EEG / EMG signals. The system described herein provides a low-cost alternative to EEG / EMG scoring of mouse sleep behavior, enabling researchers to conduct larger-scale sleep experiments that would previously have been cost-prohibitive. Previous systems have been proposed to conduct such experiments, but have only been shown to adequately distinguish between wakefulness and sleep states. The system described herein builds on these approaches and is also able to distinguish sleep states into REM and non-REM states.

[0133] The system described herein achieves sensitive measurements of mouse movement and posture during sleep. The system has been shown to observe features that correlate with mouse breathing rate using only visual measurements. Previously published systems that can achieve this level of sensitivity include plethysmography [Bastianini, S. et al., Sci. Rep. 7:41698 (2017)] or piezo systems [Mang, GM et al., Sleep. 37(8):1383-1392 (2014); Yaghouby, F., et al., J. Neurosci. Methods. 259:90-100 (2016)]. In addition, it has been shown herein that based on the features used, the novel system may be able to identify sub-clusters of non-REM sleep epochs, which may further unravel the structure of mouse sleep.

[0134] In conclusion, the high-throughput, non-invasive, computer vision-based method for sleep state determination in mice described herein above is useful to society.

[0135] Equivalent While several embodiments of the invention have been described and illustrated herein, those skilled in the art can readily envision various other means and / or structures for performing the functions and / or obtaining the results and / or one or more advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the invention. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the particular application or applications for which the teachings of the invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. It will thus be understood that the foregoing embodiments have been presented by way of example only, and that within the scope of the appended claims and their equivalents, the invention may be practiced otherwise than as specifically described and claimed. The invention is directed to each individual feature, system, article, material, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, and / or methods is included within the scope of the present invention, if such features, systems, articles, materials, and / or methods are not mutually inconsistent. It will be understood that all definitions defined and used herein control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings for the defined terms.

[0136] As used herein in the specification and claims, the indefinite articles "a" and "an" will be understood to mean "at least one" unless clearly indicated to the contrary. As used herein in the specification and claims, the term "and / or" will be understood to mean "either or both" of the elements so conjoined, i.e., "either or both" of elements that are conjunctively present in some cases and disjunctively present in other cases. Unless expressly indicated to the contrary, other elements, whether related or unrelated to the specifically identified elements, may optionally be present other than the elements specifically identified by the term "and / or."

[0137] As used herein, conditional language, particularly "can," "could," "might," "may," "eg," and the like, is generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context in which it is used. Thus, such conditional language is not generally intended to imply that the features, elements, and / or steps are required in any manner with respect to one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps are included in or performed in any particular embodiment, with or without other input or prompting. The terms "comprising," "including," "having," and the like are synonymous and used in an inclusive, open-ended manner and do not exclude additional elements, features, acts, operations, and the like. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), so that, for example, to connect a list of elements, the term "or" means one, some, or all of the elements in the list.

[0138] All references, patents, and patent applications and publications cited or referred to in this application are hereby incorporated by reference in their entirety.

Claims

1. A method executed by a computer, comprising: receiving video data representing a target video; using the video data to determine a plurality of features corresponding to the target; and using the plurality of features to determine sleep state data about the target.

2. The method executed by a computer according to claim 1, further comprising using a machine learning model to process the video data to determine segmentation data indicating a first set of pixels corresponding to the target and a second set of pixels corresponding to the background.

3. The method executed by a computer according to claim 2, further comprising processing the segmentation data to determine ellipse fitting data corresponding to the target.

4. The method executed by a computer according to claim 2, wherein determining the plurality of features includes processing the segmentation data to determine the plurality of features.

5. The method executed by a computer according to claim 1, wherein the plurality of features include a plurality of visual features for each video frame of the video data.

6. The method further includes determining a time domain feature for each visual feature of the plurality of visual features, and determining the time domain feature includes determining one of kurtosis data, mean data, median data, standard deviation data, maximum data, and minimum data. The method executed by a computer according to claim 5, wherein the plurality of features include the time domain features.

7. The method further includes determining a frequency domain feature for each visual feature of the plurality of visual features, and determining the frequency domain feature includes determining one of the kurtosis of the power spectral density, the skewness of the power spectral density, the average power spectral density, the total power spectral density, the maximum data of the power spectral density, the minimum data, the mean data, and the standard deviation. The method executed by a computer according to claim 5, wherein the plurality of features include the frequency domain features.

8. determining a time domain feature for each of the plurality of features; and determining a frequency domain feature for each of the plurality of features. Using a machine learning classifier to process the time-domain features and the frequency-domain features to determine the sleep state data, further comprising the method executed by the computer according to claim 1.

9. Using a machine learning classifier to process the plurality of features to determine a sleep state for a video frame of the video data, wherein the sleep state is one of a wake state, a REM sleep state, and a non-REM sleep state, the method executed by the computer according to claim 1.

10. The sleep state data indicates one or more durations of a sleep state, a wake state, a REM state, and a non-REM state and / or one or more of the frequency intervals, as well as one or more changes in the sleep state, the method executed by the computer according to claim 1.

11. Determining a plurality of body regions of the subject using the plurality of features, each body region of the plurality of body regions corresponding to a video frame of the video data, and Determining the sleep state data based on changes in the plurality of body regions in the video, further comprising the method executed by the computer according to claim 1.

12. Determining a plurality of width-to-length ratios using the plurality of features, each width-to-length ratio of the plurality of width-to-length ratios corresponding to a video frame of the video data, and Determining the sleep state data based on changes in the plurality of width-to-length ratios in the video, further comprising the method executed by the computer according to claim 1.

13. Determining the sleep state data is Detecting a transition from a non-REM state to a REM state based on a change in the body region or body shape of the subject, wherein the change in the body region or body shape is a result of muscle weakness, the method executed by the computer according to claim 1.

14. Determining a plurality of width-to-length ratios for the subject, each width-to-length ratio of the plurality of width-to-length ratios corresponding to a video frame of the video data, and Determining time-domain features using the plurality of width-to-length ratios, and Determining frequency-domain features using the plurality of width-to-length ratios, wherein the time-domain features and the frequency-domain features represent the movement of the abdomen of the subject, and Further comprising determining the sleep state data using the time domain features and the frequency domain features, the method executed by a computer according to claim 1.

15. The method executed by a computer according to claim 1, wherein the video captures the subject in a natural state of the subject, and the natural state of the subject includes that there is no invasive detection means within or on the subject, and the invasive detection means includes one or both of an electrode attached to the subject and an electrode inserted into the subject.

16. The method executed by a computer according to claim 1, wherein the video is a high-resolution video.

17. Processing the plurality of features using a machine learning classifier to determine a plurality of sleep state predictions for each video frame of the video data; Further comprising processing the plurality of sleep state predictions using a transition model to determine a transition between a first sleep state and a second sleep state, wherein the transition model is a hidden Markov model, the method executed by a computer according to claim 1.

18. The video is a video of two or more subjects including at least a first subject and a second subject, and the method includes: Processing the video data to determine first segmentation data indicating a first set of pixels corresponding to the first subject; Processing the video data to determine second segmentation data indicating a second set of pixels corresponding to the second subject; Using the first segmentation data to determine a first plurality of features corresponding to the first subject; Using the first plurality of features to determine first sleep state data for the first subject; Using the second segmentation data to determine a second plurality of features corresponding to the second subject; Further comprising using the second plurality of features to determine second sleep data for the second subject, the method executed by a computer according to claim 1.

19. The method executed by a computer according to claim 1, wherein the subject is a rodent, optionally a mouse.

20. The method executed by a computer according to claim 1, wherein the subject is a genetically engineered subject.